Scenario Analysis and Incident Management for Operational Resilience

Resilience improves when forward-looking scenarios meet disciplined learning from incidents. Learn to connect both so every event strengthens your framework.
5 min read time

Operational resilience improves fastest when firms combine forward-looking testing with disciplined learning from real events. Scenario analysis helps organisations imagine severe but plausible disruptions before they happen. Incident management captures what actually happened, what failed and what needs to change next. Together, they create a powerful improvement loop.

These two disciplines are frequently run as entirely separate workstreams, often owned by different teams with limited communication between them. That separation is a missed opportunity, because each discipline makes the other significantly more valuable. Scenarios that have never been tested against real incident experience tend to drift toward generic, theoretical exercises. Incident reviews that never feed back into the scenario library tend to repeat the same blind spots, because the organisation never asks whether the event it just experienced was one it had genuinely anticipated or one that exposed a gap in its thinking.

Why scenario analysis matters

Scenario analysis is not about predicting the future with precision. It is about testing management assumptions under pressure. Firms should identify severe but plausible events, estimate impact, recovery time and dependencies, then assess whether controls, contingency plans and escalation routes are strong enough. Larger institutions may link this to ICAAP and stress testing; smaller firms can still use it to test resilience qualitatively.

The value of a scenario exercise rarely comes from getting the precise numbers right. A scenario that estimates financial impact within a tight, confident range is less useful than one that genuinely stress-tests whether the organisation's people, processes and systems would actually hold up under severe pressure. What matters is whether the exercise surfaces real gaps: a contingency plan that assumes a key supplier would still be reachable during the very disruption the supplier is causing, an escalation route that depends entirely on one individual who might themselves be affected by the scenario, a recovery time assumption that has never actually been tested against the firm's real technical capability. These are the kinds of insight a well-run scenario exercise produces, and they are valuable regardless of whether the underlying financial estimate proves precisely accurate if the scenario ever materialises for real.

Build scenarios that leaders can use

Choose plausible disruptions

Cyber-attacks, third-party failure, absenteeism events, technology breakdowns and process failures are all obvious candidates. Strong scenarios are grounded in business realities and mapped to important services, customer impact and recovery objectives. Scenarios should be severe but plausible and can often be inspired by past industry events. In fact, some firms conduct these exercises after a regulatory sanction or outage at another firm becomes public.

The most useful scenarios are specific rather than generic. "A cyber-attack occurs" is too vague to test anything meaningfully. "A ransomware attack encrypts the core policy administration system during the final week of a regulatory reporting deadline, with the primary IT lead on leave" gives the organisation something concrete to work through: who would actually be making decisions, what manual workarounds genuinely exist, how long full restoration would realistically take given the firm's actual backup and recovery capability, and what the customer and regulatory impact of that realistic timeline would be. Building scenarios around what has actually happened to comparable firms, drawing on publicly reported incidents and regulatory enforcement actions in the sector, tends to produce more credible and more useful exercises than purely hypothetical ones, because participants find it harder to dismiss a scenario as unrealistic when a near-identical event has demonstrably occurred elsewhere.

Quantify impact and assumptions

For each scenario, estimate the potential impact, including financial loss, reputational damage, operational disruption and regulatory sanction, and the time to recovery. This may involve conducting workshops with different people within the firm, financial modelling, and gathering and analysing publicly available data, such as internal loss data or relevant external statistics.

The governance value of scenario analysis is often less about the exact model output and more about the quality of the discussion and the resulting action plans. A workshop that brings together operations, technology, customer service, compliance and senior management to genuinely work through what would happen in a severe scenario tends to surface assumptions and dependencies that no individual function would have identified on its own. Documenting those assumptions, and the decisions the exercise leads to, matters as much as the headline impact figure, both for internal learning and because supervisors increasingly expect to see the reasoning behind a firm's resilience conclusions, not just the conclusions themselves.

Larger firms typically integrate this into formal capital adequacy assessment processes, where operational scenarios feed directly into capital calculations. Smaller firms can use the same underlying discipline more qualitatively, for example to genuinely test whether a disaster recovery plan would hold up, without needing the same level of financial modelling sophistication.

Build an incident management lifecycle that improves the framework

Record and classify incidents early

A central register should capture incidents and near misses quickly, with classification, severity and ownership. Near misses are especially valuable because they often show control weakness before material harm occurs.

Near misses deserve particular attention because they represent a genuinely undervalued source of intelligence in most organisations. A near miss is, by definition, an event where something went wrong but the consequences were avoided, often through luck rather than design. Capturing and analysing these events systematically gives the organisation visibility of control weaknesses before they cause material harm, which is precisely the kind of forward-looking signal that effective operational risk management is meant to provide. Firms that only log incidents with actual material impact are missing a substantial proportion of the available learning, because for every incident that causes real harm, there are typically several near misses that, with slightly different timing or luck, could have escalated into something far more serious.

Escalate, investigate and remediate

The right incidents need clear escalation routes, regulatory trigger assessment, root-cause analysis and tracked remediation. Under DORA, incident reporting standards are now more formalised in the EU, and UK regulators have also updated reporting expectations. That makes evidence quality and timing more important than ever.

Under the EU Digital Operational Resilience Act, firms must report major ICT-related incidents to their competent authority where defined thresholds are exceeded, with technical standards requiring an initial notification within four hours of detection and detailed follow-up reporting within twenty-four hours. In the UK, the FCA and PRA impose comparable incident reporting requirements. These timelines are genuinely demanding, and they make two things essential: a clear, pre-agreed definition of what constitutes a reportable incident, based on factors such as customer impact and volume, and an escalation process fast enough to actually meet the notification deadlines when something serious occurs. An incident process that takes several days to determine whether an event meets the reporting threshold is not fit for purpose under these timelines.

Root cause analysis is where the genuine learning happens, and it deserves more rigour than firms often apply under the pressure of resolving the immediate incident. Structured techniques, such as the "five whys" method or a fishbone diagram, help investigators move past the immediate, surface-level cause toward the underlying control weakness or process gap that actually allowed the incident to occur. Identifying that a payment was processed incorrectly because a member of staff made an error is rarely the useful conclusion; identifying why the control that should have caught that error did not function as designed is the conclusion that actually prevents recurrence.

Feed lessons learned back into scenarios and controls

Periodically review incident trends and root cause analysis insights at risk or audit committee meetings. Document lessons learned and integrate them into RCSAs, control enhancements, or scenario updates. If the organisation learns the same lesson twice, the process is recording events without improving the framework.

This final step is where the connection between incident management and scenario analysis becomes most valuable, and it is also where many firms fall short. An incident that was not anticipated by the existing scenario library represents a genuine gap in the organisation's forward-looking risk thinking, and that scenario library should be updated accordingly. A recurring theme across several incidents, even relatively minor ones individually, often points toward a systemic control weakness that deserves attention well before it produces a more serious event. The test of whether this feedback loop is genuinely functioning is simple, if uncomfortable to apply honestly: has the organisation ever experienced the same type of incident twice, with the same underlying root cause, within a relatively short period? If so, the lessons learned process is not actually changing anything.

Conclusion

Scenario analysis and incident management belong together because one tests what could happen and the other explains what did happen. Firms that connect the two are usually better prepared, faster to recover and more credible when they say resilience is being actively managed rather than passively reported.

Building this connected approach, including how to structure scenario workshops, design an incident lifecycle that meets current regulatory reporting timelines, and ensure lessons genuinely feed back into the framework, is covered in detail in our full whitepaper. Download the complete Operational Risk whitepaper for the full approach to resilience planning.

Next Steps

Turn scenarios and incidents into stronger resilience