What is operational resilience?
Glassbreak Team · Published 2026-08-23
Operational resilience is the ability of a firm to keep delivering its important business services through disruption, and to do so within a level of disruption it has decided in advance is tolerable. What distinguishes it from the resilience work that came before is where the assessment sits: not on whether a control was implemented, but on whether an outcome held. A firm with exemplary documentation that cannot show its payments service stays inside its stated impact tolerance during a severe but plausible scenario has not met the standard.
The regime, briefly
In the UK the framework arrived through FCA Policy Statement PS21/3, which inserted SYSC 15A into the FCA Handbook, alongside the PRA's Supervisory Statement SS1/21 and Policy Statement PS6/21. The rules came into force on 31 March 2022, with a transitional period running to 31 March 2025 by which firms were expected to have performed the mapping and testing necessary to remain within their impact tolerances, and to have remediated where they could not.
Scope is by firm type rather than by activity: banks, building societies, PRA-designated investment firms, insurers and recognised investment exchanges on the PRA side; enhanced-scope SM&CR firms, payment and e-money institutions and others on the FCA side. Firms outside those categories are not subject to SYSC 15A, however useful they may find the method.
The EU's parallel regime is DORA (Regulation (EU) 2022/2554), which has applied since 17 January 2025 and covers ICT risk management for EU financial entities. The vocabularies differ — DORA speaks of critical or important functions and ICT third-party risk rather than important business services and impact tolerances — but the underlying logic of identifying what matters, mapping what it depends on, and testing the result is closely related. Firms operating in both jurisdictions generally run one programme and map it to both sets of language.
The four obligations that do the work
Identify important business services. These are services delivered to an external end user or to the market, the disruption of which could cause intolerable harm to consumers or risk to market integrity. The unit is the service, not the system and not the business line — a distinction firms routinely get wrong on the first attempt by listing applications.
Set an impact tolerance for each. A maximum tolerable duration of disruption, expressed as a number the board owns. Its value lies in being falsifiable.
Map the dependencies. For each important business service, the people, processes, technology, facilities, information and third parties it relies on. Mapping is where most of the discomfort lives, because it is the step that makes concentration risk and single points of failure legible.
Test against severe but plausible scenarios. Not worst-case fantasy and not the routine outage the service already survives weekly — a scenario severe enough to be uncomfortable and plausible enough that a supervisor would recognise it.
Firms then document all of this in a self-assessment, which is the artefact a supervisor is most likely to ask for first.
Where third parties enter
Mapping tends to reveal that important business services depend on suppliers, and that several services depend on the same supplier. The UK addressed this concentration directly: the Financial Services and Markets Act 2023 gave HM Treasury the power to designate critical third parties, with the supervisory regime for those designated entities operated by the Bank of England, PRA and FCA. DORA does something structurally similar through its designation of Critical ICT Third-Party Providers under Article 31.
Most suppliers are not designated. They appear instead as ordinary third-party dependencies inside a regulated firm's mapping, and the firm carries the resilience obligation for the service the supplier supports. This is why supplier questionnaires increasingly ask about recovery paths rather than only about certifications: a certification describes the supplier's control environment, while the firm needs to know whether the service stays inside tolerance when the supplier is degraded.
The dependency that mapping tends to surface late
A recurring finding is that the recovery path for an important business service runs through the firm's own access infrastructure. If restoring the service requires an administrator to authenticate, and authentication depends on an identity provider that is itself part of the disruption, the recovery plan contains a loop. The same is true where the credentials needed during an incident live in a password manager that is unavailable, or in the memory of one person who cannot be reached.
None of the regimes name a mechanism for solving this, and it would be wrong to claim otherwise. What they require is that the outcome holds under a severe but plausible scenario — and "the identity provider is unavailable" is a scenario supervisors recognise as plausible, because it has happened to large providers repeatedly. A firm whose scenario testing assumes its own authentication always works has not tested the scenario that most often makes recovery slow.
Glassbreak's position in this picture is narrow and worth stating plainly: it is not a regulated firm and is not directly subject to SYSC 15A or DORA. Where a regulated firm uses it, it appears in that firm's mapping as a third-party dependency supporting the recovery path for whichever services depend on emergency access to credentials. Our own third-party posture, including the DORA Article 28-30 contractual position, is published at /trust/dora, and the controls behind it at /trust/security-questionnaire.
What good evidence looks like
Supervisors ask for artefacts rather than assurances. In practice that means a documented set of important business services with rationale for inclusion and exclusion, board-approved impact tolerances, mapping at sufficient granularity to identify single points of failure, records of scenario tests including the ones that failed, and a remediation plan with dates for the gaps those tests exposed.
The tests that failed matter more than the ones that passed. A self-assessment showing only successful scenarios invites the question of whether the scenarios were severe enough — which is usually the harder question to answer.
Frequently asked questions
- How is operational resilience different from business continuity?
- Business continuity planning generally asks whether the firm can recover a system or a site, and measures success against recovery time objectives for those assets. Operational resilience inverts the frame: it starts from the service delivered to the end user or the market, assumes disruption will happen rather than planning to prevent it, and asks how much disruption to that service is tolerable before it causes intolerable harm. A firm can hold a complete set of BCP documents and still fail an operational resilience assessment if nobody has established that the service — as opposed to its components — stays within tolerance.
- What is an impact tolerance, in practice?
- It is a maximum tolerable level of disruption to a specific important business service, expressed in terms a board can be held to — most commonly a duration, sometimes with volume, timing or data-loss dimensions attached. The point of expressing it as a hard number is that it becomes falsifiable: scenario testing either demonstrates the service stays inside it or it does not. Regulators have been explicit that a tolerance set so generously that it can never be breached is not doing its job.
- Does operational resilience require a break-glass procedure?
- No regime names break-glass access as a required control. The requirement is to remain within impact tolerance during severe but plausible disruption, and mapping tends to surface the dependencies that would prevent that. Where a firm's mapping shows an important business service depends on an identity provider, a password manager, or a small number of individuals holding credentials, the recovery path for that dependency becomes part of the resilience question — because a plan that cannot reach the systems it needs to recover has not demonstrated the outcome the regime asks for.
- Does this apply to firms outside financial services?
- The UK and EU regimes described here are financial-services regimes and apply by firm type, not to businesses generally. The underlying method — identify the services that matter, decide how much disruption is tolerable, map what they depend on, then test — is not specific to financial regulation, and appears in adjacent form in NIS2 for essential and important entities in the EU. Firms outside scope sometimes adopt the method voluntarily, but should not describe themselves as subject to SYSC 15A or DORA when they are not.