Business Continuity Management: Building a BCMS That Holds

Business continuity management does not begin with emergency plans. It begins with the question of how long a business process may be down. This guide runs from the foundation of the business impact analysis through defensible time targets, continuity strategies and exercises, to the question of how the same evidence can be used against ISO 22301, BSI-Standard 200-4 and the continuity duties in NIS2 and DORA.

Pillar guideRisikomanagement13 min readLast reviewed:

What does BCM do, and where does it start?

Every protective measure can fail. Business continuity management picks up exactly where prevention stops: it does not answer how an outage is avoided, but how long the organisation can withstand one, which services it nonetheless delivers in the meantime, and how it returns to normal operation. Its reference points are business processes and the services customers, regulators and the supply chain expect, not individual systems.

That distinction sounds academic and is not. An IT restart plan describes how a server comes back up. A continuity plan describes how order intake keeps working while the server does not. Disaster recovery is therefore one output of BCM, not its purpose. Treat the two as the same thing and the first realistic exercise produces the classic finding: the system is running again and the business process is still stopped, because the paper forms, the approval rights or a service provider are missing.

BCM is equally not a sub-chapter of the information security management system. An ISMS manages risks to confidentiality, integrity and availability and works largely preventively; BCM takes over the case that happens anyway. Both draw on the same groundwork (the process map, the asset and supplier inventory, the protection requirement for availability) and feed each other, but neither substitutes for the other.

PhaseContentOutput
InitiationPolicy, scope, roles, mandate from the topBCM policy and terms of reference
AnalysisBusiness impact analysis, supporting risk analysisTime-critical processes, time targets, resource requirements
StrategyChoice of response option per resource classReasoned continuity strategies
ImplementationPreventive measures, emergency organisation, plans, alertingDocumented and available procedures
MaintenanceTests, exercises, reviews, updates on changeDemonstrated effectiveness

The cycle follows the logic of other management systems. The difference is in the last phase: it is not optional. A plan that has never been exercised is a hypothesis.

Why is the business impact analysis the foundation?

The business impact analysis (BIA) is the step that determines whether a continuity concept holds. The BIA asks not what causes an outage, which is the risk analysis's job, but what an outage costs, from what point it becomes intolerable, and which resources have to be available again for the process to continue. Everything else builds on its result. Without a BIA, every strategy is a guess and every target figure is a number with no provenance.

The sequence is essentially the same in every established methodology:

  1. Capture and delimit the business processes, starting from the services owed to customers, regulators and the supply chain, not from the system landscape.
  2. Assess how harm develops over time, staged by duration of outage rather than expressed as a single value.
  3. Derive the time targets and agree them with the process owners.
  4. Record the resources and dependencies required, including the external ones.
  5. Consolidate the results, resolve the contradictions, and have the leadership sign them off.

The last step is the one most often skipped and the one that matters most. A BIA contains commitments that cost money to keep: redundancy, standby capacity, on-call rotas, contractually assured recovery times from providers. Without sign-off by executive management, what you have is not a concept but a wish list.

Harm is assessed on several dimensions. The usual ones are financial consequences such as lost revenue, contractual penalties and the extra cost of contingency operation; legal and regulatory consequences such as missed notification or retention duties; degradation of service delivery; reputational damage with customers, regulators and the public; and danger to people and the environment. The time axis is what matters: after a few hours, then after a day, after several days, after a week. Only that progression shows where harm rises sharply, and tolerability ends there.

The substantive assessment has to come from the business. If IT fills in the questionnaire, what emerges is a system list rather than a process analysis. The analysis equally loses its purpose if every process turns out to be time-critical: a prioritisation that excludes nothing has not prioritised.

No fixed update cycle is prescribed anywhere. A regular review at a defined interval works, supplemented by event-driven updates on material change: new or rebuilt processes, a site move, a change of critical provider, acquisitions, or larger IT migrations. A BIA left untouched after its first pass describes, two years later, a company that no longer exists, while still carrying the time targets everyone will be measured against when something happens.

What do RTO, RPO and maximum tolerable downtime mean?

The BIA produces the figures everything else is later measured against. In an international setting you will meet two vocabularies for them, because German practice uses its own set and you will run into it in any audit involving a German site or authority.

MeasureInternational termGerman term (BSI-Standard 200-4)What it tells you
Threshold of harmMTPDmaximal tolerierbare Ausfallzeit (MTA)How long may the process be down at most before intolerable consequences arise?
Planning target, contingency operationRTOWiederanlaufzeit (WAZ)By when does the process have to be working again in contingency mode?
Planning target, normal operation—Wiederherstellungszeit (WVZ)By when is regular operation restored?
Data lossRPOmaximal tolerierbarer DatenverlustWhich state of the data may be lost when the event occurs?
Level of serviceMBCOMindestbetriebsniveauWhat share of the service actually has to be delivered in contingency mode?

The minimum business continuity objective (MBCO) is the figure most often left out, and it is also the one that makes a continuity strategy affordable in the first place. Leave it out and the full level of service silently becomes the recovery target, so provision ends up sized for a state nobody actually demands during contingency operation. Setting it is a business decision for the process owner, not something IT determines.

There are no normative target values for any of these. Neither ISO 22301 nor BSI-Standard 200-4 prescribes numbers; both require the organisation to derive its values, justify them and meet them. Any figure lifted from a template creates a commitment with nothing behind it.

Three consistency checks decide whether the numbers are worth anything:

  • The recovery time objective has to sit well below the maximum tolerable period of disruption. If the two are equal, any delay in the response is already intolerable harm.
  • Every value has to be backed technically and by people. Backup frequency, redundancy, replacement lead times, on-call cover and staff competence have to carry it. The tolerable data loss is directly a requirement on backup frequency.
  • Dependent services inherit the strictest requirement. A shared foundational service has to meet the value of the most critical process running on top of it.

The third point is the one exercises expose most often. A business application with a short recovery time is worthless if the directory service its sign-on depends on takes considerably longer, or if network connectivity to the alternate site is not available until day two. Time targets therefore have to be checked along the whole chain, not at the application boundary.

Which resources and dependencies must BCM capture?

For each time-critical process, record what it actually needs: people with particular skills and authorisations; buildings and workplaces; applications and the infrastructure beneath them; equipment and input materials; information and data holdings; and internal and external service providers.

Providers are where the gap between plan and contract is widest. Emergency plans routinely assume a provider's help without that provider being contractually obliged to anything. A service contract with response times is not a commitment on a recovery time, and a commitment in normal operation is not a commitment during a widespread disruption. So check what is contractually assured, whether the provider has reflected that assurance in its own continuity arrangements, and what happens if the provider itself is affected. Where several critical processes depend on the same supplier, a concentration risk arises that has to be assessed on its own terms.

Which continuity strategies are there?

The most common shortcut is to start writing plans. The useful order is the reverse. First decide, per resource class, which option bridges an outage; only then is a plan worth writing, because a plan is the articulation of a decision already taken.

OptionExampleWhat it is good for
RedundancySecond data centre, dual network connectivity, second supplierVery short recovery times, high running cost
Alternate and remote workingStandby workplaces, location-independent workingLoss of buildings and workplaces
Manual workaroundPaper process, pre-printed forms, telephone listShort bridging during IT outage, limited throughput
Prepared replacement procurementFramework agreements, equipment options, defined delivery routesLoss of equipment and hardware
Third-party performanceContractually assured takeover, cooperation partnerCapacity loss, specialist services
Deliberate acceptanceThe process rests until service is restoredLower-priority processes with a long tolerable outage

The last row is not a capitulation but a legitimate, documented decision. It belongs in the residual risk position and has to be carried by the leadership; otherwise it becomes a surprise at the worst moment.

The options are assessed against four criteria. Does the option meet the time target from the BIA? What does it cost to build and to run? Can it actually be operated with the staff available? And what residual risk remains? The output is a reasoned selection per resource class, not a collection of measures.

What belongs in a business continuity plan?

Plans follow from strategies. The customary split: the business continuity plan describes contingency operation, the restart plan the return to a defined operating state, and the recovery plan the return to normal operation. Alongside them sits the response organisation, with alerting, escalation routes, a crisis team, named roles including deputies, and communication routes internally, to customers, to providers and, where relevant, to regulators.

The quality of a plan depends less on its length than on some unglamorous properties:

  • Directive rather than descriptive. Someone reading during an incident needs steps, decision points and names, not prose about the objectives of continuity management.
  • Current contact details. Out-of-date phone numbers are the most banal and most frequent reason alerting fails.
  • Available outside the affected environment. Emergency documentation held only in the system whose failure it addresses does not exist when it is needed.
  • Built for deputies. Plans only the core team can execute fail during illness waves and holiday periods, and scenarios triggered by loss of staff have to be thought through explicitly.
  • Maintained after change. A site move, new applications, a new critical provider or a rebuilt process render plans quietly unusable.

How do you evidence that a BCM actually works?

Exercising is graduated. Each format has its own purpose, and effort rises with the strength of the evidence it produces.

FormatWhat it testsEffort
Tabletop plan walkthroughCompleteness, clarity and currency of the plansLow
Crisis team or command-post exerciseDecision-making, roles, escalation, communicationMedium
Functional and technical testIndividual procedures such as restore, alerting, alternate workplaceMedium
Restart testActual restart of a service against its committed timeHigh
Full-scale exercise close to a real eventInterplay of technology, organisation and business under time pressureHigh

An exercise without a defined objective is an event. Prepare the objective, a realistic scenario, a script with injects, named observers, defined abort criteria, and the question of which assumption is supposed to be disproved. The value lies in individual assumptions failing, not in a smooth run. An exercise in which everything works was usually designed too gently. The tabletop exercise script shows how roles, injects and expected decisions fit together, on a worked scenario through to the debrief.

The evaluation matters just as much. Every exercise produces a report, named deviations, actions with an owner and a date, and follow-up to closure. If an exercise shows that a committed recovery time cannot be met, there are two acceptable responses: improve the design, or correct the target and carry the changed commitment back into the BIA. Leaving the value untouched is not one of them.

No exercise frequency is normatively prescribed. Common practice is an exercise programme that combines formats and sets the interval by criticality of the process and rate of change. What counts in an examination is not the interval chosen but that it is reasoned, documented and actually kept.

ISO 22301 or BSI-Standard 200-4: which route fits?

Two reference points shape practice in Germany, and one of them is unfamiliar to most readers outside it. Both require the same core elements (analysis, strategy, plans, exercises) and differ in how you enter and in the form the evidence takes.

ISO 22301 is the international management system standard for business continuity, structured in the same clause sequence 4 to 10 as ISO/IEC 27001 and certifiable through an accredited certification body. Note the designation: it is ISO 22301, not ISO/IEC 22301; unlike ISO/IEC 27001, the IEC is not a co-publisher here.

BSI-Standard 200-4 is a methodology published by the BSI, Germany's Federal Office for Information Security. Instead of demanding a complete management system from the outset, it offers a staged route, from reactive through build-up to standard BCMS, and connects to the structural analysis and protection requirement assessment of the BSI's IT-Grundschutz. Its highest stage is designed to reach the level of ambition of ISO 22301; for the audit and attestation offering currently available, the BSI is the authoritative source.

CriterionISO 22301BSI-Standard 200-4
CharacterCertifiable management system standardMethodology with a staged structure
StructureUniform clause sequence 4 to 10, as in ISO/IEC 27001BCM lifecycle with integrated crisis management
Entry pointA complete management systemStaged model: reactive, build-up and standard BCMS
FitIntegrates with other ISO management systemsConnects to IT-Grundschutz structural analysis and protection requirements
EvidenceCertificate via an accredited certification bodySelf-declaration and examinations in the German public-sector and regulatory environment
VocabularyRTO, RPO, MTPDMTA, WAZ, WVZ

The decision follows from who is asking for the evidence. Where customers or regulators expect an internationally recognised demonstration of continuity capability, the route runs through ISO 22301. That expectation is standard in the supply chains of large industrial and financial customers, and behind a contractually committed availability. Organisations working mainly in the German public-sector and regulatory environment, or already working to IT-Grundschutz, more often choose the BSI standard, not least for its staged entry.

Which continuity duties do NIS2 and DORA impose?

Continuity provision has long since stopped being purely a matter of good practice. Several legal instruments require it expressly, without prescribing any particular framework.

NIS2. Directive (EU) 2022/2555 is transposed by each Member State separately; in Germany that is done through the BSIG, the German BSI Act. The catalogue of measures in § 30(2) BSIG names, at number 3, maintenance of operations with backup management, restoration and crisis management; numbers 1 to 10 correspond to points (a) to (j) of Article 21(2) of the Directive. The catalogue is expressly framed as a minimum. For entities in scope this means continuity provision is a subject of supervision, and the management body has to implement the measures and oversee their implementation. Which approach is chosen is not prescribed: a BCMS to ISO 22301 or BSI-Standard 200-4 is one route there, not an obligation.

DORA. Regulation (EU) 2022/2554 addresses financial entities and applies directly across the Union, with no national transposition in between. It treats operational continuity as part of ICT risk management in Chapter II (Articles 5 to 16), and gets specific there in two places. Article 11 requires a comprehensive ICT business continuity policy (it may stand alone or form part of the entity's general business continuity policy) together with arrangements, plans and procedures that secure the continuity of critical or important functions, limit losses, activate containment plans without undue delay, estimate preliminary impacts, and set out crisis communication. Article 12 requires backup policies and procedures with a defined scope and a minimum frequency set by criticality, plus restoration and recovery procedures that are tested regularly. Chapter IV (Articles 24 to 27) then adds the programme for testing digital operational resilience. The difference from general BCM practice lies less in the content than in the pressure to demonstrate it: testing here is not a recommended interval but a regulatory requirement, and the results have to be kept in a form that can be presented to the supervisor.

Neither instrument presupposes a particular framework, and both frame their requirements by reference to risk and proportionality. An organisation already running a BCMS satisfies a substantial part of the substance. What is usually missing is not the provision itself but its presentability: the evidence that the measures were decided, implemented, tested and effective.

Using the evidence more than once

BCM produces evidence that is usable well beyond its own management system. The effort is incurred anyway; whether it pays more than once is decided by how the results are filed.

BCM outputWhere else it counts
Business impact analysisAvailability protection requirement in the ISMS, prioritisation in risk management, criticality statements to customers and regulators
Continuity strategies with residual risk sign-offManagement review, documented engagement by the leadership
Emergency and restart plansInternal audit, certification audit, customer and supplier questionnaires
Exercise report and action trackingEvidence of effectiveness in the ISMS, proof of regular testing
Provider commitments and assessmentsSupplier management, concentration risk, exit strategy
Alerting and communication routesIncident handling, regulatory notification processes

For a piece of evidence to carry more than once it needs a few reliable pieces of metadata: subject, scope, status and validity, owner, and the requirement it is offered against. Without them the familiar pattern appears: a separate filing structure per framework, the same exercise described three times, and at the next examination the question of which version applies.

The second lever is a shared foundation. Processes, systems, providers and owners are the same for BCM, the ISMS and risk management; only the axes of assessment differ. Keep them apart and you maintain the same change several times over and, after the first year, hold three versions of the truth.

Why do continuity programmes fail?

The findings look much the same across industries. These are common enough to be worth checking for deliberately:

  • Time targets born of wishful thinking. The business names short recovery times because short sounds better; none of them is backed technically.
  • Plans for systems rather than processes. The restart of the application is described; how the business process continues without it is not.
  • Providers with no obligation. The plan counts firmly on external help that the contract does not provide for.
  • Emergency documentation inside the affected system. Plans, contact lists and credentials held only where the outage happens.
  • Exercises with no evaluation. Exercising takes place, but produces no named deviations, no actions and no follow-up.
  • No maintenance after change. New sites, applications or providers trigger no update to the BIA, the strategy or the plans.

How to proceed

  1. Settle scope and mandate. Determine which sites, services and legal entities are included, and obtain a written mandate from executive management.
  2. Start with a lean BIA. Cover the processes with recognisable time criticality first, rather than assessing the entire process map in one pass.
  3. Test the time targets against reality. Every value needs technical and staffing backing. Delete or correct uncovered values immediately.
  4. Map the dependency chains, in particular shared foundational services and external providers.
  5. Decide the strategies and have them approved, per resource class, with the cost and the residual risk stated.
  6. Set up the plans and the response organisation: short, directive, and available outside the affected environment.
  7. Define an exercise programme and start with the simplest format. A tabletop walkthrough this quarter is worth more than a full-scale exercise that stays planned for next year.
  8. File the evidence in a structured way and link it to the requirements it will later be produced against.

For financial entities that have to carry the same analyses, measures and evidence towards a supervisor at the same time, our DORA page sets out how that can be held on a single data model.

This article is general orientation and does not replace legal advice. What governs is ISO 22301, BSI-Standard 200-4 and the applicable supervisory law, in particular the BSIG and Regulation (EU) 2022/2554.

How the dependencies can be kept current

The guide makes the business impact analysis the foundation, and that is exactly where the follow-on problem starts: a process's resource list is the part of the BIA that ages fastest, and the next survey is a year away. In Rizzqo that list is not surveyed but read.

A time-critical process is a primary asset, and the applications, systems, sites, providers and people-functions carrying it hang beneath as supporting assets, with named owners. The dependency runs across several steps, so a sub-supplier carrying three formally independent providers becomes visible too. The process's availability requirement is inherited along that same chain, which makes recovery targets justifiable: the operator of an intermediate system can see which process ultimately depends on them.

Rizzqo carries this structure, which information security needs as well, and the follow-up: findings from an exercise carry on as tasks with a recorded risk reduction that only takes effect on completion.

Browse all entriesBack to top

Frequently asked questions

See compliance run on your real assets

Rizzqo turns framework requirements into owned tasks on the assets you already have, and prices the risk in real money.

Made in GermanyHosted in your countryMulti-framework