Originally drafted 2 September 2026. First published 7 October 2026; revised for publication.

Consider an illustrative continuity failure. A utility’s customer-service platform becomes unavailable on a Monday morning. The immediate concern is technical: restore the application, isolate the affected accounts and assess what happened. The deeper problem appears at the service desk. Staff cannot confirm whether a customer has already paid. The call centre has no agreed script for disputed balances. A field team receives different instructions on WhatsApp. The backup exists, but nobody knows which version can be used for a live decision or who has authority to approve a temporary workaround.

A cyber incident does not create these weaknesses. It exposes a service that was designed to work only when every system, identity check and communication channel is available. That is why cyber resilience is an executive concern. NIST’s Cybersecurity Framework 2.0 release announcement, dated 26 February 2024, describes the addition of a Govern function to its familiar Identify, Protect, Detect, Respond and Recover model. A resilient organisation is one that can continue or safely pause its essential services, explain the status of its records and restore accountable work under disruption.

This framing matters across African digital services, where outsourced platforms, mobile channels and agent networks create dependencies exposed when an identity system, provider or connection fails. The objective is proportionate ability to protect the service and recover it without improvising trust.

The service, not the server, is what must recover

Security programmes often begin with an asset inventory: devices, applications, networks and accounts. That work is necessary, but leaders should start one level higher. Which services cannot fail without immediate harm? Which decision must still be made when the usual application is unavailable? Who needs a reliable record, and how will they know whether it is current?

A hospital may need to identify a patient safely while a registration system is down. A lender may need to stop a suspicious transfer without losing its record of a legitimate one. A public agency may need to keep a benefits queue fair when staff move temporarily to manual handling. A logistics business may need to dispatch goods while preserving evidence for later reconciliation. Each scenario requires more than a backup server. It requires a designed relationship between people, authority, data and a degraded workflow.

Recovery has a service level. Restoring an application is not the same as restoring the ability to make a payment, serve a customer or safeguard a case. A meaningful recovery target therefore needs to name the minimum function that must return, the records required to perform it, the maximum acceptable uncertainty and the person authorised to make the trade-off. Without those choices, technical recovery may be fast while the service remains effectively unusable.

This approach makes security investment easier to prioritise. An organisation may tolerate the temporary loss of an internal analytics tool. It cannot tolerate losing the ability to verify a high-value transaction, contact affected customers or distinguish a valid instruction from a fraudulent one. Investment should follow the consequence of disruption, not the visibility of the technology.

Identity is the control plane during disruption

When systems are under pressure, access decisions become more consequential. An emergency fix can tempt teams to share passwords, grant broad administrator rights or bypass approval checks. These actions may accelerate a recovery in the moment while creating a second incident through unaudited access or fraud.

The more durable approach is to predefine emergency roles and tightly bounded permissions. A service manager may be able to place an affected workflow into a safe hold. A finance officer may approve a verified exception within a limit. A security lead may revoke sessions or isolate a supplier connection. None should need unrestricted access to every system simply because an incident has occurred.

Access is not authority. A person who can view a case may not be authorised to amend a payment instruction. A vendor who can support a platform may not be entitled to export customer data. A temporary user account may be necessary during recovery, but it should have an owner, a short lifetime and a record of the actions taken. These distinctions are basic, yet they are often missing from service-continuity plans written as if people will act calmly and consistently under pressure.

They matter in operating environments with shared devices, distributed branches, agency staff and intermittent connectivity. Strong authentication should be usable at the point of work, with secure recovery routes when a phone is lost, a local connection fails or a staff member must move to an assisted channel. A control that stops a frontline team serving anyone during a manageable outage is not automatically a strong control; it may be a design that transfers risk to an informal workaround.

Backups are evidence only when restoration is tested

Many organisations can say that they have backups. Fewer can demonstrate that the backup contains the required records, can be restored within the service’s time limit, and will not reintroduce an undetected compromise. A backup policy is therefore not a resilience outcome.

The test should be tied to an actual service journey. Can the organisation recover a selected set of customer or case records? Can it verify the restoration’s integrity? Can staff use the recovered data without duplicating decisions made during the outage? Can it reconcile activity performed through the fallback process? These questions reveal whether backup, data governance and operational practice fit together.

A record needs a status, not just a location. During disruption, staff should be able to identify whether a value is current, restored from a known point, manually captured pending reconciliation, or unavailable. That is more useful than giving everyone a copied spreadsheet that looks authoritative but has no controlled update path. Sensitive operational and personal data should remain in governed systems with access controls and audit logs; emergency extracts should be minimal, time-bound and protected.

The same discipline applies to suppliers. Cloud providers, managed-service firms and software vendors can be part of a sound service design. But each critical dependency needs a realistic answer to: what happens if the provider’s service is unavailable, its account is compromised, or support cannot be reached? This is not an argument for building everything in-house. It is an argument for knowing the route to safe operation when a dependency is unavailable.

Exercise the handover, not just the response plan

A plan written by a security team may describe technical containment well while leaving operational teams unsure what to do next. The gap is usually exposed when people practise together: operations, customer support, finance, legal, communications, technology and executive decision-makers.

The ITU–INTERPOL Regional Cyberdrill for Africa announcement set out an exercise for 1–4 July 2025 in the Republic of Congo, designed to test and improve collective response, including prevention, detection, investigation, response and recovery. That statement describes the organisers’ purpose, not independently measured results. Its focus on practice reflects an important lesson: resilience is a learned organisational behaviour. A slide deck cannot reveal an approval that takes three days, an unreachable supplier contact, an unclear customer message or an identity process that fails on a shared device.

Run one credible scenario. Select a disruption that would matter to the service: ransomware affecting a case-management platform, compromise of a supplier account, loss of a regional connection, or suspicious instructions arriving through a trusted channel. Walk the scenario through the first hours and the first working day. Record decisions, queues, data handoffs, customer communications and escalation points. Then improve the workflow and repeat the exercise.

A small institution may start with a tabletop exercise and a recovery test for one application; a larger enterprise can add role-based drills and supplier participation. The test is whether people can act with sufficient information and authority.

Governance must make trade-offs visible

NIST describes CSF 2.0 as a taxonomy of high-level cybersecurity outcomes that does not prescribe how to achieve them. It is guidance rather than a certification of this service-design approach. Its new Govern function sharpens a board-level question: who decides how much risk the organisation accepts in a critical service, and how will that decision be reviewed? Governance means assigning an owner, agreeing tolerable downtime and data loss, setting escalation authority, evaluating supplier risk, funding tests and seeing evidence that recovery works.

African organisations often face scarce specialist staff, limited bandwidth in field teams, and vendor-led projects with unclear ownership after implementation. These conditions make proportionate design more valuable: clear ownership, tested offline or assisted fallback paths, documented interfaces, secure identity recovery and local operating capability can provide more resilience than a product nobody can administer.

Cybersecurity is not a reason to centralise every record or deny every integration. Each connection needs a defined purpose, least-privilege access, logs, an owner and a revocation path. A narrow trust layer can support shared provenance or settlement; it should not become a repository for personal operational data or a substitute for verification.

Build resilience where the next decision is made

The most useful starting point is not a generic awareness campaign or a multi-year transformation programme. It is the service whose failure would cause the most immediate operational or public harm.

Discover the critical journey. Follow it from request to outcome, including manual handovers, vendor dependencies and the channels used when a person cannot complete the journey digitally. Identify the decisions that must continue and the evidence each decision needs.

Design the degraded mode. Define safe holds, limited manual procedures, record statuses, approval limits, customer communication and reconciliation responsibilities. Make these routes usable with the devices, connectivity and staffing that will actually be available.

Build and test the recovery path. Verify identities, restore a representative data set, run the fallback workflow and reconcile the results. Measure the time to restore the service, unresolved exceptions and the effort required from frontline teams.

Xelius supports organisations through service and workflow research, secure data architecture, platform design and implementation. A practical engagement begins by identifying the service that cannot afford an improvised recovery, then designing and testing the smallest dependable path through disruption.

Security becomes credible when recovery is ordinary work

The maturity of a cyber programme is not shown only by the controls it can list or the attacks it can describe. It is shown by whether a service can pause safely, preserve accountability and resume without losing the confidence of the people who depend on it.

That requires technical safeguards. It also requires decisions made before the incident: who may act, which records are trustworthy, what can continue manually, and what must wait. When those answers are built into the service, cybersecurity stops being an emergency layer around digital transformation and becomes part of how the organisation delivers reliably.