Let’s talk

Insights

The Systems Are Back. Is the Business Ready?

How to reconcile orders, payments and work completed during an outage before restarting automation.

Share this insight

A perspective worth sharing.

Your LinkedIn caption and article image are ready. Edit the caption, copy it, then paste it into your LinkedIn post.

Share on LinkedIn copies your caption and opens the article preview in a new tab. Paste the caption before publishing.

Share on LinkedIn

Three illuminated transaction panels align with a vehicle dispatch lane beyond a Gulf operations room.
AI-generated conceptual illustration of business records reconnecting with physical operations. The setting and interface do not depict a real client, incident or deployed system. Bridges

How to reconcile orders, payments and work completed during an outage before restarting automation.

Restoring an application does not establish that its records match what the business actually did while it was unavailable. Before normal processing resumes, reconcile the restored system with manual work, external transactions and queued instructions. Assign each unresolved item an owner, and release automation only when its starting state is understood.

For a COO, CFO or technology leader, the practical question is simple: what could happen twice, disappear, or resume from the wrong point when we turn the workflow back on?

This guide proposes a transaction-level recovery exercise for GCC businesses. It extends our business continuity overview with a method for resolving the work that accumulated around an outage. The examples and worksheet are illustrative recommendations, not client results or a prescribed regulatory procedure.

The missing interval in a recovery plan

Imagine a vehicle-transfer business whose dispatch system becomes unavailable. Staff continue working through an approved fallback process. A transporter leaves with cars that still appear unassigned in the restored application. A customer payment succeeds at the payment provider, but its confirmation never reaches the booking record. A cancellation is agreed by phone while a scheduled reminder remains queued.

Restoring a backup can make the application accessible without resolving any of those differences. Releasing every pending instruction could book a second transfer, request payment again or send a customer a message that contradicts an agreement already made.

The recovery window therefore starts with the earliest potentially affected transaction, not simply the time someone declared the incident. It ends when the relevant backlog and live operations have been reconciled sufficiently for a controlled restart. Different systems may have different recovery points; identify those explicitly rather than treating one timestamp as universal.

NIST’s Guide for Cybersecurity Event Recovery, published in December 2016, supports recovery planning, realistic exercises and learning from incidents. The specific worksheet below is our proposed way to make the business transaction visible within that work; it is not a NIST-prescribed template.

Begin after the environment is cleared for recovery

Transaction reconciliation belongs within the incident response and recovery plan. It does not replace containment, investigation, validation of recovery assets or the decision that an environment is safe to use. Restored data also needs to be trustworthy: NIST’s data integrity recovery practice guide, finalised in September 2020, addresses confidence in recovered data and demonstrates recovery approaches in a laboratory environment.

Once the responsible technical and security leads permit the next stage, establish a controlled boundary around the affected business flow. Identify which integrations, scheduled jobs, queued messages and AI-agent actions could change records or trigger external work. Decide which must remain paused, which can run in a restricted mode and which are unaffected. Apply those decisions through the incident lead; a blanket shutdown of unrelated services may create another operational problem.

Keep a record of what is held, where it is held and who can release it. Include a recovery checkpoint before any reconciliation writes. Use approved, access-controlled evidence storage and retain the original records needed for investigation. An analyst’s untracked spreadsheet should not become the only record of the recovery.

Build a reconciliation record around business events

A backup shows one system’s state. A useful reconciliation record brings that state together with evidence of what happened elsewhere. Use a stable business reference wherever one exists: a booking, order, invoice, vehicle movement or payment reference. When references differ across systems, record the relationship instead of assuming that similar names or amounts prove a match.

For each affected item, capture the restored status, evidence of external or manual activity, the proposed next action and its owner. Record event time and time zone as well as the time the evidence was received. That distinction helps when confirmations arrive late or teams operate across UAE and Saudi time zones.

The following is an illustrative starting worksheet. Its rows represent different situations, not mandatory fields for every business.

What the restored system showsEvidence to establishControlled next action
Transfer awaiting assignmentDispatch reference, vehicle identity and confirmed movement recordIf the movement is verified, update its state through an approved path and suppress the duplicate assignment
Payment pendingProvider reference and authoritative payment status, matched to the orderReconcile the existing payment; hold another collection attempt until the outcome is known
Booking confirmedAuthorised cancellation and any agreed financial treatmentResolve the cancellation and dependent instructions before releasing reminders or fulfilment
No booking recordApproved fallback record and checks against other channelsCreate or recover the record once, preserving its original business reference and history
Several records appear to matchOriginal references and accountable owner reviewKeep the items unresolved until identity and intent are established

Do not collect more personal or payment data than the exercise needs. Prefer references to controlled source records over copied documents, card information or screenshots scattered across chat threads.

Decide what to replay, reconcile or hold

Three practical categories help the recovery team make its next decision.

Reconcile an action that already happened. A vehicle has departed or a payment has completed. Bring the application record into agreement with reliable evidence, using an approved correction process. Completing the missing record should not perform the external action again.

Resume work that has not happened. The original request is still valid, its dependencies are satisfied and there is evidence that the intended action has not completed. Resume from a known point, with the appropriate duplicate protection and monitoring.

Hold a disputed or incomplete case. Conflicting evidence, uncertain payment status or an unverified cancellation needs an owner and a decision time. A timeout is evidence that a response was not received; it does not prove that the action failed.

For technical teams, that last distinction is central to retry design. The Amazon Builders’ Library explains how idempotent APIs use request identity to help repeated attempts avoid additional side effects. Ask the team to establish the actual guarantee for each relevant integration, including its scope and retention period. A key accepted by one service is not a guarantee for the entire business workflow.

If an unwanted action has already completed, a compensating business action may be necessary. Microsoft’s Compensating Transaction pattern explains why distributed work cannot always be undone by restoring an earlier state. A correction may need its own rules and may itself fail. Treat refunds, cancellations and physical movements according to their real business consequences, with the appropriate authority.

Rehearse one vehicle-transfer journey

Illustrative exercise; no client or production incident is represented. Prepare a test booking in an isolated environment with simulated external services. The starting record says that a transfer is confirmed, payment is pending and a vehicle has not been assigned.

Give three participants separate evidence: finance receives a successful payment confirmation; operations records that the vehicle left under the approved fallback procedure; customer service receives a request to change the delivery location. Then restore an earlier application state and introduce a delayed payment event and a duplicate dispatch instruction.

Ask the team to establish what is true before any write or replay. The expected response is to link the completed payment and movement to the booking, prevent a second collection or transfer, and route the destination change to the person authorised to decide it. The exercise should preserve the evidence behind each decision.

The destination change makes this more useful than a simple restore test. A technically valid update may no longer be operationally possible. The car may already be on a transporter, the destination may affect the agreed price, or an external party may need to confirm the change. These are test conditions to explore, not assumptions that software should silently resolve.

Run the journey again with one missing confirmation. The team should demonstrate how it holds that case, communicates the uncertainty and continues only the work that can be separated safely. A successful exercise does not require pretending every exception has disappeared.

Restart in bounded groups, with business sign-off

A restart decision should identify the scope being released: a service, queue, integration, site or class of transactions. Releasing one group does not establish that every affected process is ready.

Before release, ask the technical owner to confirm the relevant controls and monitoring. Ask the business owner to confirm that the reconciled state supports the intended next action. Finance, operations and customer service may each own different consequences. Name the person authorised to decide when their evidence conflicts.

Use a small, controlled first group to observe the effect of the restart. Confirm that intended actions complete, duplicates are contained and unresolved items remain held. Define a stop condition and a route back to controlled processing before increasing volume. The suitable group size depends on the impact of a mistake; there is no universal safe number.

After a cyber incident, these steps remain subject to the incident command and security recovery decisions. NIST’s April 2025 incident response recommendations place response and recovery within wider risk management. They do not establish a GCC-wide legal requirement. Local obligations and contracts need their own assessment.

Where AI agents can help—and where their authority ends

An AI assistant could help organise evidence, propose reference matches or summarise the outstanding cases for an operations lead. Those are possible uses to evaluate against a known test set; they are not a claim of proven reconciliation accuracy.

Keep proposed matches separate from approved corrections. The agent should expose the source references behind a recommendation, flag ambiguous cases and operate only within permissions agreed for recovery. A fluent summary is not enough to authorise a new payment or dispatch.

A restored agent workflow may itself contain stale instructions. Recheck its queued actions, current permissions and business context before allowing it to resume. Someone may have completed the task manually while the agent was offline. Our guide to defining an AI agent’s authority provides the complementary governance starting point.

Measure the work remaining, not only the systems restored

For the rehearsal, agree definitions that the business and technical teams can both use. Track the number of affected items in scope; those matched to evidence; those corrected and verified; and those still disputed, with their age and owner. Keep customer or financial exposure separate from a simple item count.

Record time to restore the technical service and time to authorise the relevant business flow as separate milestones. Neither is a substitute for the organisation’s recovery objectives. The difference can reveal work that a technical availability report leaves out.

Count duplicate actions detected during replay testing and distinguish prevention from correction after the event. Record unplanned manual decisions and missing evidence as improvement work, then repeat the affected test after fixing them. These are proposed measures for the exercise, not claimed Bridges results or industry benchmarks.

Questions leaders should settle before the next rehearsal

Is a successful backup restore enough to resume normal operations?

It establishes part of technical recovery. The business still needs to verify affected transactions, work performed through fallback channels and external actions that may have completed independently.

Who should own reconciliation?

A named business process owner should own the outcome, supported by the people who control the systems and evidence. Incident leadership coordinates the release. The IT team should not have to infer the commercial meaning of a disputed transaction alone.

Must every exception be closed before anything restarts?

No. A business may release a verified, separable group while holding exceptions. That decision needs explicit scope, accountable acceptance of the remaining exposure and controls preventing the held work from being processed accidentally.

What is the smallest useful first step?

Choose one important customer journey. Identify the work that could continue outside its main application, build the reconciliation record and rehearse one completed action, one genuinely pending action and one uncertain case. Document the release decision and the evidence it required.

Make business recovery part of the design

Start the conversation with the service people need to deliver and the evidence that confirms it can operate. Then connect the recovery plan, integration design and operational responsibilities around that service.

Bridges’ Disaster Recovery, Business Continuity & Resilience, Managed IT & Cloud Operations and Data Engineering & Analytics services are relevant to this work. A useful starting brief is one business flow, its recovery dependencies and the transactions the organisation cannot afford to lose or repeat.

Talk to Bridges about operational recovery readiness.

Share this insight

Related perspectives & expertise