Make recovery a verified transition
By the end, explain the diagram in your own words, solve the case and justify the correction.
Prerequisites : Operational controls for digital assets
Level 3 · Advanced →Reading path · 13 / 17 · Advanced
An alert becomes an incident through investigation, not through its severity label alone.
The essentials
An alert becomes an incident through investigation, not through its severity label alone. Establish what is known, what is uncertain and which services may be affected. Preserve relevant logs and instruction identifiers before changes erase evidence. Avoid exposing keys or unnecessary client information in the investigation record.
How it works
Containment is a decision with operational consequences. Pausing new instructions can limit exposure, while already signed or broadcast transactions may remain executable. Define authority to pause, communications ownership and criteria for restoring service. A dashboard switching to green is not sufficient recovery evidence.
What to watch
Recovery should reconcile outstanding instructions, verify access changes and test the restored service under supervision. Record residual risks and the person accepting them. A review should identify root causes and measurable improvements rather than merely confirming that the service restarted.
Build your analysis
Maintain a timeline that distinguishes observed events from assumptions. Capture the last known good state, affected instructions, access changes and relevant transaction identifiers. Containment should reduce further harm while preserving the information needed to investigate. Restarting a service can change evidence and must not be treated as a universal first response.
Extend the workshop
For the workshop, classify each of the ten instructions as delivered, pending, failed or unknown, and name the evidence for that state. Define a limited restart scope, a person authorised to approve it and checks after resumption. Finally, turn the underlying cause into a corrective action with an owner and a verification step.
Understand the details
Run a tabletop exercise with timestamps: detection, scope assessment, containment, reconciliation and controlled resumption. Inject uncertainty, such as an unreachable signing service and a late transaction receipt. Teams should explain their decisions and evidence without treating every timeout as a failed transfer.
Boundaries and common mistakes
Incident reporting requirements depend on the service and applicable framework. This lesson provides an operational exercise, not a universal legal notification deadline.
The mechanism at a glance
- Detect and preserve
- Contain exposure
- Reconcile uncertainty
- Resume with evidence
Apply the lesson to a case
A signing API times out after accepting ten requests. Monitoring later finds seven transaction hashes. Before resuming, classify the seven known and three uncertain instructions. Explain why resubmitting all ten is unsafe and which evidence is needed for the remaining three.
Reconcile original request identifiers, signer evidence and chain observations. Retry only under a defined duplicate-prevention procedure. An unknown result is a state to investigate, not proof of non-execution.
Prepare a correction note
Describe the passage and the proposed correction. This creates a local note for you to share; it sends nothing. Do not include personal or confidential information.