Skip to content

Run Failure Drills

Run only in isolated non-production with restorable data and secrets.

Establish green readiness, understood backlogs, tested alert routing, recovery owner, stop condition, backup, and baseline canary/audit results.

  1. Select one fault: block Redis/Postgres, isolate Gateway from STS, pause Audit, delay revocation, or activate invalid test policy.
  2. Inject only that fault with a time limit.
  3. Record readiness, traffic, alerts, logs, lag, and detection time.
  4. Remove the fault.
  5. Execute Recover from Failures.
  6. Record readiness, drain, freshness, canary, and audit recovery times.

Alert reaches the owner; services fail closed where expected; recovery follows the runbook; readiness and canary evidence return without unexplained queues.

If a stop condition is exceeded, terminate injection, preserve evidence, restore last known-good, and open an incident.

Update thresholds/runbooks, then validate Back Up and Retain Data.