Illustrative operations runbook; commands are intentionally abstract.
Use when the approved service-health alert fires for API error rate above the operational threshold. Do not use for confirmed security incidents; follow the security incident process.
Prerequisites
On-call operator access, incident channel, health dashboard, dependency dashboard.
Steps
1. Confirm the alert is current and identify affected region/service. Expected result: scope is known.
2. Check dependency health. If a shared dependency is degraded, open or join the dependency incident and do not restart local components repeatedly.
3. If dependencies are healthy, compare current release and capacity indicators with baseline.
4. Apply only the approved reversible recovery action for the observed condition.
5. Validate error rate, latency, and one synthetic transaction.
Stop/escalate
Escalate after one unsuccessful recovery attempt or if data-integrity indicators are abnormal.
Closure
Record action, result, linked incident, and follow-up item.
Why it works: The runbook defines when not to use it, one recovery attempt, validation, and escalation instead of encouraging uncontrolled retries.