Example text
Service Outage Response SOP
SOP-OPS-010: Service Outage Response
Owner: Engineering On-Call
Version: 1.0
Trigger
Monitoring alert, customer reports, or internal detection that a customer-facing service is degraded or unavailable.
Steps
- Acknowledge the alert and open an incident channel.
- Assign an incident commander and a communications lead.
- Confirm impact: which services, regions, and customer segments.
- Mitigate with the fastest safe action (rollback, failover, scale, feature flag).
- Post a status update within 15 minutes of declaration.
- Restore service and verify with synthetic checks plus customer-path testing.
- Resolve the incident only after error budgets and core journeys are healthy.
- Schedule a blameless retrospective within three business days.
Customer communication template
We are investigating an issue affecting [service]. Current impact: [summary]. Next update: [time].