incident response sop

Service Outage Response SOP

An outage runbook for declaring, mitigating, and communicating production service disruptions.

service outage sopoutage response runbookincident response procedureproduction incident

Example text

Service Outage Response SOP

SOP-OPS-010: Service Outage Response

Owner: Engineering On-Call
Version: 1.0

Trigger

Monitoring alert, customer reports, or internal detection that a customer-facing service is degraded or unavailable.

Steps

  1. Acknowledge the alert and open an incident channel.
  2. Assign an incident commander and a communications lead.
  3. Confirm impact: which services, regions, and customer segments.
  4. Mitigate with the fastest safe action (rollback, failover, scale, feature flag).
  5. Post a status update within 15 minutes of declaration.
  6. Restore service and verify with synthetic checks plus customer-path testing.
  7. Resolve the incident only after error budgets and core journeys are healthy.
  8. Schedule a blameless retrospective within three business days.

Customer communication template

We are investigating an issue affecting [service]. Current impact: [summary]. Next update: [time].