Status Page Update
Initial status update within 5 minutes of confirmation. Updates per cadence based on severity. Resolution post. Postmortem published per timeline.
Before you start
- Status page tool with admin access (Statuspage, BetterUptime, etc.)
- Incident severity definitions documented
- Communication template library
- Detected or reported incident details
- Affected components or regions
- Engineering team's diagnosis or update
The steps
- Confirm the incident is real and impactful — Before posting publicly, verify: monitoring alerts triggered, engineering confirms impact, multiple customer reports. Don't post for false alarms — false posts erode trust in the status page.
- Classify severity — Map to documented levels: P1 (full outage, all users affected), P2 (degraded service, partial users), P3 (specific feature down, low impact). Severity determines update cadence and communication tone.
- Post the initial 'investigating' update — Within 5 minutes of confirmation: post a brief, neutral update — what's affected, that we're investigating, when next update will be (typically 15-30 min). Avoid speculation. Don't promise a fix time you don't know.
- Maintain update cadence per severity — P1: every 15-20 minutes. P2: every 30-45 minutes. P3: every hour. Even 'no update' is an update — 'still investigating, no new info' beats silence. Customers fill silence with assumptions.
- Post resolution and impact summary — When resolved: confirm with engineering, post 'Resolved' update with what happened (high-level), how it was fixed, and that you'll publish a postmortem within 5 business days for P1/P2.
- Trigger postmortem workflow — For P1 and P2 incidents: trigger the postmortem template. Engineering owns the technical detail; support owns the customer-impact section. Postmortem published within 5 business days — public for P1, internal-only for many P2 cases.
If it goes wrong
Status page lags behind reality (customers know before page shows)
Tighten the alert-to-post pipeline. 5-minute SLA for initial post. Customers tweeting before the status page shows is a credibility hit.
Updates are vague and unhelpful
Specificity matters: 'investigating slow loading on dashboards in EU region' beats 'investigating an issue'. Use customer-facing language, not internal jargon.
Postmortems never get published
Hard SLA: 5 business days post-incident. Track postmortem publish rate as a leadership metric. Customers losing trust is the cost.
All OpenLabor playbooks