Say Only What the Evidence Supports in the Postmortem
Weeks after an outage, the live updates have been superseded and the incident channel has scrolled away. The postmortem is the account that remains. It gets filed, forwarded, and read by people who were never on the bridge, and it carries an accountable owner and named reviewers.
During the incident, the question was what the team could say now. Afterward it becomes what conclusion the review has earned. The method is to match each sentence to the evidence that sentence requires, then state the strongest conclusion that evidence supports, plainly. A postmortem can name a cause, declare the incident over, and describe a control that addresses the failure mode, and each of those is the right sentence once its evidence is on the record.
A typical first draft looks like this:
The outage was caused by a connection pool exhausted by a deployment that doubled worker count without raising the pool ceiling. The issue is now fixed and there was no customer impact. No data loss occurred. This was human error during the change window, and the new pre-deploy check will prevent recurrence.
Four sentences, six claims, and nothing in the paragraph lets a reader tell what was established from what was assumed. The sections below work through what each of those claims requires, and the paragraph gets rewritten near the end.
Part one of this series, Say Only What the Evidence Supports: Incident Language for the People Who Decide, covers the incident while it runs, from labeling an impact statement to a long set of rewrites for live updates. This piece picks up with the document written after the incident ends.
Like part one, this is written from the IT side, and the tables are working patterns to adapt to your own process. It offers no legal advice, and questions with legal, contractual, or HR weight belong to the people in your organization who own them.
The postmortem is a different document
In a live incident, the next update supersedes the last one. That property does more work than people notice. A live update that overstates something can be corrected in the same active channel, while the incident is still being worked.
A postmortem can be corrected too, but the correction has to be recorded where later readers will find it and sent to the people known to rely on earlier copies, because those copies keep circulating. The postmortem is written after the pressure is off, then revised, forwarded, and relied on outside the incident flow. It gets attached to a ticket, quoted in a quarterly review, cited by someone in a different org six months later who missed the incident and may never see the amendment. Each claim deserves more scrutiny when it goes into that document.
It often gets less. During the incident, nobody knows the answer and everybody knows nobody knows. Afterward there is a document with headings, and a required Cause field pushes reviewers to fill it before the evidence is sufficient.
Choosing how much review an incident gets
Every declared incident should leave a record proportionate to its significance. The three weights below are one workable model, and your incident policy sets the real thresholds, including when an existing ticket is enough. Deciding the size prevents two failures: a full review of a five-minute blip that nobody reads, and a two-line note for an event that deserved a real look.
| Weight | When it fits | What it contains |
|---|---|---|
| Full review | A major capability was unavailable, the incident repeated, data or security involvement was confirmed, a specialist assessment of suspected involvement calls for full review, an automated action caused material impact, or a recovery action failed and extended the outage | A labeled timeline, each claim with its evidence, findings, prevention status, and action items, plus a business annex when business impact was material |
| Structured review | Bounded impact, a missed published service commitment, a significant near miss, a false recovery, or a monitoring gap that let the incident run longer | Timeline, what helped and what hindered, findings, owned action items |
| Learning note | A contained fault, an alert that fired wrongly, or a near miss with no impact, and no data or security involvement | What was expected, what happened, what changed, and what gets tested next |
The choice belongs in the record, with a line on why, so a later reader sees that the depth was decided. A near miss can deserve more review than its impact suggests, because it shows a failure path while that path is still cheap to close.
Someone owns every sentence
Some teams now write the first draft with machine help. An assistant summarizes the incident channel, a model reads the logs and proposes an explanation, an analysis tool assembles the timeline. The appeal is less drafting for a tired team, and the output arrives fluent, confident, and formatted like a finished document.
Fluent output hides where each sentence came from. Unless the workflow labels and checks them, a generated summary can write a hypothesis in the same voice as a measurement, fill the Cause heading because the heading is there, and report an absence of errors in logs it was never given. The sentences read well, and a reader six months later may have no way to tell which ones anybody checked.
None of that changes who answers for the document. Give the postmortem an accountable owner and named reviewers. When a director, an auditor, or a customer's account team asks what a sentence rests on, those people are the ones expected to point at something. What that accountability means in formal or legal terms is a question for the people in your organization whose job that is, and this piece takes no position on it. The engineering side is simpler. Every material claim, figure, and conclusion should trace to someone who checked it, and so should every material timeline entry derived from more than one source. Raw entries cite their system of record, and the document owner is accountable for the rest.
Four habits keep machine-drafted text in check. Treat generated analysis as inference until a person has checked it, and label it that way. Name the input, because an analysis that found nothing is describing the data it was handed. When machine analysis shapes a finding, keep that input and the generated draft linked to the record, as far as your retention and handling policy allows. Name the reviewer, so the reader knows who checked the figures, and name the owner accountable for each one where that is someone else. And treat a material gap in the data as a finding or a stated limitation, because automated summaries can smooth over gaps in telemetry unless the workflow makes missing input explicit. When checking finds the draft wrong on figures or sequence in more than one place, rewrite from the primary sources instead of patching the draft.
When a tool drafted or analyzed it
| Instead of | Try |
|---|---|
| The AI identified the cause. | An automated analysis flagged the cache change as consistent with the timing. The on-call lead checked that timing against the deployment log and cache metrics and agrees it should stay on the list. The network change in the same window is still open. We are testing it now. |
| The summary below was generated from the incident channel. | The summary below was drafted with an assistant from the incident channel. The incident lead checked every figure against its source system and owns it. |
| The analysis shows there was no impact. | The analysis found no errors in the logs it was given. Those logs cover the web tier only, and the payments tier was outside its input. |
| The model says the failover worked. | The secondary site took traffic at 14:31 and has served checkout at baseline error rates since. Replication lag back to the old primary was 4 seconds and falling at 14:40. An assistant summarized the metrics, and the storage lead confirmed the numbers. |
| Nothing unusual in the telemetry. | Telemetry for the window has a twelve-minute gap between 14:10 and 14:22. The analysis covers the rest, and the gap is recorded as a limitation of the analysis. |
What each claim requires
The useful question to ask of a postmortem sentence is what you would point at if someone challenged it. That question is mechanical, and the answer can be written down in advance.
| Claim | What you would point at |
|---|---|
| A named cause | The scope you examined, the proximate mechanism, the contributing conditions, the plausible alternatives you evaluated and the status of each, and the uncertainties that remain |
| Resolved | Technical verification, business service verification or an explicit not-applicable, the stability window observed, and the residual risk recorded with the authorized owner's disposition of it |
| No data loss | The population checked and what remains unchecked, integrity validation status, a reconciliation reference or an explicit not-applicable, and the evidence you reviewed |
| No customer impact | The customer populations and paths assessed, impact assessment status, who owns that assessment, and the evidence behind it |
| A dependency or provider was at fault | Evidence that distinguishes internal, shared-boundary, and provider-side explanations, the provider's own finding as corroboration where available, and review through your organization's established supplier and external-communication process, because this claim leaves your organization |
| A control addresses the failure mode | The implemented control, the validation evidence, the bounded failure mode it covers and the conditions it was tested under, and the residual risk left outside that boundary |
| An operator's action caused it | The action as a timeline fact, the mechanism linking it to the failure, the evidence that separates it from the alternatives, the conditions that shaped it, and whether the expected controls engaged. For an ordinary operational mistake, a finding that names only the person is insufficient for an engineering review, as the section on findings that stop at a person explains |
Read that table as a cost sheet of evidence requirements. Every row is a claim you can make, and the right column is the evidence it requires. If you have the items, write the sentence and cite them. If you lack them, you have three options: get them, write the weaker claim you can support, or write the unknown.
Review disposition sits next to the evidence and apart from it: who reviewed each finding, whether they accepted it, and any dissent. Acceptance records the review team's disposition, and the evidence is what makes the finding supportable. Where publishing the finding as the organization's position needs someone else's approval, record that approval separately.
Four substitutions are worth watching for. Zero failed requests stands in for zero customer impact, no identified loss for no integrity impact, no complaints for no affected users, and a successful backup job for a successful restore. Each pair looks the same in a summary and needs different evidence.
Fill the right column first. A postmortem template with those fields above the prose changes what people write, because a blank next to stability_window announces itself while a confident sentence hides the same gap.
Three claims a postmortem wants to make
A cause. This is the heading reviewers feel most pressure to fill. Early in the review, what you often have is a mechanism consistent with the evidence and one or two alternatives you left standing because the service came back and the pressure ended. Write the mechanism, name the alternatives, and say what would discriminate between them. A postmortem that says "the observed behavior is consistent with pool exhaustion; we have not ruled out the upstream timeout change, and the discriminator is whether the failure reproduces with each change applied alone" is more useful six months later than one that named a cause the evidence could not yet support, because the next person reading it can pick up the open question. A named cause can close a specific causal question while contributing conditions stay open, and if the naming was wrong, it closes that question in the wrong place.
A resolution. When the service starts serving again, people often treat that moment as the end of the incident. But resolution is a claim about a window, not a moment. Serving correctly for four minutes and serving correctly for six hours are different claims. State the stability window actually used to close the incident, measured against the exit criteria agreed for that service. If it was short because the incident closed at 2am, say that, because the person reading this in a quarterly review should be able to tell that resolved meant twenty minutes of green at the end of a long night.
A prevention. Prevention claims are especially prone to outrunning their evidence. "This will not happen again" is a prediction about every future state of a system you have just demonstrated you understand incompletely. The supportable version names a bounded failure mode, the control's current evidence, and what stays outside its coverage: "the pre-deploy check now fails any deployment where worker count times connections per worker exceeds the pool ceiling, and we tested it against the release 4.2 change." That says less. If a later incident arrives through a path the check leaves uncovered, the document still reads as a bounded claim rather than a broken promise.
When the evidence is in, say it plainly
The claim table can read like advice to hedge everything. Its point is that strong claims have an evidence requirement, and once the review has met it, the strong sentence is the right one. Leaving a reviewed finding wrapped in "consistent with" misreports the investigation in the other direction, because a reader who sees that phrase reasonably assumes the team never finished.
A finding at full strength names the mechanism, the scope it applies to, the evidence that establishes it, and who reviewed it. It can also name the conditions that let it happen and what remains open, which keeps a strong claim from sounding broader than it is. After the review, the checkout incident from the opening draft supports a much stronger version:
The review concluded that pool exhaustion caused the checkout errors that began at 14:02 in the primary region, and that the worker-count increase in release 4.2 was sufficient to produce it under replayed production request rates and concurrency. In production, pool acquisition timeouts in the application log match the failed checkout requests by request ID from 14:02 until the pool recovered at 14:18:40. After the rollback at 14:15 and the pool's recovery, checkout latency stayed elevated until 15:25 as the queue of pending order jobs drained, per the queue depth metric. Staging replayed recorded production request rates and concurrency with production's pool settings and worker counts; production's request mix and dependency latency were outside the replay. Under those conditions, staging reproduced the exhaustion with the worker-count change alone and with both changes together, while the timeout change alone left the pool within its ceiling. The timeout change's contribution in production remains unmeasured. Contributing conditions were the absence of a capacity check in the deployment pipeline and a pool ceiling sized when the service ran half as many workers. This finding covers this incident only; the later latency spikes are under separate review.
Review disposition: accepted by the review team, including the checkout service owner. No dissent recorded.
Each strong word becomes available at a specific point. Resolved is correct once the exit criteria are met and recorded. A causal conclusion is supportable once the evidence establishes the claimed relationship, addresses the plausible alternatives that could change the conclusion, and states its scope. Recording it as the review's finding is a separate step, and the disposition goes beside the evidence. When reviewers disagree, the record carries each interpretation, the evidence each one explains, and what would settle it. A prevention claim is correct at the state its evidence reaches, for the failure mode it was tested against. Once those points are met, drop the uncertainty language that no longer describes the investigation.
This piece sets out what a defensible finding has to contain. How a review weighs competing alternatives against each other is a question of causal method, and it belongs to a separate piece. What the evidence requirements describe is the minimum record an organization needs to explain why it accepted a finding over the alternatives it considered.
A finding that later proves wrong gets a dated, traceable correction under your organization's record-control process, saying what changed and on what evidence. Readers who already quoted the original need to be able to find the change.
A timeline that keeps its labels
The timeline is one place a hypothesis can turn into a finding without anyone deciding it should. Somebody in the channel wrote "looks like the pool" at 14:06. The timeline, assembled on Thursday from the channel history, records it as the thing that happened at 14:06, and by the time the document circulates it reads as established.
A timeline that tags every entry with a marker reduces that drift. Five markers cover the live entries here, and teams can add their own: observed, inferred, decided, actioned, and result. The first two say how an entry is known, and the last three say what kind of event it records. The analysis entry below adds a sixth, analyzed, for results derived after the fact. Every entry carries a marker, and an unmarked entry counts as incomplete and stays out of the timeline that circulates. Marking an observation as inferred to avoid owning it misreports it in the other direction. Each entry gets a UTC time, a marker, the statement, and where it came from.
13:51:00Z | ACTIONED | release 4.2 deployed, worker count 8 to 16 | deploy log
13:56:00Z | ACTIONED | upstream timeout change deployed | deploy log
14:02:13Z | OBSERVED | checkout error rate above threshold | checkout probe
14:06:40Z | INFERRED | consistent with pool exhaustion; timeout change also in window | on-call lead
14:09:05Z | OBSERVED | active connections at pool ceiling on all pods | pool metrics
14:12:30Z | DECIDED | roll back to release 4.1 (worker count 8); keep timeout change, reverting it needs a second deploy | incident lead
14:15:02Z | ACTIONED | release 4.1 restored, worker count 8 | deploy log
14:18:40Z | OBSERVED | active connections back under pool ceiling | pool metrics
14:18:40Z | OBSERVED | checkout error rate back under threshold | checkout probe
15:24:30Z | OBSERVED | pending order-job queue drained | queue depth metric
15:25:00Z | RESULT | checkout latency at baseline | checkout probe
21:25:00Z | RESULT | six hours at baseline | checkout probe
Analysis entries, stamped with when the analysis ran:
D+1 10:40Z | ANALYZED | pool acquisition timeouts match failed checkout requests by request ID, 14:02 to 14:18:40 | application log joined to load balancer log, platform lead
Three things show up in that format that a narrative hides. The inference at 14:06 stays an inference, with a name next to it, and the observation at 14:09 that supports it has its own line. The decision at 14:12 records who made it, what they chose to leave alone, and why, which is the detail a reviewer needs when asking why the timeout change stayed in place. And the gap between the action at 14:15 and the result at 15:25 is accounted for: the pool recovered by 14:18:40 and the order-job queue took until 15:24 to drain. Analysis done after the incident gets its own entries, stamped with when it ran, so a later finding never reads as something seen at the time. When a decision gets reversed, the reversal goes in as its own entry with its reason, and the original stays where it was. For the live-incident side, part one covers the clock problem underneath all of this: which clock the timeline follows, and what the sampling interval of each source leaves unseen.
Findings that stop at a person
Somebody made the change, and everyone in the review knows who. The timeline can record that as a fact, by role or by name as your organization's practice allows. The finding is a different sentence. A finding that ends at a person turns a fact about who acted into the explanation for why the system failed, and it stops the investigation at the point where the useful questions start. What made the change look correct at the time? What validation would have caught it? What in the interface made the wrong thing the easy thing? Those questions produce controls. A finding that names the operator produces a name in a document that outlives the incident. The blameless-postmortem literature makes the same argument: Google's SRE book describes a review that indicts an individual as stopping short of the contributing conditions that would inform a fix [1], and Allspaw's account from Etsy argues that punishing the operator discourages people from sharing the detail the organization needs to see the system clearly [2]. All of this concerns ordinary operational mistakes. Deliberate misconduct or misuse of access belongs to the organization's designated investigative and personnel processes, outside the engineering postmortem.
A finding that names a person has relatives that look more respectable. They name a role or a lapse instead of a person, and they end the inquiry at the same spot. Each has a better question behind it.
| Instead of | Ask |
|---|---|
| The operator restarted the wrong instance. | How did the naming, the target selection, or the confirmation step let two instances look interchangeable? |
| The developer made a bad config change. | Which of validation, staging, review, rollout, and rollback had a chance to stop it, and why did each one let it through? |
| On-call missed the alert. | Was the alert routed, actionable, and distinguishable from the noise around it, and what else was competing for attention during that shift? |
| The runbook was ignored. | Was it current, findable, written for these conditions, and tested against them? |
| Monitoring missed it. | Which user-facing objective had no signal, and which coverage assumption turned out to be wrong? |
| Automation failed. | Was the evidence incomplete, the policy ambiguous, the state model wrong, the action unscoped, or the verification too weak? |
| The network team broke it. | Where did the handoff between the network and application teams leave the change without an owner who could see both sides? |
A postmortem that answers the right-hand question can produce a control, or an explicit decision to accept the condition. One that stops at the left-hand statement leaves the engineering causes and controls unexplored.
Describing prevention by what the evidence shows
A manager reading a postmortem may need to know whether the control covers the failure mode and what remains exposed. The answer comes in two records: what the team intends, and what the evidence shows. Naming the state of each is most of the work.
| Record | State | What you can write | Example |
|---|---|---|---|
| Intent | Identified | We identified an improvement. | We identified the need to check worker count against pool capacity before deploy. |
| Intent | Planned | It is planned, with an owner and a date. | The platform team plans to add that check in the next scheduled release. |
| Evidence | Implemented | It is in place. | The check has been part of the deployment pipeline since that release. |
| Evidence | Tested | It was validated against a specific case. | Replayed against the pipeline, the release 4.2 change was blocked. |
| Evidence | Observed in operation | Since it went in, it has been seen working. | In the four weeks after it shipped, it blocked two changes of that shape. |
Intent and evidence advance separately. Evidence states accumulate as milestones, so record every one reached, each with its date. Expected effect travels alongside any of these states as a prediction, written as one: "it should catch any deployment where worker count times connections per worker exceeds the pool ceiling." One failure to watch for is writing an expected effect in the language of an observed one, which readers hear as a promise. When a follow-up review changes either state, append a dated entry that records the new state and the evidence behind it, and leave the original sentence as written, so the record still shows what was known when it was published.
The same discipline applies to the follow-up language around a prevention claim:
| Instead of | Try |
|---|---|
| This will never happen again. | We have identified a detection improvement that would surface this condition earlier. |
| We implemented a permanent fix. | We added a check that fails any deployment where worker count times connections per worker exceeds the pool ceiling, and tested it against the release 4.2 change. Other paths to pool exhaustion remain open. |
| We need to be more careful. | We will add a confirmation step for production targets in the deployment tool, owned by the platform team and tested before the next scheduled release. |
| This was unavoidable. | Given what the team could see at 14:12, rolling back the worker count and leaving the timeout change in place was a reasonable choice. The review will look at what would have made the problem visible sooner. |
| Closed, no further action. | The incident exit criteria are met. Two follow-ups remain open in the tracker, each with an owner and a date. |
For the live-incident version, part one covers the other common prevention sentence, the scheduled job that keeps a problem at bay, and when that job is a legitimate stopgap.
Action items that survive the meeting
The prevention states describe what a sentence can claim. The action item is the record that changes them. "Improve monitoring" has no owner, no date, and no way to tell when it is done.
An action item that survives the meeting carries six things:
| Field | What it holds |
|---|---|
| Condition addressed | The specific contributing condition this item changes, taken from the findings. |
| Owner | A named person or team, accountable for the item. |
| Due date | A committed date. Deferred work keeps its own record with a reason, a new date, and who or what authorized each deferral under your organization's process; if the risk owner named by your process accepts the risk instead, that decision goes in the residual risk section. |
| Success measure | What improves if it works: detection time, recurrence, a blocked change, a coverage gap closed. |
| Validation method | How the team will show it works: a replayed change, a test, a game day, a measured trend. |
| Closure milestones and evidence | The milestones declared when the item was created, among implemented, tested, and observed in operation, and the evidence for each one met. Effectiveness can be its own follow-up, with its own date. |
The closure field is the one most worth settling at creation. Declare which of those milestones the item needs, and track each with its own date. The item is complete when every declared milestone is met. An item closed because the change merged shows implementation only. An item closed with a replay that shows the check blocking the original change shows a tested control, and the follow-up entry can say so.
Residual risk is a finding too
When a postmortem closes with something still open, such as a failure mode the new check leaves uncovered, a dependency nobody could test, or a workaround that stays in place until next quarter, write it down. Writing it down turns an unstated exposure into a recorded decision with an owner and a review date. A residual risk entry names what remains exposed and what happens if it occurs; its historical occurrence or a likelihood assessment, with the basis; what holds it down in the meantime and where that control falls short; who accepted it; and what brings the review forward.
| Instead of | Try |
|---|---|
| We'll live with it. | The platform lead recommended accepting the remaining exposure to slow-query pool exhaustion. If it occurs, checkout errors return until the burst ends; it has happened once in the last year. The interim control is an alert on pool wait time, which fires after two minutes, so the first two minutes go unprotected. A second occurrence before review escalates to the checkout service owner under the incident procedure. The checkout service owner, who holds that authority under the documented risk-acceptance procedure, accepted it until the connection review on March 1, 2027. |
| It's low risk. | The failure mode needs a slow-query burst during a deployment, a combination seen once in the last year. If it occurs, checkout errors return for the length of the burst. The pool wait-time alert is the only interim control, and it fires after two minutes, so the first two minutes go unprotected. A second occurrence brings the review forward. The checkout owner, who holds acceptance authority under the risk process, accepted it, with a review date of March 1, 2027. |
| Known issue. | Tracked as residual risk in the incident record: the reporting job can still read stale data during failover, which would put up to fifteen minutes of stale figures in the morning report. Failover happens about twice a year. The interim control is a banner the report shows after any failover, which relies on readers noticing it. The data platform owner accepted it under the risk process, and any failover before the next test brings the review forward. Review: next failover test. |
| Nothing else to worry about. | Two residual risks remain open, each listed below with its exposure, likelihood, interim control, owner, and review date. |
The person who accepts a risk should be whoever your organization's risk process names as able to accept it, and the entry says who that was. An engineer can recommend acceptance. The decision belongs to that owner, and the postmortem records it that way.
One document, two annexes
A postmortem gets read by the engineers who will change the system and by the managers and directors who will fund the changes, and those two readers need different depth. Two separate documents can drift apart when the shared facts are maintained in both. One workable pattern, when both audiences can share access to the same evidence, is one document with a shared core and two annexes.
| Part | Holds | Written for |
|---|---|---|
| Shared core | Capability impact and duration, the main decisions and when they were made, contributing conditions across teams, action items, and residual risk. | Everyone, including readers who need the decision record without the technical detail. |
| Technical annex | The labeled timeline, telemetry, configuration and change detail, and the evidence behind each finding. | The engineers who will act on it. |
| Business annex | Affected transactions or processes, backlog and workaround use, deadlines at risk, and commitments their owners identified as affected. | Process owners, finance, and the people deciding what to fund. |
Two sections are worth including in the core whenever they are material. The first is alternatives considered: which actions the team rejected during the incident, and why. A reader who later wonders why nobody simply rolled back the timeout change finds the answer there instead of assuming carelessness. The second is what made the work harder: the missing dashboard, the runbook written for a different version, the access request that took forty minutes. Those are findings about the system, and they produce action items tied directly to system conditions.
Every figure in the business annex carries its source, unit, time window, scope, as-of time, and status. The status is observed, estimated with its assumptions stated, or unknown, and the figure is marked confirmed once the process owner verifies it. The annex is one place estimates can lose their labels on the way into a quarterly slide.
A paragraph, rewritten
Here is the paragraph from the opening, rewritten against the evidence requirements as it could stand while the review was still open, before the reproduction behind the finding quoted earlier:
Checkout errors rose at 14:02 and cleared by 14:18:40; latency returned to baseline at 15:25. The observed behavior is consistent with connection pool exhaustion following a deployment that raised worker count from 8 to 16, each holding up to two connections under load, with the pool ceiling unchanged at 20. We have not ruled out the upstream timeout change deployed in the same window; a staging reproduction with each change applied alone and with both together is under way.
Expected behavior returned for the checkout path at 15:25 and held through the six-hour stability window that closed at 21:25. Customer impact is under assessment. Between 14:02 and 14:18:40, 10,412 checkout requests returned a 5xx response, with none after that through 15:25. The count comes from load balancer logs complete for that period, deduplicated by load balancer request ID. Each client retry gets a new ID and is counted separately, and the count of distinct affected customers is still undetermined. Reconciliation against the order ledger, owned by the order data team, is due on the fifth business day after the incident, with the order data team lead escalating to the checkout service owner if it slips. No evidence of data loss has been identified in the order records from 14:02 to 15:25 reviewed so far. Payment records are still unreviewed.
The change passed the pre-deploy checks in place at the time, none of which compared worker count against pool capacity. A check that fails any deployment where worker count times connections per worker exceeds the pool ceiling has been implemented. Replayed against the release 4.2 change, it blocked the deployment. That shows the check catches this configuration, whatever the reproduction concludes about cause. It leaves other paths to pool exhaustion unaddressed.
Claim by claim, here is what changed:
| Claim | In the first draft | In the rewrite |
|---|---|---|
| Cause | "Caused by" a pool exhausted by the deployment | A mechanism consistent with the evidence, a named alternative, and the test that will separate them |
| Resolved | "Now fixed" | A recovery time and the six-hour stability window that closed at 21:25 |
| Customer impact | "No customer impact" | 10,412 failed requests, counted and scoped, with distinct customers still undetermined |
| Data loss | "No data loss occurred" | No loss found in the order records reviewed, payment records still unreviewed, and reconciliation due with an owner |
| Attribution | "Human error" | The control gap: no check compared worker count with pool capacity |
| Prevention | "Will prevent recurrence" | A tested check, with other paths to pool exhaustion named as open |
Each row trades a closed conclusion for something the next reader can check.
The review question
Part one ends with a sixty-second test for live updates, where speed is the constraint. A postmortem gets a slower check, and it is one question applied sentence by sentence: what would we point at?
Asking whether a sentence is true leaves its support unstated. Asking what we would point at shows where the evidence stops. Often the answer comes back partial: evidence that covers half the sentence, or covers it for one region only. Whatever the evidence leaves out gets cut from the sentence or gets its own evidence.
Three further checks belong to the postmortem alone. Each finding states its scope. Each open alternative says what would settle it. And the closure state matches the exit criteria recorded for this incident, with the stability window that was actually observed.
As a single list for the review meeting:
What would we point at for each sentence?
What scope does each finding cover?
Which alternatives remain open, and what would settle them?
What does resolved mean for this incident, and which window was observed?
What evidence shows each control works, and at which milestone?
What remains exposed, who accepted it, and when is it reviewed again?
References
- [1] Google SRE, Postmortem Culture: Learning from Failure
- [2] John Allspaw, Blameless PostMortems and a Just Culture, Etsy Code as Craft, 2012
Sources verified 2026-09-29.