Say Only What the Evidence Supports: Incident Language for the People Who Decide
Somebody senior asks how many customers are affected. A dashboard shows ten thousand failed requests. Those are two different numbers: one customer can fail many requests, and a count of requests leaves open how many distinct customers were behind them. The useful answer gives the measurement you have and names the one you are missing, in the same update, before anyone fills the gap with an inference.
Incident updates go wrong in two opposite directions. One claims more than the evidence carries, and the correction reaches the same people later, at a worse time. The other reports accurate detail the reader has no way to act on. Both problems need the same discipline: make the strongest claim the evidence supports, and translate it into the decision the reader owns.
This piece is written from the IT side of the table, for administrators and engineers who have to explain an incident to the people who run the business: operations directors, managers, senior directors, the VP deciding whether to call a customer. It covers wording, evidence, and what each reader needs. It offers no legal advice. Where a question carries legal, contractual, or regulatory weight, the practice here is to hand it to the people who own that question.
Most of what follows applies to what gets written down while the incident is running: the status update, the ticket, the business summary. The bridge call runs on different rules, because at minute six of a severity one it needs short sentences and fast decisions. The written record is where the labeling pays off, because it outlives the call and reaches people who were never on it. Every example here is constructed: the incidents, times, and figures are placeholders chosen to show the pattern.
This is the first of two pieces. The postmortem gets its own treatment next week in Say Only What the Evidence Supports in the Postmortem, because it is a different document with different readers and a longer life, and the claims in it need more evidence.
Evidence is the floor
Engineers who take evidence seriously can make the opposite mistake next. Having learned to say only what they can prove, they prove everything, to everyone. The director asks whether the end-of-day batch will finish and gets a paragraph about queue depth, three graphs, and a hedge. The update is accurate and leaves out the one thing the director needed: a decision, with a deadline on it.
Saying only what the evidence supports sets the floor. The message gets built on top of it, and building it means translation. In the working model used here, an update that lets a director, a senior manager, or a VP running an operation act carries four things: what is affected in terms of the business, how sure you are, what you need from them and by when, and when they will hear from you next. The evidence stays attached for whoever wants to check it. It stops being the message.
One more habit belongs on that list. Say what the team is deliberately holding off on, and why. "We are letting the snapshot finish before we restart anything, because we have yet to confirm what cancelling a snapshot mid-run does on this array" explains restraint before someone upstairs mistakes it for inaction and asks for something faster and riskier.
Five labels for an impact statement
These five labels are the impact vocabulary of this piece, and your team may already use other words for the same ideas. Three of them say how a figure was produced: observed, estimated, or unknown. Every impact figure gets one of those three. The other two are qualifiers added when they apply: confirmed means the owner of the affected process has checked it against their own records, and correlated means it moved together in time with another observation, with no cause established. Put the label in the sentence, so the number keeps its meaning after you send it.
| Label | What it means | Example | When to use it |
|---|---|---|---|
| Observed | Directly measured | "Order submission returned errors for 41 minutes." | Whenever the measurement directly represents the claim |
| Estimated | Modeled from stated inputs | "Approximately 4,500 to 7,500 sessions were affected: 3 to 5 percent of sessions failed in a four-minute sample, applied to the 150,000 sessions in the window, assuming the sample is representative of the full window." | With the inputs, the assumption, and the range stated |
| Unknown | Not yet determined | "We have not determined how many distinct customers experienced a failure." | Whenever the alternative is a guess |
| Confirmed (qualifier) | The owner of the affected process has checked the figure against their own records | "Finance checked the count against its records: 212 invoices failed to transmit." | Only after that owner confirms it, inside or outside engineering |
| Correlated (qualifier) | Moved together in time with another observation, cause not established | "Support contact volume and checkout errors rose over the same window." | When the link is visible but untested |
An explicit unknown is preferable to an unlabeled guess. A stated unknown is a fact about the state of the investigation, and it stays accurate as a record of that moment. A guess can be overturned by the next measurement, and the correction costs more than the delay would have.
A common error is presenting an estimate as observed by dropping the assumptions. "Roughly six thousand sessions failed" and "an estimated 4,500 to 7,500 sessions failed, applying the 3 to 5 percent failure rate from a four-minute sample to the whole window" start from the same calculation. The first collapses the range into its midpoint and drops the inputs. The second one tells the reader what to verify, starting with whether the four minutes were typical of the window.
What each audience is asking
Different roles need different evidence density. The work is picking what serves the decision in front of each reader, which is a different job from simplifying.
| Audience | The real question | The decision they own | What they need | What can wait |
|---|---|---|---|---|
| Engineering peers | What is happening technically? | Which bounded diagnostic or recovery step is safe next | Raw signal, logs, specific fault layer, current hypothesis | Business framing |
| Team lead | What does this cost my team and what is blocked? | Where to put people and what to pause | Service state, capacity, blockers, next milestone | Packet-level detail |
| Operations director | What is the cross-team risk? | Which priorities and escalations to set across teams | Affected capability, dependency health, objective risk | Individual command output |
| Business process owner | Can my process run, and will my deadline hold? | Whether to start the workaround and how to protect the deadline | Volume affected, backlog, workaround, deadline risk | Uptime percentages without process consequence |
| Security | Is this unauthorized activity or a security control concern? | Which containment and investigation steps to take | Indicators, scope, evidence location and preservation status, containment status | Incident detail outside their scope |
| Privacy, compliance, and risk | Was regulated data exposed, and what do the owners of that assessment need? | Whether handling, preservation, or notification steps apply | Data classification, exposure confidence, evidence preservation | Unreviewed technical hypotheses |
| Executive | Do I need to decide something? | Risk acceptance, resources, and customer-facing priorities | Customer impact, financial exposure, the status of any legal or regulatory review as its owner reports it, recovery timeline | Component names |
Organize each update around the decision its reader owns, and watch the last column just as closely. Handing an executive a stack trace leaves the translation to them. Giving risk and compliance an unreviewed hypothesis about data exposure is worse than unhelpful, because an unlabeled hypothesis can be copied into their record as fact and then repeated as though it had been established.
One fact at three altitudes
Take a single fact from the middle of an incident. Since 14:02, write latency on the volume behind the order database has been running around 40 milliseconds against a normal 5, and orders are processing slowly.
For the engineering channel:
Order DB volume write latency at 40 ms against a 5 ms baseline since 14:02, per the array performance view. Queue depth elevated on two of four paths. Current hypothesis is the snapshot job that started at 14:00. Checking the snapshot schedule and path balance next.
For the operations manager:
Order processing and the storage underneath it both slowed sharply at 14:02. The likeliest explanation is a scheduled snapshot job. We plan to let the job finish, because we have yet to verify how this array handles a cancelled snapshot. The storage lead is checking that procedure. Once the job ends, we check whether latency drops. Next update at 15:00.
For the director or VP:
Order processing has been slow since 14:02, and the 18:00 end-of-day batch is at risk. We have a likely explanation and expect to know more when a scheduled job finishes around 14:50. If latency stays high after the job finishes, the fallback is to hold noncritical jobs tonight, and that needs an answer from the order operations lead by 16:00, with 16:30 as the final cutoff. Next update at 15:00.
All three trace back to the same technical measurement. The engineer gets the number and the hypothesis, the manager gets the consequence and the next step, and the director gets the business risk and a decision with a deadline on it. The manager and director versions add business facts, such as the batch deadline, that come from the business side, and none of the three states the explanation more strongly than the evidence allows.
When the measurement was taken
Every number in an incident record was captured at some moment or over some window, by some collector, at some interval. That detail rarely makes it into the sentence, and it constrains what the sentence can claim.
A point-in-time reading describes one instant. A health check that passed at 14:00 says the service answered at 14:00, and it says nothing about 14:00:20. A persistent failure is likely to show up in the next reading. Intermittent failures behave differently. A fault that lasts fifteen seconds and recurs every few minutes can land between one-minute health checks for long stretches, and a status page built on those checks shows green while customers see errors.
Averages hide the same failure another way. With requests arriving at a constant rate, a service that fails every request for thirty seconds and serves normally for the other four and a half minutes reports a five-minute success rate of 90 percent. That is arithmetic, and on a five-minute chart it can look like a shallow dip. The customer who tried during those thirty seconds got a complete failure. Both descriptions cover the same window, and only one of them describes what the customer saw.
Collection timing shapes trends too. Suppose disk latency is sampled every five minutes, on the five, and the nightly backup runs from 02:01 to 02:04. Under that fixed schedule no sample lands inside the backup, so two weeks of data produce a flat line, and a flat line reads as evidence of stability. Every point on that chart is accurate, and the chart still says nothing about the backup. Real schedulers may drift, so a sample can occasionally land inside the window, and the result shows up as a lone spike that looks like noise.
Then there is state captured after the fact. The connection counts, queue depths, and memory figures you collect at 15:10 describe 15:10. If the failure ended at 14:40, those readings describe the recovered system, and a write-up that quotes them as the conditions during the outage is describing the wrong moment.
Finally, the sources disagree about the time itself. The application log, the storage array, the load balancer, and the incident channel each stamp events from their own clock, sometimes in their own time zone. Lining them up is part of building the timeline, and the incident record should say which clock its timeline follows and how far apart the sources were when you checked.
The fix is to make the collection method part of the claim, and three habits catch many of these failures. State the interval with the reading, because "health checks every 60 seconds showed no failures" is a claim a reader can weigh, and "the service was healthy" hides the interval. Report the worst window next to the average whenever the failure came and went. And stamp captured state with its capture time, keeping it apart from anything measured during the event.
Declare early
In the model used here, the status words in the next section start applying when someone declares an incident, and that decision is easiest to make early. An incident record, an owner, and an update schedule cost little to set up, and they are hard to retrofit once three teams have started changing things in parallel.
Common triggers for declaring, to be tuned to your own thresholds and service levels: a service objective is missed or about to be; the scope or cause is unclear enough that more than one person needs to investigate; several people or teams could make conflicting changes; a business process owner or an executive needs status; data integrity or security may be involved; or a broad action, such as a rollback or a traffic shift, needs someone accountable to approve it.
Declaring signals that coordination is now worth the overhead; what else it triggers depends on your process. If the problem turns out to be small, the record closes quickly with a short note, and the director heard about it from you first instead of from a user.
What each status word can carry
The status line is often the first thing a director reads in an update, and sometimes the only thing. The words below are a working vocabulary, and teams with their own states can map theirs onto it. Each word can carry a specific amount of certainty and no more, so the word has to match what the team can show.
| Status | What it can carry | What nobody can conclude from it yet |
|---|---|---|
| Investigating | A condition has been detected. Scope, impact, and mechanism are under assessment. | The cause, the full impact, or a recovery time. |
| Assessing impact | Part of the scope is confirmed or estimated, and the rest is named as open. | The full business, financial, or contractual effect. |
| Mitigating | Impact has been reduced through a workaround, a traffic change, or added capacity. | That the underlying condition is corrected. |
| Restoring | A bounded change is being applied or reversed, and expected behavior is being checked. | That the service is stable, or that the business process has caught up. |
| Monitoring | Technical behavior is back in its expected range, and the stability window is running. | That the incident is over. |
| Technical recovery complete | The service is back and has held through the stability window. Business recovery continues under an owner who accepted it, with a next checkpoint named in the update. It can precede Resolved, which depends on the exit criteria declared for this incident. | That the business process has caught up. |
| Resolved | The agreed exit criteria are met, including any business recovery they cover. | That every related risk is gone. |
Statuses can overlap, since mitigation often runs while impact is still being assessed, so report the primary state and name any other still active. A recurring source of status confusion is using a status to claim recovery the team has yet to show. Mitigating gets reported as over because the graph turned green, and the director who hears it cancels the manual workaround the order team was still running. Keeping each word attached to what the team can show prevents that, and it gives the reader a vocabulary they can learn after two or three incidents.
Name the action you took
The second vocabulary problem is the verb. Engineers under pressure use one phrase for everything they did, and the phrase is usually some version of "it's good now." Directors then hear the same word for a traffic reroute, a workaround, and an actual repair, and plan accordingly. A short set of action words, used the same way every time, keeps those apart. The first three describe approaches to limiting the damage, and restoration describes the return of the capability itself.
| Word | What it means | How it gets misused |
|---|---|---|
| Containment | Limits the scope of the problem or prevents further impact while the investigation continues. | Presented as the repair. |
| Mitigation | Reduces the impact on people and processes before the condition is corrected. | Taken as proof of the cause. |
| Workaround | An alternate path that lets the business keep working. | Described as though customers saw nothing. |
| Restoration | Returns the capability to a usable state. | Declared after one successful request. |
Once readers know these words, a director can tell reduced impact from restored service, and the rest of the update can say whether the workaround can come off. Claims about preventing recurrence belong to the postmortem, which is where the second piece picks up.
The order of an update
Under pressure, people write updates in the order things happened to them: what they tried, what they saw, what they suspect. Readers need a different order. A default sequence, to adapt to your own incident process:
- Time and status, using one of the status words above.
- Scope: which capability is affected, where, and for whom.
- What was observed, and from which source.
- Impact, with its label from the five above.
- The current explanation, labeled as a hypothesis, with what would strengthen or weaken it.
- What the team is doing, and what it is deliberately holding off on.
- What has to be true before anyone calls it over.
- What you need from this reader, and by when.
- When they hear next.
For a director, lines 1, 2, 4, 8, and 9 usually carry most of the value, plus one line of evidence when it bears on the decision, and the rest can sit below a break or in the linked record. For the engineering channel, lines 3, 5, 6, and 7 carry the weight. Keep the field meanings the same for every audience, even when the order or depth changes, so a reader who has seen a few of your updates knows where to look.
Line 9 is worth sending even when nothing has changed. Stakeholders left without a next-contact time often start asking for status through side channels, and those requests land on the people doing the work. Naming a time, and meeting it with "no change, next update at 15:30" when that is the truth, gives them a predictable checkpoint instead. Set the cadence by how fast the reader's decision moves, and change it out loud when conditions change.
As a fill-in template:
Time and status:
Affected capability and scope:
Observed, with source:
Impact state (observed, estimated, or unknown):
Qualifier (confirmed, correlated, or none):
Inputs and assumptions for any estimate:
Current explanation (hypothesis; what would weaken it):
Doing now, and holding off on, with the reason:
Exit criteria:
Decision needed, owner, and deadline (or: no decision needed):
Next update:
One incident, four updates
Take the storage incident from the three-altitudes example and follow it through an afternoon as a series of updates to the operations director.
At 14:10, enough is known to coordinate and far too little to conclude:
Investigating. Order processing slowed sharply starting at 14:02. Checkout and order entry are affected in all regions, and reporting looks normal. Storage latency on the order database is running about eight times its usual level. Impact on the 18:00 end-of-day batch is under assessment. Next update at 14:30.
At 14:30, the team has an explanation and says what would prove it wrong:
Assessing impact. The slowdown lines up with a scheduled snapshot job that started at 14:00 on the same storage. That is our leading explanation, and it is badly weakened if latency stays high after the job finishes, expected around 14:50. Orders are completing at roughly four times their normal duration, per the order service's own timing over the last five minutes. At that pace the 18:00 batch is at risk. We are letting the job finish instead of stopping it midway, because we have yet to confirm what cancelling a snapshot mid-run does on this array. The fallback is to hold noncritical jobs tonight. That needs an answer from the order operations lead by 16:00, with 16:30 as the final cutoff. Next update at 15:00.
At 15:00, an action has an outcome, and the outcome is still being checked:
Monitoring. The snapshot job completed at 14:52 and storage latency was back in its normal range by 14:55. Orders are processing at normal speed, and the backlog from the slow period is draining. The order team expects it cleared by 15:40. The exit criteria for this incident are latency holding normal through a thirty-minute window counted from 14:55, and the order team confirming the backlog cleared. The fallback matters only if orders stay slow. If latency holds through 15:25 and order throughput stays normal, we withdraw the 16:00 fallback request then, with a one-line note. Next update at 15:40.
At 15:40, the exit criteria are met, and the update says what remains:
Resolved. Storage latency has held normal since 14:55, and the order team confirmed the backlog cleared at 15:37. Both exit criteria are met. The batch has 12,000 orders left and is processing about 9,000 an hour, which puts completion near 17:00 at the current rate. The order team adds a thirty-minute planning buffer, to 17:30, still ahead of the 18:00 cutoff. The timing fits the snapshot explanation without proving it. What remains: an investigation item, owned by the storage team and dated in the incident record, to establish whether the snapshot job produced the contention. Any change to the snapshot schedule waits on that result. The 16:00 fallback request was withdrawn at 15:25, when latency and order throughput both held. A written review follows later this week.
Read the four in sequence and the labels move. Impact starts as under assessment and ends confirmed by the people who own the process. The explanation arrives with a condition that would weaken it, and the next update reports that latency recovered as expected, which keeps the explanation as the leading hypothesis without proving it. In the second update, the restraint comes with its reason. The decision request carries a deadline, stays open until the stability window shows it is unneeded, and then gets withdrawn out loud at 15:25, so nothing is left pending on the decision owner's side. A change to a decision request goes out the moment it happens, outside the update schedule. And the last update names what remains before anyone asks.
Asking for a decision
Line 8 of the update order, what you need from this reader, is the line most updates leave out, and it is the one a director can act on. When engineering needs a business call, such as starting a workaround, pausing a batch, accepting a slower recovery, or approving a riskier one, the request works when it carries five parts: the options, what each one costs, who owns the call, when it is needed, and what happens if nobody answers in time. The 14:30 update flagged the fallback in one sentence. The request that went to the order operations lead alongside it carried all five:
Decision needed: hold noncritical jobs tonight, or run them as scheduled.
Owner and timing: the order operations lead. Answer requested by 16:00, when escalation begins. 16:30 is the final cutoff, after which holding the jobs frees too little of the batch window to change the outcome.
Hold them: frees roughly a third of storage throughput during the batch window. On five comparable nights last month with those jobs held, the batch finished between 17:20 and 18:20, four of the five before 18:00. The nightly reports move to tomorrow morning.
Run them: the reports stay on time, and the batch finishes near 20:00, which delays the morning shipping manifests that depend on it.
Recommendation: hold them. A late batch delays shipping for every region, while late reports affect one internal team. Tonight's workload mix is unmeasured, so treat the range as directional.
No answer by 16:00: the incident lead escalates to your designated alternate, and the jobs run as scheduled until one of you decides.
The no-answer line does the most work, and it has to be written with care. Silence at 16:00 can mean the decision owner is in another meeting, never saw the message, or disagrees, so it is no substitute for approval. The safe pattern is to name who gets asked next and to keep the current, already-authorized state until someone with authority decides. A default that changes something is appropriate only when your incident procedure already grants that authority, and the request should say which procedure grants it.
Two habits keep decision requests out of trouble. Send the request to the person who owns the decision, which the audience table maps, and resist sending it to whoever is loudest in the channel. And keep engineering's recommendation on its own line, labeled as a recommendation, so the decision owner sees both the real options and what the team would pick.
Recovered has two halves
The engineering definition of recovered is technical: the service meets its objective again and has held through the stability window. The business definition asks different questions. Orders are being accepted, the backlog from the outage has cleared, the manual workaround has been switched off, and anything processed during the incident has been reconciled. The two usually finish at different times. An update that reports the first as though it were the second is how a director tells the order team to stand down while their queue is still full.
Sometimes the right first move is on the business side. When recovery will take hours and a workaround exists, starting the workaround early can protect the deadline while engineering works, provided its risks and reconciliation work are understood and the process owner accepts them. An update that proposes the workaround early does more for the director than one that reports the repair late.
When the incident might involve data
Any sign of suspected unauthorized access, unintended modification, loss of integrity, or unexplained deletion changes who should be involved and how information is handled, with the exact procedure set by the organization's owners. From that moment, the questions that matter most belong to other people: security, privacy, the data owner, and whoever handles notification. The engineering job narrows to three things. State what was observed, with times and scope. Keep the logs and records as your organization's handling procedure directs, and leave decisions about retention to the people who own that procedure. Bring the owners in early, with the facts at the strength the evidence supports.
What changes in the writing is mostly subtraction. Speculation about exposure stays out of broad channels, because a guess repeated in a large channel becomes the version people remember. Evidence stays where the handling procedure designates, instead of being pasted into chat. The update to leadership says what the owners are assessing, and leaves the conclusion to them.
| Audience | What the update carries |
|---|---|
| Engineering and security | The observed condition with times and scope, and which restricted record holds the retained evidence. Handling locations and access controls stay in that record, under the procedure its owner sets. A reminder to avoid inferring past the facts. |
| Process and risk owners | That a condition needs assessment for integrity or confidentiality, what was contained, and which owners are engaged. |
| Director or VP | The confirmed facts, that the authorized assessment is under way, the status of that assessment as its owners report it, and the next update time. |
| Customers or the public | Nothing drafted in the incident channel. Only language approved through the organization's established customer or public communication process. |
Saying it out loud
The left column here, like the library that follows, is written for the record. On a bridge call or in a fast message to an executive, the same discipline has to fit in one breath. The test for a spoken version is whether a tired incident lead can say it without dropping the scope or the next step.
| For the record | Out loud |
|---|---|
| Error rates returned to baseline at 14:19. Latency is still elevated while queued work drains, and we expect that to finish around 15:25. | Errors are back to normal. The queue is still draining; we expect it clear around 15:25. |
| The 14:02 configuration change is our leading hypothesis. It fits the timing, and it has yet to be tested. The test finishes by 15:00. | Best guess is the 14:02 config change. Untested. We'll know by three. |
| Reconciliation is incomplete. As of 15:30, no evidence of data loss has been identified in the order and payment tables reviewed so far. | No data loss found yet, and we've only checked orders and payments so far. |
| Answer requested from the order operations lead by 16:00 on holding noncritical jobs, with the five nights of run times as the basis. | We need a yes or no from the order ops lead by four on holding tonight's noncritical jobs. Holding them got the batch out on time four nights out of five last month, and tonight's mix may differ. |
The short versions keep the words that carry the evidence: still draining, best guess, untested, found yet, only checked, four nights out of five. Each one does the work of a sentence, so the spoken update stays honest at speed.
The translations
From here the piece becomes a reference library: 79 rewrites in twelve categories, meant to be skimmed now and returned to later. The categories: sounds careful but still overclaims, cause and diagnosis, impact and scope, data and integrity, action and recovery, time and commitment, vendors and dependencies, time windows and intermittent behavior, security and access, changes and maintenance, workarounds and continuity, and up the chain. Each row swaps one claim for another rather than softening it. The right column usually says less, and it can be defended when somebody checks. The useful part is the shape of each sentence.
Sounds careful, still overclaims
| Instead of | Try |
|---|---|
| We believe the root cause is the 14:02 configuration change. | The 14:02 configuration change is our leading hypothesis. It fits the timing, and it has yet to be tested. The test finishes by 15:00, and the next update will say whether it held. |
| Preliminary analysis suggests the issue is resolved. | Error rates have been at baseline since 14:40, twenty minutes so far. Resolved waits on the thirty-minute stability window, which ends at 15:10. |
| There is no indication of data loss at this time. | We have yet to check for data loss. Reconciliation of the order tables starts at 15:00 and takes about an hour. |
| Impact appears to be limited. | Errors are confirmed on checkout in the primary region. The secondary region and the mobile apps are unassessed so far, and both are next. |
| The issue has been mitigated. | Traffic moved off the failing node at 14:30, and the error rate dropped from 12 percent to 2 percent. The node itself is still failing, and the remaining 2 percent is under investigation. |
| Out of an abundance of caution, we rolled back. | We rolled back release 4.2 at 14:15 because error rates were climbing and the release was the most recent change. Whether the release produced the errors is still being tested. |
| This was an isolated incident. | Our logs show no earlier occurrence in the last 30 days, which is as far back as we retain them. Anything older is outside what we can check. |
| All systems are operating normally. | All 14 monitored services report baseline error rates and latency as of 15:30. The overnight batch jobs report on completion, so their state is confirmed at 18:00. |
Each sentence on the left sounds responsible, and each hedge word in it carries no evidence. "Believe," "suggests," "no indication," and "appears" tell the reader how the writer feels and leave out what was checked. The version on the right replaces the hedge with the check, its scope, and when the next result lands.
Cause and diagnosis
| Instead of | Try |
|---|---|
| I found the cause. | We have identified a mechanism consistent with the observed behavior. Two alternatives are still open. The checks that test each finish by 15:00, and the next update will say which explanations remain. |
| It was a bad config change. | A configuration change at 14:02 lines up with the onset. We are validating whether it is the mechanism or a coincidence of timing. |
| The deployment broke it. | Errors began eleven minutes after release 4.2 went out, and the logs show connection timeouts from that release. We are confirming whether the release is the mechanism. |
| A bug took the system down. | Testing reproduced a defect in the retry handler under the load we saw at 14:02. We are still checking whether anything else contributed. |
| We know exactly what happened. | We have established the sequence for the checkout path. The payments path is still under review. |
| I have no idea what happened. | Interface errors on the network path, and connection counts and slow-query logs on the database, stayed at baseline for the window. That weakens a failing link on that path, and connection or slow-query saturation on the database, as explanations. The cache layer is next, and that check takes about twenty minutes. |
| It's the network. | Packet loss between the web tier and the order database rose to 3 percent at 14:04, per the path monitor. We are checking whether that loss accounts for the timeouts or runs alongside them. |
| It's always DNS. | Lookups for the payment endpoint failed on two of our six resolvers between 14:02 and 14:09, per the resolver logs. The other four answered normally, and we are checking which clients used the failing pair. |
In these rows, what was observed and what it might mean get separate sentences. A change that happened before the failure is a timing fact. Whether it produced the failure is a separate claim with its own evidence.
Impact and scope
| Instead of | Try |
|---|---|
| Ten thousand customers are impacted. | Approximately ten thousand failed requests were recorded during the window. We are validating the number of distinct affected customers. |
| Everything is down. | The confirmed affected capability is checkout. Other journeys are under assessment. |
| No customers were affected. | Between 14:00 and 14:40, telemetry shows no degradation on the customer-facing endpoints we monitor. Paths outside our monitoring remain unassessed. |
| CPU is pegged, so users are impacted. | The order service is saturated and response times have doubled. We are measuring what that means for checkout completion. |
| We lost revenue. | Finance is assessing the financial effect using order data from the window. We will pass their figure along when they have one. |
| It's intermittent. | Between 14:02 and 14:40 the service alternated between normal and failing roughly every four minutes. We are comparing that pattern with the cache refresh schedule. |
| Only a few users are complaining. | Support has 14 tickets since 14:10. Tickets undercount impact, so we are measuring failed logins across all users for the same window. |
| It only affects one region. | Errors are confirmed in the primary region. The secondary region shows normal error rates, and we are checking whether any primary traffic fails over into it. |
A host metric is evidence about a machine. Impact is a statement about a business capability, and getting from one to the other usually needs evidence about the capability or its users. The revenue row applies the same split: Finance owns that figure, and the engineering update says who is producing it.
Data and integrity
| Instead of | Try |
|---|---|
| No data was lost. | Reconciliation is incomplete. As of 15:30, no evidence of data loss has been identified in the order and payment tables reviewed so far, and the remaining tables are next. |
| Nothing was exposed. | We have not identified indicators of unauthorized access in the logs reviewed. The review is ongoing and its scope is the web and API tiers. |
| The data might be compromised. | We observed unexpected reads against the customer table between 02:10 and 02:40. The security and data owners are assessing what that means, and the next steps are theirs to decide. |
| The backups are fine. | The last backup job reported success at 02:00. A test restore of that set has started and will finish by 16:00. |
| The data is corrupted. | Checksum validation fails on 38 order records written between 14:02 and 14:15. Writes to that table are paused, and the data owner is assessing what it means for those orders. |
| We restored everything from backup. | We restored the order table from the 02:00 backup at 15:10. Orders written between 02:00 and 14:02 are being replayed from the transaction log, and reconciliation will show whether any are missing. |
"No data was lost" is an unbounded claim about a negative, made early, by someone who has not completed reconciliation. "No evidence of loss in what we have reviewed" carries its own time and scope, so it remains true of what was reviewed at that time, even if something turns up later. The backup rows have the same problem in a smaller form. A job that reported success is a claim by the job. A restore that completed and reconciled is evidence.
Action and recovery
| Instead of | Try |
|---|---|
| I just rebooted the box. | We restarted the order service at 14:20, after confirming no deploy was in flight. Error rates returned to baseline; what saturated it is still open. |
| It's fixed. | The service has returned to its documented baseline. We are monitoring for recurrence. |
| We're back up. | Error rates returned to baseline at 14:19. Latency is still elevated while queued work drains, and we expect that to finish around 15:25. |
| We failed over and everything is good. | Traffic moved to the secondary site at 14:25 and it is serving checkout. We are checking replication lag and capacity before calling it stable. |
| We rolled back the bad code. | We restored release 4.1 at 14:15 and kept 4.2 aside for review. Error rates returned to baseline by 14:19, and the backlog is draining. |
| The alert cleared. | The alert condition is gone. We are watching through a thirty-minute stability window and confirming with the order team that their queue is moving. |
| We patched it. | We applied the vendor's fix to one of four nodes at 15:30, and the error rate on that node dropped to baseline. The other three follow after an hour of observation. |
| We scaled it up and it's fine. | We added four instances at 14:40, and response times returned to baseline by 14:48. Whether demand or a slowdown inside the service drove the saturation is still open, so the added capacity stays until we know. |
Each row keeps to what was done or observed, how far it reached, and what came of it, and leaves the word for "over" until verification earns it. An alert clearing tells you the threshold stopped firing. Whether the business process recovered is a second question with a different owner.
Time and commitment
| Instead of | Try |
|---|---|
| We'll have it fixed shortly. | The next update will be at 15:00, with or without a change in state. |
| Still investigating. | Checks on the network path and the database found nothing in the signals reviewed that explains the failures. We are testing the cache layer. Next update at 15:00. |
| It'll be back in twenty minutes. | The next update is at 14:30. We can give a recovery estimate once the restore test finishes. |
| No ETA. | We can't estimate recovery yet because the restore rate is still unknown. We will have that number by 15:00. |
| Nothing new to report. | State is unchanged. Since the last update we checked the load balancer: its connection counts and backend health checks were normal for the whole window, so the reviewed signals show no failure at the load balancer itself. Next update at 15:30. |
| Engineering is working on it. | The storage team is rebuilding the affected volume. They need a decision on pausing the nightly batch by 16:00. |
| We're close. | Two of three restore jobs have finished. The third is 60 percent through after 50 minutes at a steady rate, which puts it at about 16:10. Next update at 15:45 either way. |
| Give us an hour. | We will know by 15:30 whether the rollback clears the errors. If it does, recovery follows within minutes. If it falls short, the next option takes about two hours. Next update at 15:30. |
The commitment in each row is a time you control. A recovery prediction depends on things you have yet to learn, so when one is given, it comes with what it depends on.
Vendors and dependencies
| Instead of | Try |
|---|---|
| The vendor caused our outage. | Calls to the payment provider have been timing out since 14:02, and both sides of that path are under investigation. Case 88213 is open with them. |
| It's the cloud provider's problem. | The provider's status page has reported degraded storage in our region since 13:55, and our errors began at 14:02. We are confirming the link and preparing our failover option. |
| Their patch took us down. | The outage began after the vendor's firmware update. We are working with their support team to establish whether the update is the mechanism. |
| We're waiting on the vendor. | Vendor case 88213 is with their engineering tier. Our continuity option is manual order entry, and the order team is ready to switch if the case runs past 16:00. |
| The vendor says it's fixed on their end. | The provider reports their fix went in at 15:05. Our calls to them have succeeded since 15:07, and we are watching our own error rate through a thirty-minute window before calling our side recovered. |
| It's a known bug in their product. | The vendor's support case links our symptoms to a published defect in the firmware version we run. We are confirming the match with their engineer before planning the upgrade. |
What a supplier did is for the people who manage that relationship to characterize. The engineering record supplies timestamps, observed behavior, and the case number, and it keeps looking at your own side of the connection while the supplier looks at theirs.
Time windows and intermittent behavior
| Instead of | Try |
|---|---|
| The service was up all night. | Health checks every 60 seconds passed from 22:00 to 06:00. Failures shorter than a minute could fall between checks, and the request logs for that window are being pulled. |
| Availability was 99.9 percent today. | Availability for the day was 99.9 percent, measured as successful requests over total requests. The worst five-minute window, 14:05 to 14:10, failed 12 percent of requests. |
| The latency trend is flat, so storage isn't the problem. | Five-minute latency samples show no rise over two weeks. The nightly backup runs between samples, so the chart would miss a short spike during it. A capture at a one-second interval across tonight's backup will show what happens at that resolution during this run. |
| It happened a few times this afternoon. | The load balancer logs show three error bursts, at 13:12, 14:03, and 14:51, each under a minute. Our metrics resolve to one minute, so shorter bursts could exist that we cannot see. |
| Memory was fine during the outage. | Memory use was 61 percent at 15:10, thirty minutes after recovery. We have no memory sample from inside the outage window. |
| The outage started at 14:02. | The first failed request in the application log is stamped 14:02:13 UTC. The storage array logged a latency warning at 14:02:40 by its own clock, which runs about 40 seconds ahead of the application servers. Corrected for that, the array warning came first. |
| It's been fine since the fix. | Error rates have been at baseline for 25 minutes since the change at 14:40. The failure previously recurred about every 40 minutes, so we are watching through at least 16:00 before drawing a conclusion. |
| The dashboard shows everything green. | The dashboard averages over five minutes and last refreshed at 14:35. The one-minute view shows two short error spikes, at 14:31 and 14:33, that the average hides. |
Each row names the reading, its interval, and the stretch of time it leaves uncovered. The clock row matters most once someone assembles a timeline, because the order of events often carries the whole explanation.
Security and access
| Instead of | Try |
|---|---|
| We got hacked. | We observed logins to the admin console from an unrecognized address between 01:10 and 01:25. The security team is leading the assessment, and the account has been disabled while they work. |
| It was just a phishing email. | One user reported entering credentials on a lookalike page at 09:40. The password was reset and the account's sessions were ended at 09:55. Security is reviewing the account's activity since 09:40. |
| Nobody else could have accessed it. | The access logs for that share show reads from two service accounts and one user during the window. Security is reviewing whether any of them were unexpected. |
| The firewall blocked it, so we're fine. | The firewall blocked the connection attempts it logged. Security is checking the hosts that made them and whether any traffic took another path. |
Security rows follow the data rules above: state what was observed and what was done, and leave the conclusion to the team that owns the assessment. A sentence about intent or exposure belongs to them.
Changes and maintenance
| Instead of | Try |
|---|---|
| The maintenance went fine. | The upgrade finished at 02:40, forty minutes past the window. Post-change checks passed on all six nodes. Two scheduled jobs that ran during the overrun are being checked by their owners. |
| It's a low-risk change. | The change touches one configuration value on the load balancer, has run in staging for a week, and can be reverted in under five minutes. It lands in the 22:00 window, with the network lead approving. |
| We'll roll back if anything goes wrong. | The rollback trigger is an error rate above 2 percent for five minutes after the change. Rollback takes about ten minutes and has been tested in staging. |
| Nothing changed. | The change calendar shows no approved changes in the 24 hours before the incident. We are checking deploy logs and configuration history, because unrecorded changes would show up there. |
These rows cover the conversations around a change as well as the incident itself. A change described only by its risk level asks the reader to rely on a judgment without seeing its basis. A change described by what it touches, how it was tested, and how it comes back out lets the reader make the judgment.
Workarounds and continuity
| Instead of | Try |
|---|---|
| Just use the backup process. | Manual order entry can take about 40 percent of normal volume with four extra staff, and those orders need reconciliation tomorrow. The order team decides whether to start it. |
| Users can work around it. | Users can submit expense reports through the mobile app while the web form is down. The mobile path lacks receipt upload, so reports that need receipts wait for the fix. |
| Failover will handle it. | Failover to the secondary site was last tested in March and took 25 minutes. Running there limits capacity to about 70 percent of normal, so we recommend holding it unless the primary stays down past 16:00. With the 25-minute failover and about 90 minutes for the batch at reduced capacity, 16:00 is the last start that still finishes before 18:00. |
A workaround has a capacity, a cost, and an owner who decides whether to use it. Stating all three turns "just do this" into a choice the business can make.
Up the chain
| Instead of | Try |
|---|---|
| Write latency on the array is at 40 milliseconds. | Order processing is running at about 2,500 orders an hour, a quarter of the normal 10,000, per the order service's own timing over the last fifteen minutes. As of 14:35, 13,500 orders remain, so at that rate the batch is projected to finish near 20:00, two hours past the 18:00 cutoff. Engineering recommends holding noncritical jobs to recover speed, and we need an answer from the order operations lead by 16:00. |
| Here are the logs from the last hour. | Checkout failed for about 17 minutes, from 14:02 to 14:19, and is working again. We are verifying stability through 16:00. The logs are attached for anyone who wants the detail. |
| We need more disks. | Used capacity grew about 3 TB a week last quarter, and 18 TB remain, so at that rate the array fills in about six weeks. There are two options: add a shelf this month, or move the archive tier to other storage, which takes longer and carries migration risk. We need a decision by the end of the month. |
| IT is investigating. | The storage team owns the technical recovery. The order team owns the workaround and the deadline. Both get an update at 15:00. |
| We can't do anything until the vendor responds. | The provider owns its side of the investigation, and we own continuity and our own path. In the meantime we can move the 18:00 batch to manual processing, which handles about 40 percent of normal volume with four extra staff and needs reconciliation the next day. The remaining orders wait for service restoration, and we will tell the affected teams which orders fall in that group. We need your call on that by 16:00. |
| There's nothing to worry about. | The service is stable and we are watching it through the evening. One risk remains open: replication to the secondary site is running on one of its two network paths, so a second failure would stop it. The network team restores the second path tonight at 22:00. |
| The storage array is degraded. | Storage for the finance systems is running on reduced redundancy. Nothing is failing now. Under the array's current protection level, a second disk failure in the same group would take month-end close offline. The replacement arrives tomorrow morning, and we recommend deferring tonight's large import until it is in. |
| We had a P1 last night. | Checkout was unavailable for 22 minutes last night, from 23:10 to 23:32, and the logs show about 1,400 failed orders. The order team is contacting affected customers through the normal process, and the review is scheduled for Thursday. |
The sentences on the left are common shorthand, and each one hands the listener a technical fact to convert into a business decision on their own. Each version on the right does that conversion and keeps the evidence available in the incident record.
What a restart proves
A restart resets state. It is often the right action. By itself it establishes that recovery followed the reset, and leaves open why the reset worked. "We rebooted and it is working" tells a stakeholder the problem is over, when what happened is that the symptom cleared and the condition behind it is still uncharacterized. If the restart cleared a saturated pool, the pool may saturate again if the same load returns. The better version separates three claims: the action taken, the observed result, and the open question. "After the restart, the symptom cleared and the service is serving normally. We have not yet determined what produced the saturation, so recurrence is possible. Next update at 16:00."
Sometimes the restart is the plan. Some teams run a service that gets recycled every few weeks because it degrades in a way nobody has had time to chase. That can be a legitimate stopgap, provided it is described as what it is: a scheduled action that holds down a known risk while remediation, the change that corrects the underlying condition, waits. It earns that description when it is scoped to one service, runs on a schedule someone approved, is watched for whether it still works, and has an owner and a review date for the underlying condition. Written up that way, it reads as a managed risk. Written up as "the cron job handles it," it reads as a repair, and the underlying condition drops out of view unless it keeps an owner and a review date.
Agree on the words before you need them
Most of the rows above are hard to produce at minute forty of a bad afternoon. Choosing them in advance makes them easier. Agree the status and action definitions with your team, ask your managers what they want in the first message, and find out who owns customer notification, contract commitments, and data exposure questions, so the update can name them. Then put the right-hand column of the rows your team uses most into the incident template as prompts. A template that asks for impact with one of the three labels, plus a qualifier where one applies and the inputs behind any estimate, gets better answers than one with a blank line labeled Impact.
A sixty-second test before you send
When speed is the constraint, a short check still fits. Read the draft and ask whether every claim in it is one of these: an observation with its source, an inference labeled as one, a decision with its rationale, an action with its result, or an explicit unknown.
Then check for the four claim patterns that assert more than the evidence usually supports: a named cause, fixed, no data loss, will not happen again. Each one makes a strong statement that later readers may rely on.
Then check whether an executive could read only the first two sentences and know what is affected and whether a decision is needed from them.
A claim that fails the first two checks gets narrowed to what you can support, and an update that fails the third gets reordered, before it goes out.