Information Design for Crisis Response: Lessons from AI, Command Structures, and a Single Source of Truth
Does AI Make Crisis Response Faster? The Promise and Pitfalls of Automated Triage
In the first moments of a system outage, responders must quickly work out what matters among a flood of logs, alerts, metrics, and traces. The spread of SRE (Site Reliability Engineering) has brought major advances in monitoring and incident response, and in recent years AI has begun to enter this space as well.
In 2026, Google disclosed that it uses AI agents in its SRE work for incident investigation and some mitigation actions. Microsoft has also rolled out AI-based observability features that work across logs, metrics, and traces to surface likely root causes.
Where AI shines most is initial triage. It can group related alerts out of thousands, look for similarities with past incidents, and propose candidate causes. Shortening work that once took several engineers considerable time matters a great deal, especially in a crisis that cannot wait.
But there is a trap here. The faster analysis becomes, the greater the risk of treating its output as fact.
What AI presents is, in most cases, a plausible hypothesis derived from observed data; it does not establish the cause of the outage itself. If the logs it relies on are incomplete, its conclusions will be skewed, and if dependency information is outdated, it may point to the wrong component.
However capable AI becomes, as in any cyber crisis, “what is currently most likely” must be kept distinct from “what has been confirmed.”
The problem goes a level deeper when automated recovery is also handed to AI. If the root cause estimate is wrong, automated actions can widen the outage. Google itself, while describing AI as a “force multiplier” for SRE, has made clear that it intends to keep humans in control.
The point of crisis response in the AI era, then, is not to remove humans from decision-making. It is to let AI handle the high-speed work of searching, summarizing, and proposing candidates, while humans take on the role of verifying, choosing, and taking responsibility.
AI will make crisis response faster. But at what stage should information drawn from that faster analysis be treated as fact? Managing that boundary is the new challenge in crisis recovery.
Leaders Should Not Gather Too Much: A Command Structure That Delivers Both Speed and Accuracy
When a crisis hits, leaders sometimes try to gather as much information as possible, directly. They call the on-site manager, ask engineers for explanations, and take reports from every department. This can look like proactive crisis management, but in a large-scale incident it can actually throw the flow of information into disarray.
Resolving the “separation of information and authority” we discussed in a previous post is important. But the solution is not for leaders to tap every information source themselves. What matters is clearly defining the paths that connect information to authority.
A useful reference here is the Incident Command System (ICS), a framework long used in disaster response. ICS emphasizes a clear chain of command and avoids situations where one responder receives conflicting instructions from multiple supervisors. It is also designed so the organization can scale with the size and complexity of the incident.
Incident management in SRE, a discipline that originated at Google, adopts this approach as well. An Incident Commander oversees the overall response, an Operations Lead handles technical recovery, and a Communications Lead keeps stakeholders informed. Crucially, crisis-time command roles do not have to match everyday job titles, and these roles are not standing positions in the normal organization.
None of this means the leader’s role shrinks.
What leaders should be doing is not checking the logs for themselves, but making executive decisions: how far the impact currently extends, what to prioritize protecting, and when to suspend services or make a public announcement. To do that, they need to obtain the necessary information quickly through defined channels.
When leaders start contacting the front lines individually, responders are asked to handle recovery work and executive briefings at the same time. And if slightly different explanations arrive from different departments, an extra round of checking “which version is correct” begins. More information does not necessarily mean better decisions. Fast, accurate information gathering is not about collecting large volumes of data. It is about having the information needed for decisions converge, at the right level of detail, into a single chain of command.
What leaders need in a crisis is not to know everything. It comes down to clarity on who knows what, and whose information serves as the basis for decisions.
Information Design for the First 24 Hours: How to Interpret a Single Source of Truth
When a major outage occurs, multiple “facts” tend to emerge within an organization.
The engineering team’s chat holds the latest findings, the executive meeting deck reflects the situation as of an hour ago, and the PR team has an already approved external statement ready. Customer support may be working from yet another document. Each was shared with great care to be accurate at the time it was written.
Over time, however, their contents drift apart, and the organization ends up with different explanations existing side by side. This is where the idea of a Single Source of Truth (SSOT), consolidating everyone’s point of reference into one, becomes important.
In Google’s SRE approach to incident management, one of the Incident Commander’s key responsibilities is maintaining a “live incident state document.” This document records the status of the incident, significant changes, and response actions so that multiple stakeholders can refer to the same information.
That said, treating an SSOT in crisis response as “the one correct truth” creates problems, because during a crisis the facts themselves are updated in real time. At 10 a.m., the impact may be believed to be limited to domestic operations; by 11 a.m., it may turn out that overseas operations were affected too. In that case, what the SSOT needs is not to delete or hide the earlier entry, but to keep “what we understood as of 10:00” and “the new fact discovered at 11:00” clearly distinct.
In other words, an SSOT for the first 24 hours should record at least the following separately: (1) confirmed facts, (2) current hypotheses, (3) unconfirmed items, (4) impact on customers and the business, (5) decisions made, and (6) owners and the next check-in time.
It is especially important not to write facts and hypotheses in the same field. Instead of writing “caused by a database failure,” simply recording “investigating a possible link to database latency” makes it easier to keep executives and PR from drawing mistaken conclusions when the information reaches them.
Creating an SSOT is also not enough on its own. Unless you decide who can update it, who reviews it, and how often it is refreshed, the document quickly goes stale and becomes a source of confusion itself.
An SSOT in crisis response is not a mechanism for settling on a single truth. It is a mechanism for continuously synchronizing what the organization currently holds as its shared understanding, even as that understanding keeps changing.
What is truly needed in the first 24 hours is not more information. It is narrowing the gaps between the pieces of information that exist across the organization, so that engineers, executives, PR, and customer support can all make decisions from the same starting point.