Why a Routine Maintenance Task Caused a Nationwide Outage (Part 2): Change Enablement in Practice, and Communication During an Incident
Recap of Part 1
In Part 1, we laid out the nationwide outage at Telstra: its timeline and its true root cause.
The reason an outage occurred even though the engineer worked exactly as the procedure said is that the procedure did not correctly reflect a design change that had been made to the target equipment in the past. An undocumented change, an update that was never applied, a known risk that went untracked. We read, through the lens of ITIL Change Enablement, how these overlapping management gaps cascaded into a large-scale outage.
In Part 2, we turn to practical measures for preventing this kind of outage on the ground, and to how communication should work once a failure has happened.
Chapter 4: Practical Measures for Carrying Out Maintenance Safely
So what should be done, in day-to-day operations and maintenance, to prevent an incident like this one?
Of course, a large carrier like Telstra and the internal systems of an ordinary company differ greatly in the scale and configuration of their systems.
Even so, problems such as “no one knows what changed in the past,” “a patch is still outstanding,” and “the procedure does not match reality” are common to many sites.
As in our previous articles, we lay these out as concrete measures that are relatively easy to implement.
- Record every change, however small
Design changes to equipment, configuration changes, interim fixes, temporary changes made during incident response.
However small a change is, if the state of the production environment has changed, record what was done and maintain it as configuration information where necessary.
What is especially dangerous is an operating habit of “I fixed it on the spot, so I will write it up later.”
In high-urgency situations such as incident response, there are times when you have no choice but to prioritize recovery first. But when the changes made in those moments go unrecorded and time passes, “settings no one can explain” and “equipment that differs from the standard configuration for reasons unknown” gradually accumulate.
That is exactly the kind of state that became the problem in the Telstra case.
A change record is not merely a ledger for administrators. It is also a handover to the future you who will touch the same equipment six months or a year from now, and to whoever else works on it.
- Reflect the change in the related procedures
Recording the change is not enough on its own.
If the change alters how the equipment is operated during maintenance, or what needs to be checked, the related maintenance procedures and work checklists need to be updated as well.
If the state described in the procedure and the actual state of the equipment diverge, the worker cannot make correct judgments.
Worse, it leads to the kind of incident hardest for a worker to avoid: “I followed the procedure exactly, and an outage happened anyway.”
In operations, workers are often told to “follow the procedure.”
Following the procedure is, of course, important. But just as important is continually checking whether that procedure is correct for the system as it stands today.
A procedure is not a document that is complete once written. It needs to be treated as a living document, updated every time the system changes.
- Make known risks visible and manage them by priority
For software updates and patches that should be applied, and for known defects that should be addressed, keep a list so that nothing is left to languish.
Rather than simply raising a ticket and calling it done, it is important to set priorities based on factors such as impact, likelihood, the importance of the affected system, and whether a workaround exists, and to make the deadline and the owner explicit.
The idea of the “known error” covered in our earlier problem management article applies directly to patch management and equipment maintenance as well.
Merely knowing that a problem exists does not mean the risk is being managed.
You need a mechanism that turns “we will fix it someday” into “who will fix it, by when, and in what way.”
And when a permanent fix cannot be made right away, you also need to keep interim measures on record: what work should be avoided during that period, and what monitoring should be strengthened.
- Treat maintenance work as a change, and assess it in advance
Even routine work that appears to do the same thing every time needs to be considered within the scope of Change Enablement if it affects the state of the production network or production systems.
The judgment “this is work we always do, so it is fine” is, at times, dangerous.
Systems change little by little. As configuration changes, patches, equipment swaps, and changes to connection endpoints accumulate, work that was safe six months ago is not necessarily safe now.
For that reason, before maintenance work, check at least the scope of impact, the anticipated risks, the rollback method if the work fails, the timing of execution, and the monitoring method.
Especially when working with fewer people than usual, such as late at night or on holidays, advance assessment and the accuracy of the procedure become an important safety net.
Rather than relying on the worker’s experience and intuition alone, it is important to be in a state where incidents can be prevented by mechanisms set up in advance.
- Schedule with the time of day and the blast radius in mind
This outage began late at night, and then, as morning came and the number of users and the network load grew, its impact became larger.
From this angle, the “time of day” of a change also needs thought.
It is generally believed that working in the quiet late-night hours, when there are few users, is safer. That is not wrong in itself.
However, what matters is not simply “carrying out the work when there are few users.”
If a problem does occur, how quickly can the anomaly be detected? Can the cause be identified before the morning usage peak arrives? Can the necessary staff and vendors be reached immediately? Is the person who decides on a rollback available at that hour?
The work window needs to be designed to include all of this.
It is also important to keep small the range affected by a single change, the so-called blast radius (the scope of damage).
Where possible, options include changing a subset of equipment or sites in stages, confirming healthy operation before widening the scope, and building a configuration that can automatically isolate a problem if one occurs.
The idea of keeping the blast radius small, touched on in an earlier article, is not only about access design.
It applies just as directly to how maintenance work is carried out and to how change schedules are designed.
Chapter 5: Incident Response, and the Communication That Follows
Beyond the Change Enablement angle, this case offers several suggestions about communication when an incident occurs.
In the early stage of the outage, Telstra communicated the situation through its own website.
However, the impact known at that point was still limited, and the message did not reflect the large-scale impact that became clear later.
The CEO later explained, in effect, that the early information update reflected what was known at that time, and did not reflect the larger impact that emerged afterward.
In cases where the scale of an outage grows over time, this problem is hard to avoid.
If you release information at an early stage, its content may diverge from facts discovered later. On the other hand, if you withhold information until the full and accurate picture is known, you face the question of “why did you tell us nothing.”
For that reason, when communicating during an incident, rather than trying to convey everything accurately from the outset, it becomes important to clearly separate “what we know at this point,” “what we do not yet know,” and “when we will next provide an update.”
The fact that contacting the responsible minister took a certain amount of time from the first detection of the anomaly also drew attention in the reporting.
Especially for communications infrastructure that can affect emergency calls, as in this case, you cannot think only about the internal technical recovery.
For regulators, related agencies, emergency services, users, and others, it is necessary to decide in advance “when,” “who,” “what,” and “up to what point” will be communicated.
In our earlier article on incident drills, we touched on things such as “decision-making drills,” “customer-explanation drills,” and “confirming communication routes.”
This Telstra case is precisely one where such preparation was tested for how it functions during an actual incident.
While the staff investigating the technical cause concentrate on recovery, other staff carry the communication with the outside world forward. Where necessary, the situation is escalated to management and regulators.
In a large-scale outage, this kind of division of roles also matters.
For context, behind this case lies the Optus emergency-call outage that occurred in September 2025.
Drawing on the lessons learned there, carriers had come to be held to stricter requirements, including welfare checks when an emergency call fails.
Rather than letting a past outage end as a mere “accident,” it is reflected into rules and procedures to prepare for the next one.
In that sense, this case also lets us see that not only individual companies but the telecommunications industry as a whole is in the middle of ongoing, continuous improvement.
In Closing: Even When You Follow the Procedure, a Wrong Procedure Still Causes an Outage
What the Telstra outage brings home is a slightly hard truth for anyone in operations and maintenance.
The engineer followed the procedure. The outage happened anyway.
Because the procedure itself did not correctly reflect a change that had been made to the target equipment in the past.
For anyone in operations and maintenance, this is by no means someone else’s problem.
An undocumented design change. A patch still sitting unapplied. A procedure that has not been updated. A configuration change left in place after being added temporarily during incident response.
Taken one by one, each is a mundane problem.
And when you are caught up in daily inquiries, incident response, and routine work, these are exactly the tasks you are tempted to put off, on the grounds that “they are not causing trouble right now.”
But in this case, at the far end of such small management gaps accumulating, a nationwide outage occurred.
What ITIL Change Enablement teaches is that a change is not “make the change to the system and you are done.”
Record the change.
Update the current configuration.
Reflect it into the necessary procedures.
Track known risks.
Carry out the necessary permanent fix, and confirm it through to completion.
Only when you go that far is the change you made in a state where it is properly managed within the organization.
Put another way, a change must not be left “inside the equipment” alone. The fact of the change, its reasons, and its cautions need to be kept on the side of people, documents, and operational processes as well.
At its root, this is the same as what we have discussed in previous articles.
Build mechanisms that do not depend on individual attentiveness alone, so that even if someone makes a mistake, it is unlikely to lead to a major outage.
And when an outage does occur, rather than ending with “who failed,” consider “why that failure developed into an outage,” and turn it into improvements to processes and mechanisms.
Over the long run, that accumulation is what protects the system.
Is there a “change no one truly understands” sleeping inside the equipment or systems you manage right now?
Is there a patch still sitting unapplied?
Is there a procedure whose contents no longer match the current state?
Is there equipment about which people say “it has been this way for ages, but no one knows why it is set like this”?
If even one of these rings true, it is something better checked in calm, normal times than investigated in a panic after an outage.
Even short of a nationwide outage, an incident with the same structure can happen at any site, large or small.
Before it becomes the kind of outage that makes you break out in a cold sweat, why not take stock of your own environment once?
References
Telstra Exchange (2026). “Triple Zero Senate Inquiry, opening statement from Vicki Brady.” (Opening statement to the Senate committee by Telstra’s CEO) https://www.telstra.com.au/exchange/triple-zero-senate-inquiry–opening-statement-from-vicki-brady
Telstra Exchange (2026). “Our recent mobile network outage has been resolved: here’s what happened.” https://www.telstra.com.au/exchange/some-mobile-calls-and-data-services-are-affected-today–here-s-w
Australian Government, Minister for Communications (2026). “STATEMENT – Telstra Outage.” https://minister.infrastructure.gov.au/wells/media-release/statement-telstra-outage
ABC News (2026). “Time-keeping technology triggers Telstra nationwide outage. Here’s what we know.” https://www.abc.net.au/news/2026-07-08/telstra-nodes-time-keeping-technology-causes-outage/106892560
The Register (2026). “NTP server that traveled back in time caused massive Aussie mobile outage.” https://www.theregister.com/networks/2026/07/17/ntp-server-that-traveled-back-in-time-caused-massive-aussie-mobile-outage/
Information Age (ACS) (2026). “Telstra ignored bug before 8.8 million-user outage.” https://ia.acs.org.au/article/2026/telstra-ignored-bug-before-8-8-million-user-outage.html
piyolog (2026). “オーストラリアで起きた全国的な通信障害についてまとめてみた” (A summary of the nationwide telecommunications outage in Australia, in Japanese). https://piyolog.hatenadiary.jp/entry/2026/08/22/033313