The recent Telstra outage that disrupted mobile services across Australia was caused by a neglected software update on a key time-keeping system, CEO Vicki Brady told a Senate inquiry. The failure highlights critical vulnerabilities in network infrastructure and the importance of timely software maintenance.
What Caused the Telstra Outage?
The outage originated from one of Telstra's three Network Time Protocol (NTP) servers, a Microchip SSU 2000 model manufactured in 2011. During routine maintenance to replace faulty backup power in Melbourne, the server was shut down and restarted. Due to an underlying software configuration issue, it reset to the year 2006 instead of the correct date.
Get the #1 Wireless Door Camera
REOLINK Bestseller: 2K Weatherproof Video Doorbell, No Monthly Fees.
This incorrect date then spread across the network over several hours, causing authentication certificates in other servers to become invalid. As a result, 45% of all calls and data sessions were affected, leaving customers unable to authenticate onto the network.
Redundancy Wasn't Enough
Telstra insisted that the network had redundancy in place, with three NTP servers located in Sydney, Melbourne, and Perth. However, the design change meant that the servers were not communicating properly, and the redundancy did not prevent the outage. The company admitted that maintenance teams were unaware of the design change that affected how the server would reset.
Key Details from the Senate Inquiry
During the hearing, executives confirmed that the manufacturer had alerted Telstra in both 2022 and January 2024 about the need to update the software. Had the update been applied, the outage may have been avoided. The server, which costs $30,000 to replace, was still under support from Scientific Devices.
| Factor | Impact |
|---|---|
| Software update neglected | Server reset to 2006 date |
| Design change unknown | Redundancy failed to prevent outage |
| 45% of calls/data affected | Nationwide chaos for hours |
Lessons for Network Reliability
This incident underscores the need for rigorous software patch management and clear communication between maintenance teams and network engineers. Companies must ensure that all critical systems receive timely updates, especially when manufacturers flag potential issues.
- Software updates must be applied promptly to avoid cascading failures.
- Network redundancy requires proper configuration and awareness of design changes.
- Maintenance procedures should include checks for software version compliance.
FAQ
What caused the Telstra network outage?
The outage was caused by a neglected software update on a Network Time Protocol (NTP) server, which reset to the year 2006 during maintenance, leading to widespread authentication failures.
How many customers were affected by the Telstra outage?
Approximately 45% of all calls and data sessions were affected, causing nationwide disruption for Telstra mobile customers.
Could the Telstra outage have been prevented?
Yes, Telstra was alerted by the manufacturer in 2022 and January 2024 to update the software. Applying the update would likely have prevented the outage.
This incident serves as a stark reminder for all telecommunications companies to prioritize software updates and ensure that network redundancy plans are fully understood by all teams. For more insights on network reliability and business continuity, stay tuned to our business section.