Google AI
The Times Australia

Times Media Advertising

Optus has revealed the cause of the major outage. Could it happen again?

  • Written by: Mark A Gregory, Associate Professor, School of Engineering, RMIT University
Optus has revealed the cause of the major outage. Could it happen again?

Around 4.05am on Wednesday November 8 2023, Optus suffered a nationwide network outage lasting well into the evening, more than 12 hours later.

Now, Optus has released some information on what happened, stating[1] “we now know what the cause was and have taken steps to ensure it will not happen again”.

As a telecommunications expert, I believe we should have no confidence in this statement, because the poorly worded explanation leaves many questions unanswered.

Could a similar outage happen again? We don’t know – but there are ways to make it less likely.

Read more: Optus blackout explained: what is a ‘deep network’ outage and what may have caused it?[2]

How did the outage unfold?

The Optus outage caused all services to go offline. Landlines, mobile phones, home internet, small business and enterprise, and cloud connections all dropped out.

The most serious impact of the outage was that Optus landlines couldn’t dial 000[3] and Optus mobile phones were unable to connect to the 000 emergency call service unless the connection occurred through Telstra or Vodafone infrastructure.

More than 10 million Optus customers were affected by the outage that brought Melbourne’s trains to a halt[4] and left Optus’s small business customers unable to carry out EFTPOS transactions[5].

So, what went wrong with Optus?

Optus has revealed that a “routine software upgrade” triggered a cascading failure in the Optus internet protocol[6] (IP) core network – the central backbone of their network that authorises device access and provides customer management.

Optus has provided a brief answer on why the entire network went offline[7]:

“These routing information changes propagated through multiple layers in our network and exceeded preset safety levels on key routers which could not handle these. This resulted in those routers disconnecting from the Optus IP Core network to protect themselves.”

Routing information is used to find a path from one location on the internet to another – a router is a device that manages the traffic flows.

The explanation provided by Optus points to human error. This confirms what industry experts suspected had happened[8]. The resulting flood of “routing information changes” overwhelmed[9] key routers in the core network causing them to disconnect, thereby bringing the entire network to a halt.

Read more: Explainer: what is the 'core network' that was crucial to the Optus outage?[10]

Photo of a white wifi router on a desk with a person working on laptop in the backround
Your internet router is a home version of a device that manages data traffic flow. Teerasan Phutthigorn/Shutterstock[11]

Should the outage have been preventable?

Outages of this kind are not uncommon – human error has led to major companies going offline in the past.

But an entire telecommunications network going offline is unusual. The network should be designed in such a way that redundancy (backups) and resiliency are built in from the outset.

Before a software upgrade occurs, there should be modelling, testing and several layers of sign-off.

In case something goes wrong, there should be infrastructure and system redundancy. An automated or manual procedure should exist to ensure the redundant systems become operational within a few minutes.

In 2021, Facebook, WhatsApp and Instagram[12] disappeared from the internet for roughly six hours due to an incorrect routing configuration.

Meta’s lengthy and informative statement[13] at the time provides an example of the level of detail that we should expect Optus to provide.

With the Optus outage and similar incidents at other companies that have led to major outages, in nearly every case the outage was preventable and highlighted deficiencies in the organisation.

Read more: In a crisis, Optus appears to be ignoring Communications 101[14]

What should Optus do now?

The national outage means the Optus network is not fit for purpose[15].

It can be assumed Optus has a number of deficiencies, such as problems with engineering capability, testing, procedures, network redundancy and resilience.

Optus states they are “committed to learning from what has occurred” and will continue to work to “increase the resilience” of their network.

For this to lead to an effective outcome, Optus will need to carry out a review and put in place new processes, infrastructure and systems to prevent a similar outage in the future.

How do we know a similar outage won’t happen again?

We don’t.

We need enhanced government regulation of the Australian telecommunications network operators to provide improved visibility of the redundancy and resilience of their networks. The Senate has commenced an inquiry[16] into the Optus outage.

Telecommunications is an essential service. Australians should be able to connect to the 000 emergency call service at all times. Reliable access to medical services, EFTPOS and the internet are vital.

If necessary, penalties should be introduced into the Telecommunications Act 1997[17] to ensure telecommunications network operators implement and maintain “best practice” related to network operation, redundancy and resilience.

References

  1. ^ stating (www.optus.com.au)
  2. ^ Optus blackout explained: what is a ‘deep network’ outage and what may have caused it? (theconversation.com)
  3. ^ couldn’t dial 000 (www.sbs.com.au)
  4. ^ trains to a halt (www.9news.com.au)
  5. ^ EFTPOS transactions (www.abc.net.au)
  6. ^ internet protocol (www.cloudflare.com)
  7. ^ why the entire network went offline (www.sbs.com.au)
  8. ^ suspected had happened (www.scimex.org)
  9. ^ overwhelmed (www.theguardian.com)
  10. ^ Explainer: what is the 'core network' that was crucial to the Optus outage? (theconversation.com)
  11. ^ Teerasan Phutthigorn/Shutterstock (www.shutterstock.com)
  12. ^ Facebook, WhatsApp and Instagram (blog.cloudflare.com)
  13. ^ Meta’s lengthy and informative statement (www.facebook.com)
  14. ^ In a crisis, Optus appears to be ignoring Communications 101 (theconversation.com)
  15. ^ not fit for purpose (www.abc.net.au)
  16. ^ commenced an inquiry (www.aph.gov.au)
  17. ^ Telecommunications Act 1997 (legislation.gov.au)

Read more https://theconversation.com/optus-has-revealed-the-cause-of-the-major-outage-could-it-happen-again-217564

Times Magazine

Offshore vs Inshore Centre Console Boats: Which One Should You Buy?

Centre console boats have become one of the most popular choices among modern anglers. Their open ...

Why Australian Enterprises Are Rethinking Their Core Communication Technologies

The corporate landscape in Australia has undergone a permanent structural shift over the past few ...

Road safety risk: New data reveals almost 2 in 3 Australian drivers are letting car maintenance slide as cost of living pressures bite

Australians are putting off vehicle maintenance and new research released on the eve of National R...

Woodroffe footy club BBQ legend crowned in national Bunnings search

Bunnings has found its latest community hero, naming Brent Tanner from Darwin Buffaloes Football C...

VoltX Energy expands into Victoria & ACT to meet surging home battery demand

Leading Australian energy solutions provider VoltX Energy and premier sponsor of the NRL Manly Wa...

Victorian Drivers To Receive 20% Rego Rebate From June 1 In Major Cost-Of-Living Measure

Victorian motorists will begin receiving significant registration savings from June 1 as the Allan...

How Australian Businesses Are Using AI To Cut Costs And Improve Efficiency

Artificial intelligence was once viewed by many small business owners as something futuristic, exp...

Quickest Way of Getting Rid of Your Old Cars in Brisbane?

If you are done searching for a practical solution for quickly getting rid of your old car, this w...

The Human Supplement Craze Has Officially Gone to the Dogs (Literally)

Australians’ appetite for supplements is no longer limited to their own vitamin cabinets. New reta...

The Times Features

Pauline Hanson at the National Press Club: A Defining P…

For almost 30 years, Senator Pauline Hanson has been one of the most recognisable and controversia...

Covid: The pandemic has ended but the health story hasn…

Covid is no longer the daily emergency it was in 2020 and 2021. The fear, lockdowns, border closur...

Macca’s introduces new McSmart range with more choice f…

Macca’s is launching its new-look McSmart range from Wednesday,1 July, with  three new meals at thre...

Why Australia Was Hoping For Another Interest Rate Cut

When the Reserve Bank considers interest rates, the focus is often on inflation, employment and ec...

$100,000 A Year: Where Does That Put You In Australia?

For many Australians, earning $100,000 a year remains an important financial milestone. It is a s...

The Kennedy Center and the Trump Name: A Battle Over Hi…

The removal of Donald Trump's name from part of Washington's famed Kennedy Center has become far m...

The Times Guide to Sydney's Beaches

Winter may still have a grip on Sydney, but anyone who has lived in Australia's largest city knows...

How Australia's Childcare Crisis Is Taking a Toll …

Australian mums and dads are increasingly anxious, exhausted, and distrustful of Australia’s childca...

The Economics of a Cup of Coffee: Is Your Daily Cappucc…

For many Australians, a morning coffee is no longer a luxury. It is a ritual. A quick stop at the ...