The AT&T outage of 2024
A configuration error during a routine network expansion disconnected around 125 million devices for most of a day, and blocked thousands of calls to emergency services.
A configuration error during a routine network expansion disconnected around 125 million devices for most of a day, and blocked thousands of calls to emergency services.
On 22 February 2024, starting at about 02:45 Central time, AT&T's mobile network failed across the United States. Phones showed "SOS only". Calls would not connect. Data would not work. It lasted around twelve hours for most affected customers.
The FCC's subsequent investigation put the scale at approximately 125 million devices, more than 92 million blocked voice calls, and over 92,000 failed calls to 911.
AT&T was carrying out a routine network expansion - Adding a new network element. The configuration applied to that element was incorrect, and it was not caught before deployment.
The misconfiguration caused devices to disconnect and immediately attempt to reconnect. That is the cascade: Millions of handsets simultaneously trying to register produced a registration storm that the network could not absorb, which caused more disconnections, which produced more registration attempts.
AT&T removed the faulty element within about three hours. Restoring service took considerably longer, because the network had to absorb the reconnection of a hundred million devices without collapsing again. Recovery was staged deliberately for that reason.
Three things separate this from a website going down, and they compound.
You cannot report it. Every other outage leaves you a channel to find out what is happening. A mobile network outage removes the device you would use. People drove to places with Wi-Fi to work out what was going on.
Emergency calling is affected. The 92,000 blocked 911 calls are the reason this was treated as a regulatory matter rather than a service failure. Mobile phones are the primary way most people would call for help, and a network outage removes that for everyone in its footprint simultaneously. Devices are meant to fall back to any available network for emergency calls, and the FCC found that fallback did not work as intended.
The dependencies are physical. Payment terminals on cellular connections, alarm systems, fleet tracking, medical alert devices. A mobile network is infrastructure that other infrastructure runs on.
The FCC's report identified a chain of process failures rather than a single technical one.
The change was not adequately peer reviewed. The configuration went out without the scrutiny AT&T's own procedures required.
There was no adequate rollback plan. Recovery meant identifying and removing the faulty element under pressure rather than reverting a known-good state.
Testing did not model the scale. The configuration may have behaved acceptably in a test environment; the cascade only appears at production scale.
Controls were not applied consistently. The safeguards existed on paper.
The FCC subsequently reached a settlement with AT&T including a $950,000 payment and a commitment to specific compliance measures.
This was not unique. T-Mobile's June 2020 outage lasted around twelve hours and affected voice and messaging nationwide, beginning with a single fibre circuit failure that cascaded through a routing misconfiguration - The FCC's report described it as preventable.
Rogers in Canada, in July 2022, was worse in some respects: A configuration change during a network upgrade took down both mobile and fixed-line service for a large share of the country for about nineteen hours, affecting payment terminals, emergency calling and government services. Rogers customers could not even switch to another network, because their phones had no service at all.
The common shape: A routine change, inadequate review, a cascade at production scale, and a recovery slowed by the need to restore service without triggering the cascade again.
Less than you would like, but not nothing.
Wi-Fi calling. If the fixed-line internet is working, Wi-Fi calling routes voice over it and bypasses the mobile network entirely. It needs to be enabled in advance - During an outage you may not be able to download anything or change settings that require a network.
Know that 911 should fall back. Handsets are designed to place emergency calls over any available network regardless of carrier. It does not always work, as this incident showed, but it is worth attempting.
Have a second path. For anyone whose work depends on connectivity, a second carrier - Even a cheap prepaid eSIM - Is meaningful redundancy, because carrier outages are single-carrier events.
Note that a landline is not obsolete. A traditional copper line works during a power cut and during a mobile outage. It is one of the few genuinely independent paths left.
The lesson generalises. This was not a hardware failure, an attack or a natural disaster. It was a change, applied correctly, that had not been reviewed properly, in a system whose behaviour at full scale nobody had modelled.
That is the normal shape of a large outage in every industry. What made this one exceptional was not the cause but the consequence - That for twelve hours, a hundred million people could not call for help.
A configuration error applied during a routine network expansion. It caused devices to disconnect and immediately try to reconnect, producing a registration storm the network could not absorb. The FCC found the change had not been adequately peer reviewed and that there was no adequate rollback plan.
The FCC put it at approximately 125 million devices, with more than 92 million blocked voice calls and over 92,000 failed calls to 911 during roughly twelve hours of disruption.
Phones are designed to place emergency calls over any available network regardless of which carrier you subscribe to. That fallback did not work reliably during this incident, which is why more than 92,000 emergency calls failed, but it is still worth attempting.
Enable Wi-Fi calling in advance - It routes voice over your home internet and bypasses the mobile network entirely. A second carrier, even a cheap prepaid eSIM, is real redundancy, because carrier outages affect one network at a time.
The real causes of website outages - Configuration pushes, BGP withdrawals, expired certificates and dependency cascades.
How to confirm an ISP outage, why cable outages affect one street and not the next, and what to do while you wait.
Typical outage durations by cause, why most incidents resolve in minutes, and the warning signs that an outage is going to be a long one.
What happened during the October 2021 Meta outage, why a routing change removed the company from the internet, and why it took six hours to fix.