Major outage history
Every large outage is a lesson in how the internet is actually wired. These are the ones worth understanding.
The internet fails in a small number of recurring shapes: A routing change that withdraws a network from the world, a configuration push that reaches everywhere in seconds, a dependency nobody drew on the diagram, and a certificate nobody renewed.
Each account below reconstructs one event - What the public saw, what actually happened underneath, how long it took, and what changed as a result.
If you are not sure where to start, the 2021 facebook outage is the one most people need first, and the us-east-1 problem follows on from it directly. If you would rather just find out whether a particular site is working, check it from our servers or look at what is failing right now.
All 7 guides in major outage history
The 2021 Facebook outage
What happened during the October 2021 Meta outage, why a routing change removed the company from the internet, and why it took six hours to fix.
The us-east-1 problem
Why AWS us-east-1 has an outsized blast radius, the major incidents it has caused, and what makes a regional failure global.
The CrowdStrike outage
What happened in the July 2024 CrowdStrike incident, why the recovery took days rather than minutes, and what it revealed about update distribution.
Cloudflare outages
How Cloudflare outages happen, why global configuration pushes are structurally risky, and what changed after each major incident.
The AT&T outage of 2024
What happened in the February 2024 AT&T outage, why carrier failures are different from website outages, and what the FCC investigation found.
The Fastly outage
How one configuration change at one Fastly customer triggered a global CDN outage, and what it showed about latent bugs in shared infrastructure.
The 2021 Roblox outage
How Roblox stayed down for 73 hours in October 2021, why subtle performance problems are harder to diagnose than crashes, and what recovery required.