Independent downtime-cost research, read by SRE and reliability teams.Sponsor this site →

Case Study

AWS us-east-1 December 2021: seven hours, and what they cost per hour

On the morning of 7 December 2021, an automated scaling activity in AWS's internal network triggered a cascading failure across us-east-1 services. Most major consumer apps and SaaS platforms hosted in the region were unavailable or severely degraded for approximately seven hours of acute impact, with downstream recovery tails extending past 30 hours for customers with implicit single-region dependencies. No aggregate cost figure for this event has ever been published: AWS disclosed none, and no analyst or insurance-modelling firm estimated one. So this page gives you the two things that are real, the timeline and the duration, and then does the only arithmetic anybody can check.

Timeline

What happened, hour by hour

Time (Eastern)Event
07:30 ETAutomated scaling activity triggers internal-network instability
07:30 to 08:30Internal API failure cascades; public AWS APIs begin returning errors
08:30 to 09:00AWS Service Health Dashboard updated; many customer-facing services impacted
09:00 to 13:00Peak impact period; most us-east-1-hosted services unavailable or severely degraded
13:00 to 15:00Recovery begins; AWS engineers restore internal-network capacity
~15:00 ETAWS declares broad service restoration; some downstream services continue to recover
Following 30+ hoursCustomers with single-AZ dependencies inside us-east-1 see extended recovery tails

Timeline from AWS's official post-incident summary and contemporaneous status-page updates.

Affected Services

Selected major customers impacted

The list below is not exhaustive. Thousands of smaller services hosted in us-east-1 were also affected. The selection here illustrates the breadth of consumer-facing impact across streaming, finance, food delivery, IoT, transport, and messaging.

Company / serviceObserved impact
NetflixStreaming impaired across multiple regions
Disney+Login and playback failures
RobinhoodTrading platform issues during US market open
CoinbaseTrading and account-access issues
SlackMessaging delays and partial outages
RingDoorbell and camera notifications stopped
Roomba (iRobot)Cloud-controlled robots unresponsive
TinderApp failures during peak evening hours later
VenmoPayment processing issues
Amazon retail siteSome product pages and customer-account features impaired
DoorDashOrder processing degraded
McDonald's appMobile ordering down
United AirlinesBooking and check-in issues
Delta Air LinesSome booking pathway issues

Root Cause

The internal-network scaling cascade

Per AWS's official summary, the trigger was an automated scaling activity at 07:30 ET on the internal network that hosts AWS's internal networking devices. The scaling triggered an unexpected behaviour in the clients of the internal network, which began a connection-storm against the internal network. The connection-storm consumed the network's remaining capacity, which produced cascading failures across the internal services that the public AWS APIs depend on.

Because the AWS Service Health Dashboard itself partly depends on the same control-plane services, customers experienced an extended period during which they could see something was wrong but could not get reliable status information. This compounded the operational impact: customer engineers were debugging blind for the first hour or more.

Recovery required AWS engineers to restore the internal-network capacity carefully without re-triggering the connection-storm. The recovery process took approximately five hours from the start of active mitigation to broad service restoration, with a long tail of downstream impact as individual customer services worked back to nominal state.

Economic Impact

Nobody published an aggregate cost. Here is the arithmetic instead.

AWS does not disclose customer-impact figures, and for this incident no analyst or insurance-modelling firm published one either. That makes December 2021 different from the two AWS outages that do carry an external estimate: the February 2017 S3 outage, which cyber-risk modelling firm Cyence estimated cost S&P 500 companies about $150 million and US financial-services firms a further $160 million, and the October 2025 DynamoDB outage, for which CyberCube published an insured-loss range of $38 million to $581 million. For December 2021 there is nothing to cite, so any confident aggregate total you find quoted for it, including on sites that look authoritative, is somebody's unpublished model.

What can be done honestly is the division. The acute outage ran about seven hours. A year holds 8,760 hours, so a company's gross revenue run-rate over the event is its annual revenue divided by 8,760, times 7. That is the whole method, and the table below is just that sum at four scales. Every figure in it is reproducible on a phone, which is the point: it is a yardstick, not a claim about anyone's losses.

Annual revenuePer hour7 hours (acute outage)30 hours (long recovery tail)
$10,000,000$1,100$8,000$34,200
$100,000,000$11,400$79,900$342,500
$1,000,000,000$114,200$799,100$3,424,700
$10,000,000,000$1,141,600$7,990,900$34,246,600

Gross revenue exposure, and a deliberate upper bound: it assumes the outage took you fully offline for the whole window. Subtract the orders that were deferred rather than lost and the margin you never earned on them; add idle payroll, recovery engineering and any contractual penalties of your own. The calculator does that properly with your inputs.

Two structural observations survive the absence of a total. First, exposure concentrated in large consumer-facing customers, because that is where revenue per hour is highest and where a failed request is a lost transaction rather than a retried one. Second, whatever customers lost, the AWS SLA did not return it: credits are 10 to 25% of the service fee, not a share of revenue impact, and a single seven-hour event frequently leaves monthly uptime inside the SLA threshold, so many customers were owed nothing at all.

For why SLA credits return so little, see our SLA credit asymmetry analysis. The us-east-1 case is a textbook example.

Architectural Lessons

Why "multi-region" is not the same as "regionally independent"

Many customers who believed they had multi-region resilience discovered, during the December 2021 incident, that their architectures had implicit single-region dependencies. Three patterns recurred. First, IAM and certain Route 53 features have control planes that are anchored in us-east-1, so a us-east-1 incident can affect identity and DNS operations even in other regions. Second, cross-region replication has its own control-plane dependencies that can fail open or fail closed in surprising ways. Third, many customer deployment pipelines, monitoring stacks, and operational tooling lived in us-east-1 because it was the original AWS region, so customers could not deploy fixes to their other regions during the incident.

The practical lesson is that true regional independence requires explicit testing through game days (controlled regional failover exercises) rather than just on-paper architecture diagrams. Customers who had recently run a us-east-1-loss game day generally recovered fastest. Customers who had multi-region architecture only on the deployment diagram discovered missing dependencies during the actual incident.

For the cost-benefit math on multi-region active-active versus single-region with strong backup, see our business case builder. The us-east-1 December 2021 incident is the most commonly cited reference point for the "why a single AWS region is not enough" argument, even though pure regional failures are still rare in absolute terms.

Recovery Tail

Why some customers were still recovering 30 hours later

AWS declared broad service restoration by approximately 15:00 ET on 7 December 2021. For many customers, the actual return to nominal service took much longer. The pattern was uneven. Customers with predominantly stateless workloads (read-mostly web services, content delivery) recovered quickly after AWS APIs returned. Customers with stateful workloads (databases, queues, event-streaming pipelines) took longer because they had to drain backlogs, reconcile inconsistent state, and unwind partial-failure conditions accumulated during the outage.

Some customers reported continuing partial impact past 30 hours after the initial incident. These were typically customers with deep single-AZ dependencies inside us-east-1 (services that ran in a single Availability Zone, depended on storage volumes in that AZ, and relied on operational tooling that also ran in that AZ). This is why the headline seven hours understates the event for part of the customer base: for those long-tail customers the effective outage was 24-hour-class or worse, which is the reason the table above prices a 30-hour window alongside the acute one.

Frequently Asked

Common Questions

What caused the AWS us-east-1 outage of 7 December 2021?
Per AWS's official summary, an automated scaling activity on the internal network at 07:30 ET triggered an unexpected client behaviour that produced a connection-storm. The connection-storm consumed the internal network's remaining capacity, which produced cascading failures across the internal services that power the public AWS APIs. Recovery required carefully restoring internal-network capacity without re-triggering the storm.
How long did the AWS us-east-1 December 2021 outage last?
Approximately 7 hours of acute outage (07:30 to ~15:00 ET on 7 December 2021), with recovery tails extending past 30 hours for some customers. Many customers experienced extended partial-impact periods because their architectures had stateful components that took time to drain backlogs and reconcile inconsistent state.
How much did the outage cost in aggregate?
Nobody knows, and no credible aggregate figure exists. AWS disclosed none and no analyst or insurance-modelling firm published an estimate, unlike February 2017 S3 (Cyence, about $150M to S&P 500 companies) or October 2025 (CyberCube, $38M-$581M insured). Any confident total quoted for December 2021 is somebody's unpublished model. What is checkable: the acute outage ran about 7 hours, so a business on $100M of annual revenue was exposed to roughly $79,900 of gross revenue ($100,000,000 / 8,760 hours x 7), and one on $1B to roughly $799,000 - both upper bounds. Service credits returned very little either way: the AWS SLA pays 10 to 25% of the service fee, not a share of revenue impact, and a single 7-hour event often leaves monthly uptime inside the SLA threshold.
Which major services were affected?
Netflix, Disney+, Robinhood, Coinbase, Slack, Ring, Roomba, Tinder, Venmo, McDonald's mobile ordering, DoorDash, and many more. The official Amazon retail site itself was partly affected. Thousands of smaller services hosted in us-east-1 were also impacted. The breadth reflected the historical concentration of new-service deployments in us-east-1 plus the control-plane anchoring of certain AWS services in that region.
Did SLA credits compensate the cost?
No, not meaningfully. The AWS SLA returns 10% of the monthly service fee when uptime falls below 99.99%, rising to 25% for deeper breaches. For a customer running $100,000 per month on the affected service, that is a $10,000 to $25,000 credit. Against a $10 million business loss the credit covers 0.25%. The credit is signalling, not compensation, at any meaningful outage scale.
What is the architectural lesson?
True regional independence requires explicit testing through game days, not just deploy-twice. Many customers who believed they had multi-region resilience discovered implicit single-region dependencies during the actual incident: IAM and certain Route 53 features anchored in us-east-1, cross-region replication control planes, deployment pipelines and monitoring stacks that lived in the affected region. Multi-region in architecture diagrams is not the same as multi-region in operation.

Related

Updated 2026-04-27