Перейти к содержимому
Security

8.5 Million Blue Screens: the Day a Security Update Broke the World

No attacker required — one content push grounded airlines, delayed surgeries and froze banks.

·4 мин чтения ·vev.dev

On 18 July 2024, CrowdStrike — the security vendor whose software runs on computers inside enterprises that operate critical services — released an update to that software. Within hours, machines around the world stopped working.

Nobody attacked anything. There was no ransomware, no stolen password, no intrusion. A trusted supplier sent a routine update to software whose entire purpose is protecting those machines, and the update was defective.

Microsoft estimated that the update affected 8.5 million Windows devices — less than one percent of all Windows machines. That proportion is the part worth sitting with. Less than one percent was enough to produce broad economic and societal impact, because the software that failed was installed precisely where it mattered most: inside enterprises running critical services.

The bill for one bad update

Parametrix, the firm behind the estimate, put direct losses to Fortune 500 companies at about $5.4 billion, Microsoft excluded. Roughly a quarter of that list — 124 companies — were affected. Healthcare absorbed an estimated $1.94 billion, with three-quarters of the Fortune 500 healthcare companies hit. Banking absorbed $1.15 billion. Airlines carried the worst cost per company, an average of $143 million each, and all six Fortune 500 carriers were caught. Delta cancelled thousands of flights and took longer than its peers to restore operations; Southwest said it was not directly affected and saw only minimal disruption.

Insurance covered little of it. Cyber policies were expected to absorb only ten to twenty percent of the losses, with CyberCube projecting insured losses of $400 million to $1.5 billion — potentially, on that firm's reading, the worst single loss the cyber insurance sector has seen in twenty years.

The recovery was the second lesson. Microsoft deployed hundreds of engineers and experts to work directly with customers, shared what it was seeing with other cloud providers including Google Cloud Platform and Amazon Web Services, and published manual remediation documentation and scripts; CrowdStrike in turn helped it build a scalable solution to accelerate fixes across Azure infrastructure. Manual is the expensive word in that sentence. A machine that will not start cannot be reached by another update; a person has to get to it. Multiply that by thousands of machines across dozens of sites, and the gap between the airlines that recovered quickly and the one that did not starts to make sense. Microsoft's own conclusion was that the episode showed how interconnected the ecosystem has become — cloud providers, software platforms, security vendors and the customers who depend on all of them at once.

What this means if you are building something

Ask every vendor how their updates reach your production systems. Not whether they test — everyone tests. Ask whether updates arrive automatically, whether you can delay them, whether they go to all customers at once or in waves, and how quickly the supplier can pull one back. If nobody at your company can answer that for your hosting provider, your payment processor or your security agent, that is the finding.

Roll out in stages, your own work included. Nothing should reach every machine, every user or every customer simultaneously. A release that goes to five percent first and waits an hour is not bureaucracy; it is the difference between an incident and a headline.

A rollback plan is one you have actually run. Everyone believes they can revert. Far fewer have timed it, under pressure, on a Friday evening. Test the reversal on staging the same way you test the feature, and write down how long it took.

Map what you depend on. Jonathan Hatzor, co-founder and CEO of Parametrix, put it plainly: "Companies should thoroughly map their service providers and assess their dependency on each." Most businesses cannot list on a single page the outside services their website or their order flow stops without.

Our reading: this was not a security story. It was a change-management story that happened to involve a security vendor. The failure shape — trusted supplier, automatic update, no staged rollout, slow reversal — is available to anyone running a website, an online shop or an internal system, and it does not require an attacker. Asking the questions above costs an afternoon. July 2024 is the reason to spend it.

Sources

Поделиться
Контакты

Есть идея для проекта? Запросите предложение!

Есть проект? Напишите нам, если хотите поработать вместе над чем-то интересным. Или вам нужна наша помощь? Не стесняйтесь обращаться к нам.