On 20 October 2025 a fault in the internal name resolution of a single American data centre region brought down large parts of the network. In its own analysis the operator describes a rare race condition in its automation that made a database endpoint unreachable and cascaded from there into numerous further services. The incident lasted around fifteen hours; fault reports covered more than a thousand offerings worldwide, including in Europe.
The obvious reaction, to run everything in-house again, falls short. Innovation, security updates and AI functions today come from continuously maintained services, and specialists for safe in-house operation are scarce.
The better question is one of spread: in which region does the business-critical data sit, which services hang on a single provider, and what happens in daily operations if that provider is unavailable for a day?