There is no such thing as the cloud. There is only someone else’s computer.
Unless you’ve spent the last few months living in an unmapped subterranean cavern, you have seen the headlines about the war in the Persian Gulf. The United States decided to poke the hornet’s nest with a very large stick, and Iran stung right back. Somewhere beneath the military briefings and geopolitical posturing, the cloud developed a rather inconvenient problem with being made of actual stuff.
A staggering amount of military and geopolitical chaos has unfolded, but for those of us who live and breathe systems infrastructure, one headline stood out like a siren in a server room: the physical damage inflicted on Amazon Web Services datacenter infrastructure in the Middle East.
As Tom’s Hardware reported in its September 16 update, AWS was directing customers toward other regions after attacks damaged facilities in the UAE and Bahrain. AWS could not restore access to resources and data hosted exclusively in the UAE’s mec1-az2 availability zone. Recovery work was continuing, with no firm service-restoration timeline.
Somewhere, an architecture diagram still has a cheerful little cloud icon where all of that used to work.
If you have ever built an architecture on AWS, their Solutions Architects have pounded the same gospel into your skull until your ears bled: deploy multi-AZ, mirror your databases across secondary regions, build for failure. For years, cynical CFOs and budget-conscious engineering leads rolled their eyes, assuming Amazon was just manufacturing excuses to double-dip on compute and egress fees.
But the brutal, unvarnished reality is that hardware breaks, fiber cuts happen, and there is no such thing as 100% uptime. Sometimes the failure involves a building you can no longer use and data you can no longer retrieve. You are going to need something more substantial than a stern email to your account manager.
Questioning a cloud bill is perfectly reasonable. Treating every dollar spent preparing for a disaster as waste because the disaster hasn’t happened yet is how you arrive at the emergency meeting where everybody wants a second copy of everything, somewhere else, available immediately. The purchasing department has been asked to acquire yesterday.
The Trap of the Centralized Internet
I wrote an article back in 2025 titled The Consuming Cloud – Centralizing the Internet warning about this structural failure mode. My argument wasn’t just about our social platforms and media feeds coalescing into a handful of monolithic silos; it was about the dangerous centralization of the underlying physical infrastructure.
We spent decades encouraging organizations to abandon on-premise hardware, shutter regional server rooms, and migrate critical databases, financial ledgers, and hospital records into three corporate landlords: AWS, Microsoft Azure, and Google Cloud. There are real efficiencies in that arrangement. There are also consequences when you herd the operational machinery of civilization into concentrated geographic clusters: you are painting giant, radiant bullseyes on physical territory.
The abstractions are lovely. The buildings still have addresses.
Meanwhile, the industry has arrived at the galaxy-brain pitch: let’s put datacenters in low-Earth orbit! Google’s Project Suncatcher is exploring space-based AI infrastructure. Because apparently, if managing cooling systems on Earth is tricky, adding radiation, orbital logistics, and the inability to send a technician over in a van will be an absolute breeze.
Personally, I would like to see a successful restore exercise here on Earth before giving the backup administrator a launch window.
The Asymmetric Drone Revolution
War has always been an endless, brutal pendulum: a new defense emerges, someone invents a weapon to bypass it, which forces a new defense, forever. The drone problem adds an ugly economic question: how long can you afford the exchange?
The attacker’s equipment does not need a price tag remotely comparable to the infrastructure it damages. And defending against repeated threats has to be sustainable. NATO’s military leadership has explicitly discussed the need for lower-cost defensive options. Winning the interception and losing the budget is a deeply stupid way to discover that arithmetic also participates in warfare.
There are countermeasures, and they are evolving. The infrastructure lesson is that nobody gets to assume an expensive building is safe because damaging it ought to be prohibitively expensive. Your datacenter’s construction budget does not function as a force field.
Personally, I think the people selling us the ethereal future should spend a little more time explaining how the concrete survives. But what do I know? I just study cybersecurity.
The Cloud Is Strategic Infrastructure
A regional datacenter concentrates an extraordinary amount of operational dependence behind some remarkably boring walls. Banking, logistics, telecommunications, port shipping manifests, and emergency coordination can all depend on systems hosted there. The fact that the exterior looks like a warehouse for discount patio furniture does not make the consequences of losing it small.
If an artillery shell or explosive drone punches through a roof, you aren’t just replacing a few motherboards. You may be rebuilding liquid chillers, backup generators, fiber-optic distribution frames, and industrial high-voltage switchgear. The recovery involves actual construction, actual equipment, and people who have to work at an actual site.
You cannot redeploy a transformer with Terraform.
A region that hosts your AI cat-photo generator can also host services somebody needs to get paid, move medicine, or coordinate an emergency. There is a lot of real life behind those boring walls, and the consequences do not disappear because the outage happened inside a windowless building.
That building also does not automatically become a lawful target. The ICRC’s guidance on datacenters in armed conflict makes clear that civilian datacenters are protected; military-objective status requires individual assessment, and proportionality and precautions still apply.
For the people designing systems, the practical problem remains: a physically damaged facility may be unavailable far longer than the business can tolerate. The recovery plan needs somewhere else to go.
Architectural Postmortem: How We Survive Kinetic Cloud Outages
So where does this leave us? We cannot put the digital genie back in the bottle. We aren’t going back to storing accounting records in filing cabinets or hosting corporate intranets on a beige Pentium tower sitting under the receptionist’s desk.
Surviving this era requires a ruthless, unsparing shift in how engineers, CTOs, and systems architects model risk. We have to stop treating “The Cloud” like a divine, ethereal sky kingdom and start recognizing it for what it actually is: reinforced concrete, copper piping, and fragile silicon sitting on the geopolitical fault lines of planet Earth.
That gives us three practical requirements.
1. Recovery outside the affected region
Multi-AZ within a single region still matters. AWS availability zones are designed with independent power, networking, and other infrastructure to reduce correlated failures. That protection does not eliminate the need to plan for a disruption that affects an entire region or makes it inaccessible.
Real resilience means having a recovery path outside the failure you are trying to survive. That might involve another region, another provider, local bare-metal infrastructure, or some combination. It depends on how long the service can be down, how much recent data it can afford to lose, and what your team can actually operate.
Active-active across multiple clouds is one option. It is also a fine opportunity to build something so complicated that only one exhausted person knows how it works. Choose the architecture against your recovery requirements, then prove it works. Three vendor logos on a slide do not count as proof.
2. The resurgence of sovereign edge computing
Critical services such as water control systems, municipal emergency routing, and access to essential healthcare records need explicit plans for losing their remote dependencies. A local service should not discover during an outage that it needs permission from a computer three countries away to do the thing it was installed to do.
Where safe and feasible, that means local systems capable of disconnected “island mode,” with the data, authorization, and capacity required to keep essential functions operating. Some functions may need to stop safely. Decide which ones, test how long the system can remain disconnected, and work out how records will reconcile when connectivity returns.
Putting a container on a small computer does not answer those questions. It does give you a small computer with a container on it. Congratulations on completing step one.
3. Independent, protected recovery copies
Replication is not a complete backup strategy. AWS’s disaster-recovery guidance distinguishes continuous replication from protection against corruption or destruction. Logical damage can propagate; independently retained recovery points give you somewhere to return.
If physical disks are destroyed, surviving replicas do not obediently erase themselves. The questions are whether a usable copy survived, whether it is outside the same disaster, and whether you can actually recover from it.
Use geographically separate recovery copies with retention, isolation, and access controls appropriate to the workload. Immutable, air-gapped, and offsite describe different protections. None makes storage “kinetic-proof.” You cannot negotiate with an explosion by explaining that the backup product has an excellent compliance brochure.
A copy in the company’s local equipment room may be useful during a cloud outage. It needs another layer of protection if that room is part of the same disaster you are planning around.
Now Try the Damn Restore
Pick an essential service. Assume its primary region is unavailable. Then have the team restore it using only the resources the plan says will survive.
Can they get the encryption keys? Can they authenticate without the missing environment? Is the configuration available? Does the restored service have usable data? How long does the whole thing take while somebody hunts for the one person who understands the permission error?
Measure the result against the recovery time and data-loss limits you promised. Fix the failures while they are still part of an exercise.
A green backup dashboard is reassuring. So is a smoke detector with a little light on it. Occasionally, you should find out whether the fucking thing works.
We have spent years building threat models full of zero-day exploits, ransomware gangs, and deployment scripts that should never have survived code review. Keep all of that. Then make room for the possibility that the next major threat to your uptime arrives through the roof.
It might be shrapnel. Plan your architecture accordingly.



