A box on a desk, a rack on a hyperscale floor, and the long line that now runs between them.
Enterprise AI compute is not consolidating into the cloud or collapsing onto the desk. It is stratifying into five tiers, and the money, along with the fight, has moved to the seams between them: the routing layer, the colocation floor, the residency rule, the network port.
The takeawayAI PCs are now the shipment majority, but the desk competes only at the shallow end. The work is multiplying faster than any one tier can hold.
In June 2026, AMD opened pre-orders for the Ryzen AI Halo: a 5.9-inch box holding 128 gigabytes of unified memory, fast enough to run a serious model locally with no cloud account, no API key, and no meter ticking on every token. NVIDIA got there first with the DGX Spark, then nudged it to $4,699 when memory grew scarce; AMD strolled in at the old price and threw in native Windows. Verified
It is a tidy story and a useless map.
Cue the headlines about the desk coming for the data centre. The box is the most visible edge of something slower and far more consequential: the AI compute stack is stratifying into distinct tiers that each do a different job, and the money, along with the fight, has migrated to the seams between them.
Lay the tiers out and the argument settles itself, from the AI PC on the desk to the hyperscale campus at the summit. Everyone wants to call this a ladder, with work forever climbing toward the cloud or being dragged back to the edge. It is not a ladder. It is a web. Each tier has a job, the economics that pin it there, and a class of work that belongs to it, and the movement that matters is sideways.
The numbers refuse to cooperate with the war narrative: the AI infrastructure market closed 2025 at $337bn and is on track for $1.2tn a year by 2030, while the on-premises slice alone climbs from $34.5bn toward $171.7bn. Nobody is eating anybody. The work is multiplying faster than any one tier can hold. Estimated
Here is the part they skip at the launch event. The little box competes only at the shallow end – a model fine-tuned with adapters, retrieval over a folder of contracts, a piece of content too sensitive to send anywhere. What it actually does is collect the work that never reached a data centre at all, the analysis that died in a spreadsheet because the cloud was too dear, too slow, or too leaky to trust.
The desk is not stealing demand. It is conjuring it. And the border between the first tier and the second is drawn not by raw compute but by a network port – the Spark ships with a 200-gigabit fabric to cluster, the Halo a lone 10-gigabit jack – which is a very revealing place to find the front line.
By 2026 that front line has moved fast. AI PCs, defined by an NPU meeting Microsoft’s 40 TOPS Copilot+ bar, reached an estimated 54.7 per cent of global PC shipments, up from 31 per cent in 2025 and running near 143 million units, roughly four times the 2024 level. The desk-side tier is no longer a premium niche; it is the shipment majority, even if the high-memory devices this piece opened with remain a narrower slice of it. Estimated
The takeawayThe old cloud-first and on-prem-or-nothing doctrines are dead. The winners build a routing layer that sends each workload to the tier its economics demand.
If the stack is a web, routing is how you use it. The blunt doctrines of the last decade – cloud-first, or on-prem-or-nothing – are being retired for what the trade now calls cloud-right: put each workload where its requirements send it, not where the IT department finds it convenient.
The economics have stopped being a matter of taste.
On-premises kit can pay for itself against cloud rental in under four months on a busy workload, and a million tokens on private hardware can cost ten to fifteen times less than the same tokens through a frontier API over a few years. Cloud still wins below roughly two thousand seats, where its tidiness is hard to argue with.
Past five thousand, once you count compliance, latency and the quiet larceny of egress fees, on-prem and on-device pull ahead. The sharpest operators have stopped choosing edge or cloud at all; they build a routing layer that reads each request and posts it to the tier that fits, and some already run four parts in five at home.
Data centres won the training war. The inference war is still being fought, in every tier at once.
Within two or three years, the consensus goes, about 85 per cent of enterprise AI work will be inference rather than training – precisely the workload that takes happily to the edge and the on-premises floor. Training is a big, concentrated prize; inference is a big, scattered one, climbing from $101bn to $532bn by 2030 and outrunning training by a comfortable margin. Bain calls inference the new centre of gravity, and for once the phrase earns its keep. Estimated
The takeawayThe tier quietly winning enterprise AI is not the cloud, the server or the AI PC. It is colocation, now backed by a hard capacity number.
Ask which tier is actually winning enterprise AI in 2026 and the honest answer is the one nobody toasts at the conference. Not the public cloud, not the on-premises server, not the shiny new AI PC. It is colocation, the least glamorous box on the org chart.
It is colocation, the least glamorous box on the org chart.
The reason is unromantic and decisive. Persistent AI carries the lowest three-year cost and the steadiest bill on dedicated infrastructure, and colocation hands you the power density, the cooling for racks past 100 kilowatts and the compliance paperwork enterprise AI now demands, all without making you pour a foundation. It has become the hinge of the whole stack, the place where on-premises control shakes hands with cloud reach. It will not get a magazine cover. It is winning anyway. Estimated
There is now a number under the claim. A McKinsey Global Institute report in late June 2026 projected AI-related data centre demand growing three and a half times, from about 44 gigawatts to 155 gigawatts, and reaching roughly 70 per cent of total data centre demand. The tier that gets no magazine cover is the one absorbing most of the build. Estimated
The takeawayRegulation is not a handbrake on the build. Residency rules redistribute demand downward, lighting up several tiers at once.
The force tightening every strand of the web is regulation, and most people have it exactly backwards. They treat it as a handbrake. It is an accelerator.
They treat it as a handbrake. It is an accelerator.
The EU AI Act is in full force, residency regimes are breeding across the emerging markets, Gartner expects sovereign cloud spending to near $80bn in 2026 – up more than a third in a year – and Forrester reckons half the G20 will soon insist on home-tuned AI for public services.
Every one of those rules is a demand signal, and it fires across the whole stack at once: it shoves compute away from foreign hyperscalers and toward domestic colocation, on-premises servers and, at the light end, the AI PC, where a small model chews through a regulated document with no API call and no third-party processor to vet. A new residency law does not shrink the infrastructure market. It redistributes it downward, lighting up several tiers in the same stroke. Verified
The spending behind that redistribution is now sized. Gartner puts sovereign cloud near $80bn in 2026, growing 35.6 per cent in a year, with China ($47bn) and North America ($16bn) the largest markets and the Middle East and Africa, mature Asia-Pacific and Europe the fastest-growing regions. It expects a fifth of current global-cloud workloads to migrate to local or sovereign providers, and Europe to overtake North America on sovereign spend by 2027. Verified
The takeawayThe hyperscalers are spending like nations (~$725bn in 2026) and are becoming power-hungry infrastructure utilities, not asset-light software houses. Verified
At the top, the business is changing its nature. The five biggest hyperscalers will spend north of $600bn on infrastructure in 2026, about three-quarters of it on AI, and Goldman Sachs puts their combined 2025-to-2027 outlay above $1.15tn (both since revised up sharply; see the mid-2026 update) – more than double the prior three years.
These are no longer asset-light software houses; they are turning into power-hungry infrastructure utilities.
They are paying for it increasingly with debt and building their own silicon to wean themselves off the merchant chip vendors for internal work. Microsoft has re-signed a nuclear plant for twenty years at a price comfortably above the going renewable rate, which tells you everything. And that, more than any product launch, is what opens the door below: as the giants turn inward and upward toward frontier scale and custom chips, NVIDIA and AMD pivot to the sovereign and enterprise tiers the giants’ own silicon will never reach.
The centre is the number everyone quotes. The corners are who writes the cheques; the edges are who makes them. Every capex figure below has been revised up at least once through 2026. Verified · company guidance, Morgan Stanley, mid-2026
The interesting question about the AI stack in 2026 is no longer which tier wins. It is where the value pools, and the answer is the boundaries between them.
Every tier grew. The desk-side node became the shipment majority, colocation quietly absorbed most of the enterprise build, sovereign platforms turned regulation into demand, and the hyperscale summit kept spending at a scale that reads more like a country than a company. But growth at every layer is not the story. The story is that orchestration, routing and residency, the logic that decides which workload runs where and under whose rules, is now the scarce, priced thing. The tiers are commodities. The seams are the business.
That is a harder market to own than a single winning layer would have been, and a more durable one. A hyperscaler can be undercut; a colocation floor can be overbuilt; a sovereign mandate can be repealed. The coordination layer that spans all of them only gets more valuable as the stack fragments, which is exactly the direction regulation, economics and physics are all pushing it.
My own read goes one layer deeper than the software. The thing that will decide who actually wins the seams is not a smarter orchestrator; it is whether the hardware underneath can be built and fed. And in 2026 the binding constraint on that hardware is memory. The same HBM and DRAM squeeze that repriced the desk-side node is a tax on every tier of this stack, and it is the subject of the next study. If The Continuum is about where compute lands, The Memory Supercycle is about whether there is enough of the one component that makes any of it run. Read that next.
The Continuum is a Sociogencia case inside The Next Hotspot, an interactive read on where the world actually builds AI infrastructure. Chapters land through the summer, the full edition on 4 August, and new original analysis is already going live.
figures and forecasts are drawn from Entelligencia’s research brief, Enterprise AI Infrastructure: The Continuum from AI PCs to Data Centers (June 2026), and its underlying citations, including 451 Research, Gartner, Forrester, Goldman Sachs, Bain, KPMG, MIT, PwC, Grand View Research, and vendor disclosures from AMD, NVIDIA, HP, Qualcomm, Dell, Lenovo, Microsoft and Equinix. Charts are directional, built from the cited endpoints and ranges, not precise measurements. Tier hardware references include the shipping HP ZGX Nano (NVIDIA GB10 Grace Blackwell) and the announced Qualcomm Dragonfly C1000 data-centre CPU, unveiled at Qualcomm’s June 2026 investor day, with general availability targeted for 2028.