Booking GPU capacity right now is a bit like calling the hottest restaurant in town. The reservation book is full into 2028, and the biggest party at the door just had two more people show up. AWS plans to deploy another two million Nvidia GPUs across 2027 and 2028, on top of an earlier commitment of more than one million. Nobody in the alternative Cloud world is shocked by the appetite, but the real question is what happens to everyone standing behind that party in line.
The queue is real, and it’s long
Nvidia’s latest earnings call made the situation plain: supply, rather than demand, is what’s holding growth back, and the company doesn’t expect that to change overnight. It said supply commitments jumped from $119 billion to $279 billion in a single quarter, warning that shortages stretch from memory and silicon vendors to the suppliers building AI data centers.. Meanwhile, the hyperscalers and fast-growing GPU specialists such as CoreWeave and Nebius are scaling at a pace few others can match.
When supply is tight, allocation tends to follow scale. Big buyers place big, long-dated orders and get served first, which leaves smaller providers facing longer lead times, less predictable deliveries, and less room to negotiate.
It isn’t only a GPU story
The AI build-out is also pulling on memory and storage supply across the wider server market, and the effects show up well beyond the accelerator racks. Hosting providers including Hetzner and OVHcloud have raised prices this year, pointing to rising hardware costs and limited component availability. Even a provider that never sells a single GPU is feeling this in its refresh cycle and its margins.
For the sector as a whole, that adds up to a period of infrastructure inflation. Hardware costs more and takes longer to arrive, and every provider has to decide how much of that to absorb and how much to pass on. Those are uncomfortable conversations at renewal time; but the cause is easy to explain, which helps.
Where alternative Clouds have room to move
Here’s the more encouraging angle. When hyperscaler capacity is stretched and largely pre-booked, customers who can’t get what they need, or can’t stomach the price, start looking elsewhere. Alternative Cloud providers are well placed to capture that demand, especially when they offer what the big platforms find harder to deliver at scale: smaller, more flexible deployments, regional data residency, hands-on support, and simpler pricing.
It also helps that plenty of AI work doesn’t need the newest silicon. Inference, fine-tuning, and smaller task-specific models often run happily on previous-generation hardware, so not every customer needs the latest chip. The GPU shortage rewards providers who match the workload to the right hardware.
What it means for the wider sector
Zoom out and the shortage is reshaping the structure of the market. Capacity is concentrating with the biggest buyers, which widens the gap between providers with guaranteed supply and those without. That gap pushes customers toward whoever can offer certainty. It also sharpens the case for regional and sovereign options, where data location and local control matter as much as raw performance.
This also puts pressure on the rest of the ecosystem. Buyers are looking more closely at alternative accelerators and custom silicon, and MSPs and resellers who build AI services on top of Cloud capacity are learning that their roadmaps depend on someone else’s supply chain. Providers who can be a reliable partner in that chain, rather than the cheapest line item, are likely to gain ground.
How CSPs can respond
Nobody can conjure GPUs out of thin air, but plenty is within a provider’s control:
• Plan procurement earlier. Longer lead times call for forecasting further ahead and building relationships with more than one supplier.
• Get more from what you have. Scheduling, multi-tenancy, and better utilization stretch every GPU further.
• Right-size the offer. Reserve top-end hardware for the workloads that need it and steer the rest toward alternatives.
• Revisit pricing and contract terms. Build in ways to handle component cost swings so they don’t land entirely on your margin.
• Sort out power and cooling early. These often take longer to arrange than the hardware itself.
• Be straight with customers. Clear timelines build more trust than optimistic ones.
The bottleneck will ease eventually as new capacity comes online across the supply chain. The habits it forces, from planning ahead to using hardware efficiently, are worth keeping either way.
CloudFest speakers saw this coming!
These questions were front and center at CloudFest 2026. The AI-Powered Cloud Solutions track brought together speakers including Ditlev Bredahl, Julian Chesterfield, and Ivan Mudryy, who covered lessons from the first year of running AI Clouds, how to turn GPU capacity into a viable service, and why right-sized AI infrastructure can suit CSPs and MSPs better than copying hyperscale designs.
Stay in the loop
CloudFest Global 2027 takes place March 15–18 at Europa-Park in Rust, Germany, and the GPU queue will still be a hot topic.
Subscribe to the CloudFest newsletter (below) to be first to hear when registration opens (hint, it’s very soon or may have happened by the time you read this). Subscribers also get free ticket offers, VIP upgrade opportunities, and the latest State of the Cloud Market Report.

