Skip to main content
All Insights

Infrastructure Economics

Questions Before Approving AI Infrastructure Spend

Photo of Todd Smith
Todd Smith · July 3, 2026 · 6 min read

You do not need to understand every aspect of the engineering to hold an AI infrastructure spend to account. The economics are enough. A short list of plain questions tells you whether a proposed spend buys output the business could not get cheaper or sooner some other way, or whether it pays a 'scarcity premium' for capacity nobody checked the company already owned. Each question below points at a place where money gets committed without being tested. Each has an answer a sound proposal can give, and a different one a weak proposal reaches for instead.

How to use this

Bring these questions to any decision that ends in signing for compute: a vendor proposal, a renewal, a capacity-expansion request, a line in a board deck.

Three words carry the weight, so let's pin them down first. 1) A microchip is the expensive accelerator that does the heavy AI math, whatever the brand or approach. 2) Utilization is the share of the hardware you already pay for that is doing real work. 3) A workload is one whole piece of AI work, and it runs in stages, only some of which need the more expensive chips.

That is the entire vocabulary these questions ask of you.

The six questions

1. What share of the capacity we already pay for is doing useful work?

Across real production fleets the expensive chips do useful work about 5 percent of the time, so a company typically holds around twenty times the capacity it uses.¹ A good answer to this first question is a specific utilization figure, measured over a normal stretch with nights and weekends in it. A less reliable answer is a peak number. No number at all means the spend reached the table before anyone checked how much paid-for capacity is already idle.

2. Are we buying new capacity to do work the fleet we own could already do?

New capacity is the costly fix for a shortage the company often created itself. If the fleet runs near a twentieth of its potential, the work the new spend would carry can probably run on hardware already in the racks.¹² A good answer has ruled that out first and can show why the existing fleet cannot absorb the work. A weak one never measured the existing fleet before proposing to grow it. Those are not the same proposal in terms. One did the homework; the other skipped to the purchase order.

3. What does this capacity cost, how soon does it run, and is the price set by scarcity?

The expensive chips are scarce, so renting and building are both priced at a premium and counted in long timelines. Rental rates climbed hard into 2026 with on-demand capacity sold out; building runs millions of dollars per megawatt and waits years on power.³⁴ Ask for the price per unit and the months to running in the same breath. A good answer gives you both and highlights the scarcity premium paid. A weak answer hands you a rate with no timeline, or a timeline with no rate.

4. If this is owned hardware or a multi-year commitment, how much of the earning window sits idle?

Capacity bought outright, or committed on a one-to-three-year cloud term, is a fixed block of earning-time, and the calendar spends it down whether the hardware runs or not.⁵ A good answer accounts for how much of that window will actually carry work, because idle time on a committed asset is not deferred value. It is missed value now gone. A weak answer treats the commitment as fully used while the rest of the proposal quietly assumes it will not be.

5. What would it cost us to leave?

A commitment that is expensive to exit sets the price the supplier can charge at every renewal, and it piles the risk of one outage onto one provider.⁶ A good answer knows the switching cost in engineering time and dollars, and treats keeping that cost low as worth paying for. A weak answer never worked out how the business would walk away, which represents lock-in and future cost commitment.

6. Does this raise our output per dollar, or just our output?

Output is easy to buy. Spend more, get more; that was never the hard part. The test that matters is whether each dollar buys more output than it does today. Recovering idle capacity clears that bar, because it adds work at almost no new cost. Paying scarcity premiums to run work an idle fleet could already carry fails that test. A good answer shows the spend lifting output per dollar above the status quo. A weak one raises output only in step with the money spent, which is scale dressed as progress, and usually the dearest road to the same place.

Read the pattern, not the reply

No single answer decides it but the overall pattern can. A sound proposal a) knows its current utilization, b) has ruled out recovering idle capacity before asking to buy more, c) quotes price and lead time together, d) accounts for the earning window and the cost to leave, and e) shows the spend raising output per dollar. A weak answer skips the utilization number, opens with buying or building, quotes a flattering rate with no clock on it, and treats a multi-year lock as the price of doing business.

Line them up and the difference is plain. The first proposal spends to produce more. The second pays a scarcity premium for capacity the company may already own but not utilize. These six simple questions tell those two apart, without you having to know virtually anything about how the hardware works.

A note on the numbers

The figures here are covered in full across the rest of this series; each is summarized only enough to make the question land. The near-5-percent average is a measured fleet-wide number across tens of thousands of production systems, current to Q2 2026.¹ The pricing, power-wait, and reserved-term conditions are Q2 2026 market facts and will move, so they are reported as ranges with the spread intact, as supply-driven evidence rather than forecasts.³⁴⁵ The switching-cost and outage points come from independent estimates and reported incidents, not from any vendor's own pitch.⁶ As across the series, the argument does not rest on any single number. It rests on the shape of all of them together: capacity already paid for and mostly idle, new capacity priced at scarcity and years out, and a spend decision that owes it to the business to test the cheaper path before committing to the dearer one.

References

¹: Average GPU utilization of about 5 percent measured directly across tens of thousands of production clusters, indicating roughly 20 times over-allocation, with billions in committed compute sitting idle. Cast AI, 2026 State of Kubernetes Optimization Report; ITBrief; Data Center Knowledge (May 2026). cast.ai/reports/state-of-kubernetes-optimization | itbrief.co.uk/story/cast-ai-report-finds-5-gpu-use-in-kubernetes-clusters | datacenterknowledge.com/infrastructure/ai-demand-surges-as-billions-in-compute-remain-locked

²: AI compute demand inside firms outpacing supply, with workloads queuing and the cost of innovation rising when new work waits on capacity. SiliconANGLE (Mar 2026); Vexxhost GPU capacity analysis (Mar 2026). siliconangle.com/2026/03/05/infrastructure-bottleneck-enterprise-ai-needs-hyperspeed-pivot | vexxhost.com/blog/gpu-capacity-crisis-ai-infrastructure-2026

³: H100 one-year rental contract pricing up almost 40 percent from late 2025 into early 2026 with on-demand capacity sold out across providers, and the same chip several times more expensive on the big clouds than on specialist providers. SemiAnalysis H100 rental index (Apr 2026); Spheron GPU cloud pricing comparison (May 2026). newsletter.semianalysis.com/p/the-great-gpu-shortage-rental-capacity | spheron.network/blog/gpu-cloud-pricing-comparison-2026

⁴: Standard data center construction at roughly 11 to 12 million dollars per megawatt as a one-time build cost (20 million or more for AI-optimized), gated by grid-connection waits of four to seven years and the fastest self-powered sites at 24 to 48 months. Aptly Tech (Mar 2026); Tech-Insider on grid waits (May 2026); Archdesk (Apr 2026). aptlytech.com/data-center-buildout-cost-complete-pricing-guide | tech-insider.org/us-ai-data-center-delays-cancellations-7gw-capacity-crisis-2026 | archdesk.com/blog/global-ai-data-center-construction-2026

⁵: Reserved cloud commitments of one to three years requiring payment for the term whether or not the capacity is used, and owned chips depreciating on a calendar as new architectures arrive on a 12-to-18-month cadence. Petronella Cybersecurity (Mar 2026); Introl GPU depreciation analysis (Jan 2026). petronellatech.com/blog/ai-workstation-vs-cloud-cost | introl.com/blog/gpu-depreciation-strategies-asset-lifecycle-optimization-guide-2025

⁶: Enterprise AI platform switching costs estimated at roughly 2.3 to 5.7 times the original implementation cost, and a single-supplier outage (the October 2025 AWS US-EAST-1 disruption, over fifteen hours, roughly a thousand platforms) showing the concentration risk of one provider. Capgemini Engineering via Zenodo (Feb 2026); ThousandEyes outage analysis (Oct 2025). zenodo.org/record/18620726 | thousandeyes.com/blog/aws-outage-analysis-october-20-2025


Six plain questions test whether an AI infrastructure spend buys output the business cannot get cheaper or sooner, or pays a scarcity price for capacity it already owns and is not using.