The AI Compute Reset Already Happened (And Your 2027 Budget Is Priced Off the Wrong Year)
Introduction: Anthropic Is Paying $1.25 Billion a Month for One Building in Memphis
Anthropic pays SpaceX $1.25 billion a month for a single data center. The building is Colossus 1 in Memphis: more than 300 megawatts of energized capacity and over 220,000 Nvidia GPUs, taken whole, on a contract to May 2029, its terms public through SpaceX's Form FWP. Fifteen billion dollars a year against a 300-megawatt floor puts the ceiling at $50 million per megawatt.
That figure is usually read as a warning that AI compute is about to get more expensive. The warning is late. The repricing has already run: the one-year contract series rose from October 2025 to March 2026, and the spot market peaked in early May 2026. The latest independent reading points the other way: JPMorgan put the average H100 on-demand rental in the non-hyperscale cloud market at $2.68 per GPU-hour in July 2026, down 1.1% on the month and the first decline after seven consecutive increases. It came via a market-news aggregator, not the primary note, and is provisional until that note is in hand.
The durable statement is narrower. The floor under AI compute moved up in late 2025 and is not returning to 2024 levels, the rate of increase is already moderating, and memory constrains supply now while power constrains it from 2027. A 2027 budget built on the 2024 price collapse is priced off the wrong year.
The Price Rise Everyone Is Quoting Is Real, and the Contract Series Ran From October 2025 to March 2026
SemiAnalysis maintains an H100 one-year rental price index built from a monthly survey of more than 100 market participants and validated against transaction data. It recorded $1.70 per GPU-hour in October 2025, a break above $2.00 in late January 2026, and $2.35 at the end of March, almost 40% above the October low. Downstream re-reports of those figures all trace back to SemiAnalysis and count as one source.
Three other readings move the same way without sharing that origin. Ornn's live-transaction index put Blackwell spot at $2.75 per GPU-hour in mid-February 2026 and $4.08 by mid-April, a 48% rise. Lambda raised H100 SXM on-demand from $2.99 to $3.99 in June 2026. Verda moved the same card from $2.29 to $3.25 in early 2026. One independent index and two price cards agree on direction.
Four Different Prices for the Same H100, and Why Commentators Mix Them Up
The same GPU carries at least four distinct prices. The published on-demand rate is sticky by design and lags the market. The one-year contract rate is where most volume transacts, the bulk of it in commitments of six months or longer. Spot is the most volatile, and a composite index blends them. Reserved commitments run 20% to 40% below on-demand. Treating a move in one series as a statement about the market is the error underneath most commentary.
Sold Out Is Not a Figure of Speech: Booked Through September on Half the Providers Surveyed
SemiAnalysis described Blackwell availability as very tight, with capacity arriving through August and September already booked and roughly half of surveyed providers completely sold out. Observed B200 rates in August 2026 span $2.71 to $27.04 per GPU-hour, with GCP's a4-highgpu-8g at $16.11. That is the supply condition every number below sits on.
The Index That Didn't Move: Still 57% Below the 2023 Peak
Across the exact window in which the contract rate rose almost 40%, the composite spot-contract index barely moved: $2.84 in October 2025, $2.80 in March 2026, $2.82 in April. Its second-half 2023 peak was $6.62, leaving the composite 57% below the top of the last cycle.
The contract series rose off a distressed floor while the composite still carries spot pricing and legacy contracts rolling down from higher levels. The on-demand range confirms the direction independently, moving from $1.45–$1.95 to $2.10–$2.70 over the same months. On the composite, anyone buying H100 capacity in 2026 pays far less than a buyer in 2023 or 2024.

Then July Happened: The First Monthly Decline in Eight Months
JPMorgan's July 2026 on-demand reading of $2.68 per GPU-hour broke a run of seven consecutive increases. The bank read it as the most strained phase beginning to ease as supply responded, with utilization still above 70%. Ornn places the peak in early May and describes the months since as normalizing. There is precedent: H100 rates fell from roughly $8 an hour to roughly $2 across 2023 and 2024 as supply caught up.

The Trackers and the Bank Cannot Both Be Right
Several August 2026 trackers report H100 approaching $3 and B200 at $5.50 to $5.80, up from below $5 in January. JPMorgan and Ornn report a May peak and a decline since. The likeliest reconciliation is that lagging contract and hyperscaler list prices are still ratcheting while non-hyperscale on-demand turned over between May and July. Settling it needs the authoritative series, and the public index page rendered values only through April 2026.
The $2.60 That Doesn't Exist, and the 52 Weeks That Belong to Memory
The most-quoted number here is an H100 contract rate of $2.60 per GPU-hour in August 2026, from a widely shared chart posted by a paid analyst at Milk Road PRO, a subscription research product that runs published portfolios; Milk Road's public account has urged buying CoreWeave and Nebius. The 40% move and the $1.70 base are correctly taken from SemiAnalysis; the August date and the $2.60 endpoint appear in no source. Use the endpoints that exist: $1.70 in October 2025 to $2.35 in March 2026.
The companion claim, GPU lead times of 36 to 52 weeks, fails the same way. It circulates across at least six 2026 industry analyses and traces to no primary statement from Nvidia, TSMC or any named analyst house. Value Add VC attaches 52 weeks specifically to HBM backlogs, suggesting transposition rather than measurement.
Attributable facts explain the shortage better. TSMC's CoWoS advanced packaging is fully allocated, supply-chain analyses report SK Hynix and Micron sold out of their entire 2026 HBM production, and HBM costs rose roughly 30% in the fourth quarter of 2025 alone. Industry analyses put order backlogs near 3.6 million GPUs in April 2026, with demand running 1.4 to 1.6 times supply over 18 to 24 months.
What a Megawatt Earns: $2–3 Million at a Colo, $8–11 Million at a Neocloud, a $50 Million Ceiling at Colossus
Three businesses sit on the same megawatt. Colocation sells space and power while the tenant owns the silicon. A neocloud sells GPU-hours from debt-financed hardware in leased space. A hyperscaler owns the power and increasingly the chips, and sells the platform above them.
CoreWeave's first-quarter 2026 revenue of $2.08 billion annualizes to $8.32 billion against roughly 1 gigawatt active at 31 March 2026, about $8.3 million per megawatt. Nebius guides to $7–9 billion of ARR against 800 megawatts to 1 gigawatt connected, roughly $7–11 million. IREN targets over $4 billion against about 480 megawatts, roughly $8 million. Digital Realty's FY2026 revenue guidance of $6.6–6.7 billion against 2,408 megawatts of UPS-backed capacity gives $2.75 million. Colossus 1 gives up to $50 million.

Contracted Gigawatts Are Not Energized Gigawatts
Every figure above is denominator-sensitive. CoreWeave held roughly 1 gigawatt active against 3.5 gigawatts contracted at 31 March 2026. Nebius shows the same gap, over 3.5 gigawatts contracted against 800 megawatts to 1 gigawatt connected. Dividing by contracted power understates revenue per megawatt threefold or fourfold. Run-rate compounds it: an exit-velocity number over a base still growing. A figure quoted without its denominator cannot be checked.
Why the Colo Tier Isn't Measured the Same Way as the Other Two
The colocation band that circulates, $3.5 million to $4.4 million per megawatt, is a rate card rather than realized revenue. It reproduces tier-1 retail pricing near $370 per kilowatt per month times twelve, traceable to a single Q1 2026 pricing guide and not independently corroborated. The filings-derived figure for Digital Realty is $2.1–2.75 million, because it divides revenue by all operating capacity, including unleased space. Comparing the two exaggerates the gap. Equinix is absent for a plainer reason: no comparable megawatt denominator could be located against its FY2026 guidance of $10.144–10.244 billion.
The Top Tier Is a Scarcity Rent, Not a Moat
The $50 million ceiling holds because frontier-generation clusters are sold whole, under scarcity, to buyers with no alternative. The Google agreement with SpaceX runs at $920 million a month from October 2026 to June 2029 for 110,000 GPUs across more than 100 megawatts, or $71–110 million per megawatt on a floor denominator. None of that describes structurally superior economics. It describes a rent, and a rent that size is collectible only while the input stays scarce.
The Input That Actually Reprices: PJM Wholesale Power Up 75.5% in a Year
PJM wholesale power averaged $77.78 per megawatt-hour in the first quarter of 2025 and $136.53 in the first quarter of 2026, a rise of 75.5%. The capacity auction cleared 6.8 gigawatts short of the reliability requirement, the third consecutive failure. About $6.3 billion of $16.4 billion in capacity charges for the 2028/29 delivery year is attributable to data-center demand, $29.4 billion across four auctions.

Gartner Said This in November 2024, and the Part Nobody Quotes Is the Price Part
Gartner predicted on 12 November 2024 that power availability would operationally constrain 40% of existing AI data centers by 2027. Bob Johnson, the named VP analyst, argued utility expansion would not keep pace from 2026. The clause almost nobody quotes is the one most on point: Gartner forecast the cost of power itself would rise significantly as operators used their bargaining position to secure supply. A handful of outlets carried one release, and the forecast is now 21 months old.
The Bill Arrives Whether or Not You Rent a GPU
Retail electricity in data-center-heavy states has already repriced. IEEFA and E&E News put Illinois near 23.85 cents per kilowatt-hour, up about 28% year over year, and Virginia near 17.61 cents, up about 15.4%. Gallup finds 71% of respondents oppose a data center in their neighborhood. Local consent is part of the 2027 timeline, not separate from it.
Power Binds in 2027. Memory Binds Now.
As reported in August 2026, Chamath Palihapitiya argued that power is the single binding constraint on AI compute, ranking hyperscalers above neoclouds above model makers. The fact pattern is sound. The timing is not.
Named principals point elsewhere for 2026. Sam Altman and Brad Lightcap identified memory supply, not power, as the primary bottleneck for training and inference in March 2026. SK Hynix and Micron are sold out of HBM for the year, and TechInsights reports Nvidia designing newer inference silicon around HBM limits. The constraints bind on different clocks: memory determines how many GPUs ship in 2026, power how much capacity can exist from 2027.

The Anchor Tenants Are Building Their Own: Colossus, 5 Gigawatts of Trainium, and On-Site Gas at Stargate
On 20 April 2026 Amazon committed up to $25 billion to Anthropic against Anthropic's commitment of more than $100 billion over ten years to AWS and up to 5 gigawatts of new Trainium capacity. Anthropic separately expanded with Google and Broadcom for multiple gigawatts.
That deal is the weaker illustration, since buying custom silicon from a hyperscaler is not becoming a neocloud. The stronger ones are Colossus 1, where a model maker took an entire facility, and Stargate, a $500 billion program with more than 9 gigawatts planned across seven sites, at least three running on-site gas to bypass interconnect queues. OpenAI and SoftBank each put $500 million into SB Energy, an equity position in generation. Epoch AI's tracking supplies the corrective: 0.3 gigawatts is live at Abilene, and 9 gigawatts is a plan. The customers whose volume subsidized the rental market are leaving it.

The Case That All of This Unwinds in 2027
Man Group, an institutional asset manager with its own positions in the trade, argues that power infrastructure built for 2024 and 2025 demand risks becoming stranded by 2027 and 2028, with a first wave of defaults as initial lease terms renew. CoreWeave's 49 data centers are essentially all leased from third parties, the silicon depreciates on a four-to-six-year clock, and the debt behind it frequently outlasts the contracts it was underwritten against.
The honest counterweight is that no current indicator points negative. Chip revenue, cloud sales and pre-leasing all remain strong. This is a forecast about 2027 and 2028, not a 2026 observation. The depreciation schedules the neoclouds use to carry that silicon are themselves contested.
What a Mid-Market Operator Should Actually Do About Any of This
The spread across procurement modes is wider than the spread across time. Hyperscaler on-demand H100 runs $3.90 to $12.29 per GPU-hour, with AWS at the top of that range and GCP between $9 and $11.50. CoreWeave reserved runs about $2.06 to $2.49, specialist marketplaces $1.49 to $2.69. Reserved sits 20% to 40% under on-demand, Verda spot 65% to 70% under. The observed 2026 shift is hybrid: a reserved baseline, spot for spikes.

Your Real Problem Is Probably 5% Utilization, Not a 40% Rate Move
An unnamed 2026 industry report puts average enterprise GPU utilization near 5%, driven by over-provisioning and inefficient scheduling. That figure traces to no named publisher and should be read as directionally right and specifically soft. Better established, though the sourcing is FinOps-vendor-adjacent: inference takes roughly 85% of enterprise AI budgets, most teams find spend running two to three times expectation, and agentic workloads consume 5 to 30 times more tokens per task than a chatbot call. A buyer at single-digit utilization overpays by an order of magnitude whatever the rate.
Match Contract Duration to Workload Certainty, Not to Price Fear
Locking three years of capacity at a cyclical peak, into a market with a credible 2027 oversupply scenario and hardware that obsoletes in four to six years, is its own way to get hurt. So is budgeting on continued deflation. Neither is solved by predicting the price. Both are solved by matching contract duration to workload certainty and splitting a reserved baseline from burst capacity.
Two Exchanges Are Racing to List Compute Futures, and You Can Use the Index Today
ICE and Ornn announced cash-settled GPU compute futures on 19 May 2026, referenced to the OCPI and pending regulatory approval. CME Group and Silicon Data announced a competing market on 12 May. Neither is live, and ICE has given no timeline. The usable part today is the index, on Bloomberg Terminal since 2 April 2026, giving buyers a public benchmark to negotiate against instead of a vendor quote. Two exchanges competing is evidence the volatility is real. A working forward market is also how a squeeze gets priced and dissipated.
If Your Compute Bill Is an API Bill, This Is a Different Story Entirely
None of the above explains a rising token bill. The API market has split. Commodity inference fell roughly 80% between early 2025 and early 2026, GPT-4o input going from $5.00 to $2.50 per million tokens in twelve months. Frontier pricing has doubled since January 2026, with GPT-5.6 Sol at $5 per million input tokens and $30 per million output. Add the agentic multiplier noted above and the cause is model mix and volume. Attributing it to neocloud rents would be a factual error.
Conclusion: The Floor Moved in Late 2025 and Isn't Going Back
The era of AI compute as a deflating commodity ended in late 2025. What followed was a cyclical upswing off a distressed trough, layered on a multi-year decline that has not reversed, and it is already moderating. Memory binds now, power from 2027, and the largest buyers are exiting the rental market by building their own.
$50 million per megawatt caps what the scarcest input in this industry is worth to a buyer with no alternative. The operators who get hurt over the next two years, the 2027 and 2028 window Man Group flags, will be those who budgeted for continued deflation, or who locked multi-year capacity at the top of a spike and then ran it at 5% utilization. Neither mistake requires a price forecast to avoid.