Tokens per megawatt-hour: the energy economics of frontier AI
At a typical industrial tariff, one megawatt of AI computing uses about $78 of electricity an hour, a small share of what it costs to run and what its tokens earn. The real hurdle is getting a grid connection on the date the chips arrive.
On this page · 8 sections
How Much Infrastructure Is Getting Built#
For frontier AI companies, the labs training the most capable general-purpose models, infrastructure is as much the business as software. Training builds a model on a cluster of AI chips over weeks. Once trained, inferencing is a separate workload: each request arrives as input tokens, and the model generates output tokens one at a time until the answer is done. Labs price both per million tokens, and size the infrastructure that produces them in megawatts.
A public count of that infrastructure comes from Epoch AI, an independent research group that tracks frontier data centres from satellite imagery, permits and filings. Its September 21, 2026 snapshot estimates 13.1 GW of operational IT capacity across 86 sites, with another 22.3 GW projected on those sites. Announced plans are larger still. OpenAI has announced chip agreements, some as letters of intent, for 26 GW: 10 GW with NVIDIA, 6 GW with AMD and 10 GW with Broadcom, all in 2025, and in 2026 an 8 GW campus in Pike County, Ohio, leased for 20 years under NVIDIA’s guarantees. Anthropic’s 2026 deals are stated in gigawatts: up to 5 GW with Amazon, nearly 1 GW of it online by the end of 2026; 5 GW with Google and Broadcom from 2027; up to 2 GW of AMD chips from 2027; and up to a gigawatt on Azure agreed in late 2025.
Building it is expensive. Chips, networking and the building come to about $40 million per megawatt, most of it chips. Sites are sized by the power they can deliver to their computers: at the 10.2 kW rating of an eight-GPU DGX H100 server, one megawatt of server capacity holds about 800 of NVIDIA’s H100 chips, or fewer of the newer, more powerful ones. Epoch’s record for IT power at a single site has doubled roughly every ten months, from 237 MW in 2024 to 946 MW today. That site is xAI’s Colossus 2 in Memphis, whose newer chips do the work of about 1.1 million H100s.

Energy per Token#
A megawatt-hour is an abstraction until you know what one answer costs. A short chat answer from a frontier model can use about 0.3 watt-hours, roughly one second of a microwave oven. One detailed estimate comes from a 2026 Microsoft study published in the journal Joule. Its frontier medians come from three open models above 200 billion parameters, DeepSeek-R1, Llama 3.1 405B and Nemotron Ultra 253B, on H100 servers: assumed node power and throughput fitted to published benchmarks determine the energy allocated to each answer, plus the overhead for cooling and the rest of the facility.
| Symbol | Meaning | Value used |
|---|---|---|
| E query | energy per query, watt-hours | 0.31, the paper’s chat median |
| PUE | overhead for cooling and the rest of the building, on top of what the computers draw | 1.21, or 21% extra, approximately the lognormal median of the paper’s 1.05 to 1.40 range |
| P node | assumed steady-state draw of an eight-GPU node, kilowatts | 7.14, or 70% of the 10.2 kW rating |
| L out | median output tokens in each query regime | 300 for chat, 5,000 for reasoning |
| TPS | aggregate tokens per second in the worked case | 2,322, back-calculated from the chat median |
| 3.6 | converts kilowatt-seconds to watt-hours, since one watt-hour is 3.6 kilojoules | constant |
Input tokens are missing from the equation because it is an output-dominated approximation for short prompts. Prefill reads the whole prompt in one pass, while decoding writes the answer token by token, so the authors approximate a query’s energy by its output length and hold the prompt at 500 tokens. A longer prompt still costs something: it lowers the throughput term and raises the energy per answer. The authors flag the limit themselves. Prefill is under-modelled for long contexts, at 100,000 input tokens it would be substantially higher, and orchestration, tool calls and context management sit outside the study. Most agent input is not new, either. Serving systems keep the processed prompt in a cache, so the next turn reads it back instead of computing it again, and production traces show 90 to 96% of agent input tokens arriving as cache hits. The agent case below therefore counts cached and fresh input separately, at ratios taken from ML.ENERGY’s phase-split measurement and Irminsul’s cache-hit energy.
Chat is cheap in energy; reasoning costs about 13× more in the paper. Across its models and query distributions, chat with a median 300 output tokens uses a median 0.31 Wh, and reasoning with 5,000 uses 3.91 Wh. The 2.9 Wh per ChatGPT query from de Vries (Joule, 2023) used non-production assumptions; Oviedo et al. put such estimates 4 to 20× too high. A megawatt-hour, a thousand units on a power bill, corresponds to about a billion output tokens in this worked H100 case. The actual rate depends on request shape and the serving stack: long contexts lower throughput, larger batches raise it, and tight latency targets cap it. MLPerf measures throughput per system at a stated latency target; the figure here is back-solved from the paper’s medians. These figures describe large generative models; recommender, vision and classical machine-learning workloads use different amounts of energy per task.
| Task | Output tokens | Wh per task | Output tokens per MWh, H100 worked case |
|---|---|---|---|
| Chat query, 500 in | 300 | 0.31 | 0.97 billion |
| Reasoning query, 500 in | 5,000 | 3.91 | 1.28 billion |
| Agent scenario, 60 LLM calls, 4.8 M in (90% cached) | 40,000 | 153 | , |
Provider figures are in the same ballpark. Google’s full-stack measurement of a median Gemini Apps text prompt is 0.24 Wh in May 2025, of which active accelerators account for 0.14 Wh; its June 2026 environmental report repeats the figure rather than re-measuring. Sam Altman put the average ChatGPT query at 0.34 Wh in June 2025, without stating a method, and OpenAI has published nothing since. The independent academic ML.ENERGY benchmark finds reasoning responses at 25× the energy of chat. Its longitudinal analysis shows why throughput and energy diverge: fifteen months of serving-stack upgrades gave 3 to 5× more tokens per second on the same H100s but only 15 to 41% less energy per token, because GPU power rises with utilisation.
The agent row uses the reasoning calibration and adds input energy at assumed ratios informed by other serving measurements: an uncached input token at one-tenth of an output token, near the median phase split ML.ENERGY measures for Llama models on H100s, and a cached read at a quarter of that, from the 14 to 37% of prefill energy a cache hit costs in Irminsul’s 4,096-token tests. Applying those to an 80,000-token frontier context is an extrapolation. The 153 Wh includes node and facility overhead, so it is not directly comparable with GPU-only measurements of agent episodes, which span about 3 to 34 Wh across configurations in a metered SWE-bench study. The task shape is illustrative, informed by production traces of coding agents from GitHub Copilot and TraceLab.
Where is the next big cut in energy per token likely to come from?
From three places at once. The Joule authors put the combined opportunity at 8 to 20× against today’s energy per query, from techniques already in use or close to it, hardware included, and some of it comes from writing fewer tokens. They assess each category separately, then combine selected gains rather than multiplying the full ranges:
- Models, 1.5 to 10×: distillation, training a smaller model on a larger one’s outputs, has cut energy 5 to 10×. Quantisation, running the model at lower numerical precision, adds more, as do mixture-of-experts designs that fire only part of the model for each token and controls that make reasoning models write up to 40% fewer tokens. Small models use one to two orders of magnitude less energy on narrow tasks such as maths and coding.
- Serving software, 1.5 to 5×: scheduling to each service’s latency target, caching, splitting prompt reading from answer writing, a small model drafting tokens for the large one to check, and routing each request to the smallest model that can handle it.
- Chips and facilities, 1.5 to 2.5×: newer chips such as NVIDIA’s Blackwell (B200/GB200); liquid cooling; and power capping. NVIDIA’s MLPerf v5.1 results from September 2025, the last round to run Llama 3.1 405B on both generations, put GB200 NVL72 at 3.0× the per-GPU throughput of an eight-GPU H200 system offline and more than 5× on the interactive test. For Vera Rubin, NVIDIA projects up to 10× more tokens per megawatt than GB200 NVL72 on one reasoning model; its first MLPerf preview in September 2026 posted up to 2.5× GB300 NVL72 per GPU on DeepSeek-R1. These are single benchmarks, not universal multipliers.
The range looks reachable. Google cut energy per prompt 33× between May 2024 and May 2025, mostly through model changes. It reported over 3.2 quadrillion tokens a month at I/O 2026, 7× a year earlier, and the IEA puts the growth in AI-focused data-centre electricity use at 50% in 2025. Each answer gets cheaper while the fleet consumes more.
Energy per Training Run#
Training is the other big use of power, and some of it is publicly disclosed. Meta’s model card discloses 30.84 million H100 GPU-hours for Llama 3.1 405B. At the chips’ 700 W rating, that implies 21.6 GWh, or about 35 GWh with assumed node and cooling overhead of 1.6×. At an assumed $2.50 per GPU-hour, the rental-equivalent compute cost is $77 million; the estimated electricity, at $92 per MWh, is $3.2 million. It remains the largest fully disclosed run: Meta’s Llama 4 card gives 7.38 million GPU-hours across two smaller models, and no frontier model released in 2026 discloses its hours.
| Run | GPU-hours | Estimated energy incl. overhead | Rental-equivalent compute cost | Electricity at $92/MWh | Electricity as share of compute cost |
|---|---|---|---|---|---|
| Llama 3.1 405B, GPU-hours disclosed | 30.8 M | 35 GWh | $77 M | $3.2 M | 4% |
| 100,000 Blackwell GPUs, 100 days | 240 M | 480 GWh | $660 M | $44 M | 7% |
These sites are gigawatt-scale. Stargate Abilene is planned at about 1.2 GW, roughly 10 TWh a year at full draw, with 421 MW of IT running there in September 2026. A training run holds a flat load for weeks and can pause at a checkpoint. Interactive inference follows users, so it cannot pause. According to the Joule authors, the pressure on grids mostly comes from the three things mentioned below, and not from the energy used to give a single answer:
- Training loads: a run draws near-full power for weeks and arrives as one large block, not as gradual growth.
- Pace of adoption: demand is growing faster than transmission can be planned, permitted and built.
- Concentration: gigawatts land at a handful of substations rather than spreading across the network.
Cost of a Megawatt-Year#
Take one megawatt of server capacity and add up what it costs for a year. Filled with H100 servers, rated at 10.2 kW for eight GPUs each, that is 98 servers and 784 chips. The servers average an assumed 70% of that rating, so 0.7 MW, and PUE 1.21 adds cooling and the rest of the building to make 0.847 MW at the meter. Over 8,760 hours that is 7,420 MWh a year. Own it and depreciation dominates the annual budget.
| Item | $ million per MW-year | Share |
|---|---|---|
| GPU and network depreciation | 5.65 | 74% |
| Facility | 0.75 | 10% |
| Electricity, 7,420 MWh at $92 | 0.68 | 9% |
| Operations | 0.60 | 8% |
| Total | 7.68 | 100% |
At $50 per MWh electricity is 5% of the total; at $120, 11%. A B200 megawatt costs about the same, near $8 million, with 546 GPUs at 1.83 kW of system power each and $45,000 apiece, though each of those chips does several times the work. Facility overhead varies too: Google’s fleet averaged PUE 1.09 in 2025, against the 1.21 assumed here. Renting at assumed on-demand rates of $2.50 to $4 per H100-hour would cost $17 to 27 million a year. For context, CoreWeave’s second-quarter 2026 revenue of $2.58 billion works out to about $6.9 million a year for each of the 1.5 GW it had active at the end of June, a little below what our budget says the same megawatt costs to own. Its annualised second-quarter net interest bill of $640 million adds roughly $1.7 million per megawatt-year, the financing our budget leaves out. These are company-wide ratios, not H100 rental quotes.
Spread the annual budget over output and the unit cost appears. At an assumed 75% of the calibrated throughput, the H100 case produces 5.4 trillion output tokens a year. Dividing $7.68 million by that output gives about $1.43 per million output tokens. This is a simplified accounting cost; a full levelized cost would also include financing and the other omitted costs.
The cost per task can keep falling without the hardware budget falling. New chips, better serving software and smaller models can deliver more useful work per megawatt; shorter reasoning and better caching can reduce the work each task needs. The Joule paper’s combined 8 to 20× opportunity includes these effects, so it is not all extra tokens from the same machine. About nine-tenths of this budget is depreciation and other fixed charges, so an idle megawatt costs nearly as much as a busy one. The cost per token depends on how much work the megawatt actually does.
Revenue of a Megawatt-Year#
Revenue is what a megawatt earns by running; the next section prices what is lost when it does not.
At the list prices of Anthropic’s Sonnet 5, the worked H100 case would sell for about $72 million a year. That assumes three-quarters productive output, selling output tokens at $10 per million and the accompanying 500 input tokens per 300 output at $2. Free tiers, subscriptions, cached input and discounts change what is collected. Dividing reported run-rates, which annualise a recent period’s revenue, by an assumed 2.5 GW of live compute gives a different view in the bottom two rows. Anthropic said only that run-rate revenue crossed $47 billion in May. Bloomberg reported in August that Anthropic passed $65 billion at the end of July and that OpenAI reached $40 billion; TechCrunch summarises both reports. Two cautions. Anthropic books the full token price paid through Amazon’s Bedrock marketplace as its own revenue and pays Amazon separately, so the run-rate includes money that goes to the cloud (SemiAnalysis, May 2026), and the two labs may not define run-rate the same way. Both figures also cover whole companies on mixed hardware, so dividing by assumed capacity gives a scenario, not what an inference megawatt earns.
| Price basis | $ million per MW-year | $ per MWh bought (tariff: $92) | Multiple of $7.7 M budget |
|---|---|---|---|
| Sonnet 5 list, $2 / $10 per M tokens | 72 | 9,680 | 9.3× |
| Haiku 4.5 list, $1 / $5 | 36 | 4,840 | 4.7× |
| Gemini 3.8 Flash list, $0.75 / $3.75 | 27 | 3,630 | 3.5× |
| Anthropic, top-down | 26 | 3,500 | 3.4× |
| OpenAI, top-down | 16 | 2,160 | 2.1× |

Break-even sits below that $1.43 because input tokens are sold too: $1.07 per million output tokens, once 500 input tokens per 300 output are charged at a fifth of the output price. Flash’s $3.75 price is 3.5× that threshold, though its actual serving cost depends on the model. Smaller models can support lower prices. In the Anthropic and OpenAI company-ratio scenarios, electricity is 2.6% and 4.3% of revenue.
The Sonnet-price scenario implies an 89% margin over the simplified budget. Reported lab margins use different cost boundaries: Anthropic’s inference gross margin was in the mid-sixties per cent in early 2026, by SemiAnalysis’s estimate, and OpenAI’s compute margin on paid products, revenue minus the cost of running its models, was reported at about 70% in October 2025 (The Information, via SaaStr). By May 2026 Anthropic’s gross margin was reported at about 70% on first-quarter results (SaaStr, third-party). Free tiers, flat-rate subscriptions, discounted cached and batch tokens, enterprise deals, cloud rental costs and differences between models all affect the outcome.
A healthy margin on serving does not settle company profitability. Training the next model, research compute and researchers’ pay must also be funded, and capacity is bought ahead of demand. The two labs now differ here: Bloomberg reported in August that Anthropic was profitable on an adjusted operating basis in the second quarter of 2026, on $11.5 billion of revenue, while OpenAI said in July that its planned infrastructure spending through 2030 had reached $750 billion (Wall Street Journal, via TechCrunch). The scale helps explain the capital raises: OpenAI closed a $122 billion round in March 2026, and Anthropic raised $65 billion in May.
What Waiting Costs#
One hour of that megawatt draws 0.847 MWh, costs $78 in electricity and produces about 615 million output tokens. The company-ratio scenarios put revenue at about $1,800 to $3,000 an hour. If demand is waiting and the work cannot be recovered elsewhere or later, the costly input is time without power.
On a 100 MW campus, one month without power puts $133 to 217 million of sales at risk on those assumptions. Monthly depreciation and other fixed charges of $58 million continue; they must not be added to lost sales. A whole year of cheaper electricity, $50 instead of $120 per megawatt-hour, saves about $52 million.

And the waits are long. The IEA puts data-centre construction at one to three years, while new grid infrastructure can take five to fifteen years to plan, permit and complete. JLL reports average grid-connection waits exceeding four years in primary markets. In Texas, ERCOT was tracking more than 474 GW of large-load requests in June 2026, about 90% of them data centres, against a record peak of 86 GW. By September, ERCOT’s first batch held about 496 GW across 735 projects, of which 66 GW was conditionally classified as base load, 128 GW as studied and 302 GW excluded. On 3 August, Texas’s governor ordered an audit of all data centres advancing through ERCOT’s interconnection process before further approvals could move forward.
The strain already shows in capacity prices. PJM’s capacity auction went from $28.92 per MW-day for 2024/25 to $269.92 for 2025/26 and has cleared at its cap in every auction since, most recently $325 for 2028/29 in July 2026, and PJM says the 5,250 MW increase in forecast load behind the 2027/28 auction is largely due to additional large loads.
Buyers are accepting premiums and long commitments to secure power. Talen describes its Amazon contract as carrying anticipated premium prices, AEP reports take-or-pay minimums of up to 90% of contracted demand on terms as long as 20 years, and Entergy says the additional transmission for Meta’s Louisiana expansion is funded directly by the customer.
What Can Pause#
Grids sometimes ask big customers to use less for a few hours, on the hottest or coldest days. Whether a data centre can say yes depends on its work, which comes in three kinds:
- Training: building a model. A run lasts weeks, saves checkpoints and can pause and resume.
- Batch inferencing: requests that can be answered within hours, such as bulk document processing or agents working in the background. Anthropic and OpenAI sell it at half price, with results within 24 hours.
- Real-time inferencing: answering people as they wait, in chat, search or a coding assistant. It cannot simply pause without affecting service, though routing or power management can sometimes reduce local demand.
Utilities and grid operators pay for this through demand-response programmes, at rates that vary by market and contract. One CAISO-based study, by Du et al., models a $150/MWh reward and a $375/MWh penalty for falling short. Interrupting valuable real-time work can cost more; deferring training or batch work can cost less if it is completed later. Checkpointing, deadlines and catching up still have costs.

It has been demonstrated, so far at modest scale. Emerald AI, Oracle and NVIDIA cut power at a 256-GPU cluster by 25% for three hours during two real utility peaks in Phoenix, meeting service targets that allowed some workloads to slow down (Nature Energy). By March 2026 Emerald AI counted five live demonstrations, including a site in Hillsboro, Oregon, that cut load by a quarter or more on three consecutive days while replaying a historical heat-dome scenario. Google has written 1 GW of such flexibility into its utility contracts to support faster connections. Texas law requires curtailment readiness for specified new transmission-connected loads, with exceptions, and a separate competitive demand-reduction service for loads of at least 75 MW. Duke University estimates that if new loads accept partial cuts for about 177 hours a year, giving up roughly 0.5% of their annual energy, the modelled headroom on the existing US grid is 98 GW without building new generation. The estimate ignores transmission limits. A February 2026 follow-up from the same institute models the decade ahead and finds that flexibility, even when limited to a small share of hours, can cut system costs by tens of billions of dollars.
The flexible share depends on the mix. Half a training load plus all batch work can step back, and how much of a campus that is depends on how the work divides. More interactive work lowers that share; more background agents raise it, and so would training energy overtaking inference, which the Lawrence Berkeley National Laboratory’s 2024 report projected by 2028. What a campus can actually shed is lower again, because idle power, networks and cooling keep drawing. Lower precision offers another possible lever: Du et al. at Tsinghua and HKU simulate it against an FP16 baseline, with quality-loss assumptions of 1 to 8% calibrated from other studies. The benefit depends on quality targets and how much quantisation is already in use.
How long to pause
How long does a pause have to last? Demand-response events typically run a few hours. Chen and Zheng model deferring data-centre load and find the grid gains little once a deferral runs past three hours, and the Phoenix trial ran for exactly that. So three hours is the case used here. Batteries only have to cover the gap between what the grid asks for and what the workload can shed. On a 100 MW site where 25 MW can pause, a call to cut a quarter is covered by pausing that work alone, and a call to cut half needs the other 25 MW from storage for three hours, or 75 MWh. Sizing must also allow for losses, reserves, longer events and recharging. At fleet scale the IEA projects 20 to 25 GW of batteries and 15 to 27 GW of on-site gas serving data centres by 2030.
Conclusions#
- 01Electricity is a small cost share. It is 9% of the ownership budget at $92 per megawatt-hour, and 2.6 to 4.3% of the company-ratio revenue scenarios.
- 02Megawatts on a date matter. A month’s delay outweighs a year of cheaper power, $133 to 217 million against $52 million on a 100 MW campus, when the delayed work is lost.
- 03The flexible share depends on the work. Training and batch work can step back; real-time answers cannot. That flexibility is what buys a faster connection, and batteries and on-site generation cover what the workload cannot shed.
Frontier AI has made electricity a small line on a very large bill. What a premium still cannot guarantee is a connection on time, and that is where the competition has moved.



