Lesson 20 · Chips & LLMs

Datacenter Power & Energy Economics

The bottleneck has moved off the chip and onto the grid. This lesson follows a watt from the substation to the silicon, names the metrics that decide who can actually build, and opens a whole layer of investable names beyond your core seven.

Builds on: L8 (inference), L13 (distributed training), L19 (the power wall) Grounds: why power is the new CoWoS

For three lessons you've watched bottlenecks migrate: from litho (L17), to HBM and CoWoS packaging (L12), to networking (L13). In 2025 the binding constraint moved again — and this time it left the building. You can buy the GPUs; you still can't power them. The scarce resource is now electricity delivered to a place, on a schedule. This lesson makes that constraint quantitative, so you can read a "10 GW buildout" headline the way an engineer reads a node name.

Core thesis: The end of Dennard scaling (L19) pushed the industry toward ever-denser, ever-hotter accelerators. A GB200 die now dissipates 500–600 W/cm² and a single rack pulls ~120–130 kW. Stack thousands of those and a frontier cluster needs 300–500 MW — the load of a small city — on a grid where a new connection takes 7–10 years. So the new moat isn't FLOPS; it's secured power, fast. That reprices utilities, grid equipment, cooling, and nuclear — a value-capture layer financial media is only now waking up to.

01 — The Wall Moved Off the Chip

From the Power Wall to the Grid Wall

Recall L19: around 2005 voltage stopped scaling, leakage took over, and power density became the ceiling on a single chip. The industry's escape was parallelism and specialization — more chips, packed tighter. That escape works at the chip level but simply relocates the heat problem upward. The same physics that made one transistor hard to cool now makes one building hard to power.

The numbers compound at every level. A modern accelerator is ~700–1200 W. NVIDIA's GB200 NVL72 packs 72 GPUs into one liquid-cooled rack drawing ~120 kW nominal (deployed racks report 130–132 kW) — roughly 10× the density of a traditional ~12 kW air-cooled rack. [NVL72 specs] Multiply by thousands of racks and you arrive at the cluster-scale figure that now governs the industry.

500–600
W/cm²
GB200 die heat flux — 40–100× past air cooling
~120
kW / rack
GB200 NVL72 vs ~12 kW legacy rack
300–500
MW
A 100k-GPU training cluster
7–10
years
New large grid connection (US/EU)

Sources: die heat flux, cluster power, connection queues.

02 — Follow One Watt: The Power Stack

Where Every Watt Goes (and Leaks)

To reason about cost and efficiency, trace a watt from the grid to the gate. At each layer some power is lost to conversion, distribution, and — above all — cooling. The single most important accounting tool is PUE.

PUE = Total Facility Power ÷ IT (compute) Power PUE = 1.0 is the unreachable ideal (every watt reaches a chip). PUE = 1.5 means 50% overhead — half a watt of cooling/losses for every watt of compute.
The power stack — one watt from substation to silicon
Grid / Substation UPS / distribution Cooling biggest overhead Rack / GPU the "IT" in PUE Die work ← Total Facility Power (the PUE numerator) → ↑ IT Power (denominator) Best-in-class PUE ≈ 1.1–1.2 · Global average ≈ 1.5–1.6
PUE measures everything outside the green box as overhead. Driving PUE from 1.5 → 1.15 means ~25% less total power for the same compute — which is why hyperscalers obsess over it and why liquid cooling (Section 05), far more efficient than chilled air, is now a competitive weapon, not a nicety. [PUE benchmarks]
03 — The Metrics That Matter

How to Price a Buildout

Headlines quote gigawatts and hundreds of billions. Translate them with four numbers — the energy-economics analog of "density + PPA" from L19:

MetricWhat it measuresReference value (2025)Why an investor cares
PUEFacility efficiency (overhead)~1.1–1.2 best · ~1.5 avgLower PUE = more sellable compute per MW procured = better unit economics
$ / MW (or $/kW)Cost to build capacity~$10–12M / MW all-inThe capex denominator; sets the depreciation that must be earned back
$ / GW campusFull gigawatt-scale build~$45–55B / GWWhy only a handful of balance sheets can play at the frontier
Energy / tokenInference efficiency~0.3 Wh / query (GPT-4o)The marginal cost of serving; ties power to the P&L of every API call

Sources: $/GW (Turner & Townsend via BloombergNEF), energy/query.

The energy/token figure is the bridge back to L8 (inference economics). At ~3–4 joules per output token for a mid-size model, energy is a real and growing line in the cost of every served response — and the lever (FP8/FP4 quantization, better MFU, MoE sparsity) that the most efficient operators pull to widen margin. Hold this thread; it reappears in the inference-economics capstone.

04 — The Binding Constraint: Grid Interconnect

Why Power, Not Silicon, Now Gates Growth

Here is the structural fact that reorders the whole investment landscape. A hyperscale campus can demand as much power as an aluminium smelter or a mid-size city — but a new large-scale grid connection in the US or Europe now takes 7–10 years, with some projects waiting 13. In the PJM region (the largest US grid market), the application-to-operation timeline rose from under 2 years in 2008 to over 8 years in 2025, with a queue exceeding 2,600 GW of pending requests. [grid-impact study]

Meanwhile demand is vertical: AI datacenter power is projected to rise from ~10 GW (2025) to ~68 GW (2030), +160%, and over 23 GW of capacity was under construction globally at end-September 2025 — three-quarters of it in the US. [BNEF] When demand grows that fast against a decade-long connection queue, the scarce asset isn't the GPU — it's a shovel-ready megawatt.

The diligence reframe: The right question about an AI buildout is no longer "did they secure the H100/Blackwell allocation?" — it's "where is the power, is it contracted, and when does it energize?" Power-purchase agreements, behind-the-meter generation, and existing grid interconnections have become the genuinely scarce, defensible asset. This is the 2025 successor to "do they have CoWoS allocation?" from L12.

That scarcity is driving three workarounds, each an investable theme:

05 — Cooling: The Overhead You Can Engineer Away

Air → Liquid, and Why It's Now Mandatory

Cooling is the largest non-IT slice of the PUE numerator, so it's the overhead operators fight hardest. The L19 physics forces the issue: at 500–600 W/cm², a GB200 die produces heat flux 40–100× beyond what moving air can remove. Air cooling is not merely inefficient at this density — it is physically impossible. [liquid-cooling rationale]

ApproachRemoves up toWhere it's usedTrade-off
Air (CRAC/CRAH)~15–20 kW/rackLegacy & general computeCheap, simple; hits a hard density wall
Direct-to-chip liquid~120–150 kW/rackGB200 NVL72 & all frontier AICold plates + plumbing; new failure modes, leaks
Immersion~200+ kW/rackEmerging / nicheHighest density; dielectric fluid, serviceability cost

Liquid cooling does double duty: it both lifts the density ceiling and cuts PUE (water carries ~3,000× the heat per volume of air, slashing fan and chiller energy). That is why it shifted from optional to mandatory in one GPU generation — and why a discrete, fast-growing cooling supply chain (Vertiv, nVent, cold-plate and CDU makers) became a picks-and-shovels play on the buildout.

06 — Investment Implications

A New Layer of Investable Names

This is the lesson's payoff. The power constraint creates value-capture outside your core seven (TSMC, NVDA, AVGO, SK Hynix, ASML, AMD, Cadence/Synopsys) — a parallel set of beneficiaries most chip-focused investors underweight.

Power generation & utilities
The clearest beneficiaries of secured-power scarcity: independent power producers and nuclear operators with existing, dispatchable capacity (e.g. Constellation, Vistra, Talen, NRG). Their moat is a physical asset — generation already connected to the grid — that no amount of capex can fast-track. Watch PPA signings with hyperscalers as the signal.
Electrical & grid equipment
Transformers, switchgear, UPS, and gas turbines are themselves backlogged. Eaton, Schneider, Vertiv, and GE Vernova sell into both the datacenter and the grid upgrades needed to feed it — a double exposure. Long backlogs = visible, durable revenue, but watch for a buildout-cycle peak.
Cooling supply chain
Liquid cooling went from optional to mandatory in one GPU generation (Section 05). Vertiv, nVent, and cold-plate/CDU specialists ride content-per-rack growth that scales with GPU density, not just unit count. The risk: lower barriers to entry than silicon — margins could compress as competitors pile in.
Nuclear & the SMR option
A 45 GW offtake pipeline makes nuclear the marquee long-dated bet, plus uranium/fuel (Cameco) as the upstream. But separate the proven (restarts, large reactors) from the speculative (SMRs — mostly pre-revenue, pre-regulatory-approval, late-decade at best). Don't pay deployed-cashflow multiples for option value.

How this connects to your thesis: the power layer doesn't replace the core seven — it gates them. If utilities and grid equipment can't deliver megawatts, NVIDIA can't ship the GPUs into operation, which caps realized (vs. ordered) demand. So power-buildout data (interconnection approvals, PPA volume, turbine backlogs) is now a leading indicator for your existing positions — not just a separate trade. Add a power-availability falsifier to the demand side of the NVDA/AVGO thesis.

Synthesis — Follow One Dollar of Capex (power view)

Where the Energy Dollar Lands

Your recurring "follow one dollar of hyperscaler capex" trace now has a power branch. Of the ~$750B the 14 largest operators are spending in 2026 (up from ~$450B), a large and growing share never touches a GPU: at ~$45–55B per GW, the campus shell, substation, transformers, switchgear, cooling, and the power contract itself absorb the dollar before silicon does. [BNEF capex] The bottleneck (power), the captured name (utilities + electrical equipment), and the leading indicator (interconnection + PPA volume) are now explicit hops in the trace.

Primary Source

Go Deeper

Read first: SemiAnalysis — "AI Datacenter Energy Dilemma: Race for AI Datacenter Space" — the highest-signal cost-and-power model of the buildout, from the newsletter in RESOURCES.md. Then: the IEA's 2025 datacenter electricity update for the macro demand/grid picture, and BloombergNEF's buildout tracker for capex and $/GW figures.

Comprehension Check

Quiz — 5 Questions

Select the best answer for each.

1. A datacenter with PUE = 1.5 means that for every watt of compute, the facility draws:

Half a watt of total power for cooling and conversion losses
One and a half watts total — half a watt of overhead per compute watt
Fifteen watts total because of the cooling multiplier applied
Exactly one watt since modern facilities reach the ideal value

2. Why has liquid cooling become mandatory for racks like the GB200 NVL72?

It is cheaper to install than air handling units in every case
Regulators banned air-cooled datacenters in most US states recently
Die heat flux now exceeds what moving air can physically remove
Liquid cooling lets the chips run at far higher clock frequencies

3. In 2025, the binding constraint on AI capacity growth has primarily become:

GPU supply, since foundries cannot make enough advanced silicon
Securing delivered electricity, given multi-year grid interconnect queues
HBM memory, which remains the single hardest part to produce
Software talent able to write the kernels that run the models

4. The most defensible scarce asset created by the power constraint is:

A large order allocation for the newest generation of GPUs
A patent portfolio covering the latest liquid-cooling cold plates
A contracted, grid-connected, shovel-ready supply of megawatts
A long-term lease on land located near a major fiber backbone

5. How does the power layer relate to your core seven holdings (e.g. NVDA)?

It replaces them, since utilities now capture most of the value
It is unrelated, as power and silicon trade on separate cycles
It only matters for training, never for the inference side of demand
It gates them — power delivery caps realized versus ordered demand
From your instructor: The chain to remember — Dennard's end made chips denser and hotter; density became rack power (~120 kW); rack power became cluster power (300–500 MW); cluster power hit a grid that takes a decade to connect; so secured power is the new moat. PUE tells you how much of a procured megawatt you actually get to sell; $/GW tells you who can afford to play; energy/token ties it back to inference margin. Ask me anything: how PPAs and behind-the-meter generation work, whether SMRs are real or hype, or how to add a power-availability falsifier to the NVDA/AVGO thesis.