Everything you've learned — compute (L8), HBM (L12), and power (L20) — collapses into one number you can defend to a skeptic: the cost to serve a million tokens. Build it from the metal up, then watch it decide build-vs-buy and who captures the margin.
This is the lesson where the course becomes operational. You've assembled the whole stack; now you compress it into the master metric of the inference business: cost per million tokens ($/Mtok). Every API price, every "our model is cheaper" claim, every build-vs-buy decision, and the entire margin structure of AI services reduces to this number versus the price charged for it. The skill: build it bottom-up so you can spot which inputs actually move it — and therefore which companies capture the value.
Core thesis: Serving cost is cost-per-GPU-hour ÷ tokens-per-hour. The numerator is set by hardware capex + power (L20); the denominator is set by throughput, which is gated by memory bandwidth and HBM capacity (L8, L12). The gap between this cost and the API sell price is the Model-as-a-Service (MaaS) gross margin — enormous for frontier models, near-zero for commodity ones. The same math sets the build-vs-buy line (own ASICs/GPUs vs rent) — and that line is exactly where Broadcom's ASIC case, AMD's TCO wedge, and NVIDIA's pricing power are won or lost.
Strip away the noise. The cost to generate tokens is just a rate of spending divided by a rate of producing:
The art is in those two inputs, and each maps directly to a prior lesson:
Move the sliders. The left column is the rent path (cloud price, all-in). The right is the own/build path, assembled from capex + power exactly as in L20. Watch what happens to $/Mtok, to gross margin against your sell price, and to the build-vs-buy verdict.
Note how small the power line is per token at these prices (a few cents per Mtok) yet how it dominates the aggregate bill at gigawatt scale — that's the L20 paradox: negligible per token, decisive per cluster.
| Lever | Pulls $/Mtok via… | Owned by (thesis name) |
|---|---|---|
| Throughput ↑ | Bigger denominator: continuous batching, FP8/FP4 quant, MoE sparsity, speculative decoding | NVDA (software), the model lab |
| HBM bandwidth/capacity ↑ | Lifts the memory-bound ceiling (L8) & fits the model in fewer GPUs | SK Hynix; AMD (capacity wedge) |
| $/GPU-hour ↓ | Cheaper capex/power per hour: own vs rent, ASICs, lower PUE | AVGO (ASIC), utilities/cooling (L20) |
The headline trend: GPT-4-equivalent serving fell from ~$20/Mtok in late 2022 to ~$0.40/Mtok by 2026, and API prices dropped ~80% in a single year. [Epoch AI] That deflation is all three levers compounding — and it's why the volume of tokens matters as much as the margin per token.
The calculator's verdict line is the heart of the investment case. Owning (GPUs or custom ASICs) trades a big fixed capex for a low marginal cost; renting is pure variable cost. The cross-over depends on utilization and workload stability:
Diligence reframe: When a company claims an inference cost advantage, decompose it into the three levers and ask which one and is it durable. A throughput edge from a software trick (batching) is copyable in months; a structural edge from owning cheaper silicon + power at high utilization is not. That distinction is the difference between a feature and a moat.
The MaaS gross margin is (sell price − serve cost) ÷ sell price. At a $0.52/Mtok serve cost and a $15/Mtok frontier output price, that's a ~97% gross margin — extraordinary, and the reason model labs and clouds raced into the business. But two forces compress it:
Tie to your thesis: this is the quantitative spine of the whole portfolio. NVDA's gross margin is a line item in every renter's serve cost — durable only while throughput-per-dollar leadership holds. AVGO's ASIC TAM is exactly the set of workloads where the build-vs-buy line tips to "own." AMD's wedge is the memory-bound segment where capacity, not topology, decides $/Mtok. And the inference price deflation is the demand engine: cheaper tokens → more tokens → more silicon and power. Add to THESIS.md a per-name note: which lever of $/Mtok does this name own, and how copyable is it?
Read first: Introl — "Inference Unit Economics: The True Cost Per Million Tokens" for the full bottom-up build. Then: Epoch AI's inference price-trend data for the deflation curve, and SemiAnalysis (in RESOURCES.md) for serve-cost teardowns by GPU generation.
Select the best answer for each.
1. The cost to serve a million tokens is fundamentally:
2. For token-by-token decoding, throughput is usually capped by:
3. A custom ASIC tends to beat renting GPUs specifically when the workload is:
4. An inference cost edge is most likely a durable moat when it comes from:
5. As inference prices fell from ~$20 to ~$0.40 per Mtok, the effect on hardware demand was: