All insights

Physical Infrastructure Buildout, week ending 19 July 2026

Stratum Atlas · Special report · 19 July 2026 · Edition 2 evidence cut

On 16 July 2026 China’s Moonshot AI launched Kimi K3. Blind rankings put the open weight model first on Arena’s Frontend Code Arena over Claude Fable 5 and GPT 5.6 Sol. Moonshot’s own tables place it with or above leading US closed systems on several coding and agentic benches. It is the largest open weight model printed to date. Weights and the technical report land 27 July. That is the news people are reading. This note asks what that competitive print means for the AI-driven infrastructure buildout once you leave the leaderboard and meet the physical valves: compute packages, campus power, and grid long leads.

Evidence cut, blank until 27 July

This note uses The Stratum Terminal editions dated 18 July 2026 and Moonshot’s 16 July Kimi K3 launch disclosures. The Kimi K3 technical report and open weights are scheduled for 27 July 2026. Absolute K3 training compute is blank until then. Until those print, we read K3 through the 16 July disclosures and Kimi K2 as the disclosed proxy.

The call can change after 27 July. If the technical report shows a sharply smaller absolute training job than the K2 proxy implies, or if buyers mark packages, campuses, or turbine reservations against that print, we will say so in Edition 3. Backlog into 2027 is what is already booked and what still stalls delivery. It is not a promise every gigawatt and every wafer start converts.

The call, as of 19 July

Provisional call: Kimi K3 is a training recipe and serving architecture print. On current evidence it is not a physical clearance print. The AI compute, campus, and power co minima on The Stratum Terminal still bind. The book is intention heavy and sold out. Efficiency can change the intensity of future training and inference demand, and it can contribute to cancellations and deferrals inside a backlog. It does not, by itself, shorten a 128 week transformer lead or open non priority CoWoS.

The 2027 path is not locked. The physical valves have not cleared on 18 July evidence, and the launch blog alone is not enough to mark them cleared.

Reading K3 through the K2 proxy

Moonshot chose K2 as the efficiency denominator. That makes K2 the right proxy before 27 July. K2’s numbers bound what K3 likely is. They are not K3’s numbers.

What K2 disclosed

Moonshot’s Kimi K2 technical report (arXiv:2507.20534) prints a Mixture of Experts model with about 1 trillion total parameters and 32 billion activated, 8 of 384 experts per token plus a shared expert, pre-training on 15.5 trillion tokens with MuonClip and zero loss spike, on a cluster of NVIDIA H800 GPUs (8 GPUs per node, NVLink and NVSwitch inside the node). Sparse MoE, huge token budget, export controlled class accelerators, multi node high bandwidth fabric.

What K3 disclosed on 16 July

Kimi K3 prints 2.8 trillion total parameters, 16 of 896 experts under Stable LatentMoE, Kimi Delta Attention and Attention Residuals, quantization aware training from SFT with MXFP4 weights and MXFP8 activations, and serving on supernodes of 64 or more accelerators. About 2.5× overall scaling efficiency versus K2. Overall performance still trails Claude Fable 5 and GPT 5.6 Sol.

AttnRes paper (arXiv:2603.15031): Block AttnRes matches the loss of a baseline trained with 1.25× more compute on scaling runs and on a Kimi Linear 48B / 3B activated model trained on 1.4 trillion tokens. That supports the direction of the efficiency claim. It is not a K3 FLOP budget.

Proxy reading

Hold K2 fixed and apply the K3 launch facts:

  1. Larger total model. 2.8T versus about 1T.

  2. Different sparsity. 16 of 896 versus 8 of 384. Active FLOPs per token versus dense 2.8T fall. Memory residency for serving does not. Moonshot’s ≥64 accelerator supernode note is the serving floor they printed.

  3. Better conversion versus K2. The 2.5× is relative: more capability per unit of training compute than the K2 recipe. It is not a print that K3 used 40 percent of K2’s tokens, GPUs, or joules.

  4. Same lab, same optimizer family. Per Head Muon extends the Muon line K2 scaled with MuonClip.

  5. Hardware in evals, not in training. H200 and alternative GPGPU in kernel sandboxes; H20 in benchmark footnotes. The K3 training cluster SKU is blank.

Proxy conclusion before 27 July: K3 is a bigger, sparser, more efficient open frontier MoE from the same stack that trained K2 on 15.5T tokens and H800s. Absolute training compute for K3 is unknown. It could be lower than a naive 2.8T scaling of K2. It could still be a very large job.

Edition 3 recomputes the proxy when Moonshot prints token count, FLOPs or equivalent, cluster size and SKU mix, duration and energy if disclosed, and the decomposition of the 2.5× versus K2.

Launch stack, short

Scale and sparsity. First open model Moonshot places in the 3T class. Stable LatentMoE activates 16 of 896 experts. Serving still needs large high bandwidth domains.

KDA. Hybrid linear attention for long sequence flow. Prefill cache work toward vLLM. No K3 production FLOPs table at 1M context in the launch blog.

AttnRes. Depth wise selective residuals; Block AttnRes for trainable memory cost. Paper print: 1.25× compute matching versus baseline.

Quantization and routing. MXFP4 / MXFP8 from SFT. Quantile Balancing, balanced expert parallel training, Per Head Muon, SiTU, Gated MLA.

Commercial serving. API at $0.30 / $3.00 / $15.00 per million tokens (cache hit input / cache miss input / output). Mooncake cited for cache hit rates above 90 percent on coding workloads. Reasoning on at max effort at launch.

The physical book

The Terminal is global. Geography is tagged on datapoints. For AI driven power and campus delivery, the thickest backlog prints are US utility and OEM books. For compute, the co minima sit in Taiwan, Japan, Korea, and US designer allocation.

May 2026 global semiconductor sales printed $120.6 billion, up 104 percent. TSMC raised 2026 capex to $60 billion to $64 billion with Q2 gross margin 67.7 percent. ASML raised 2026 sales to EUR 43 billion to EUR 45 billion. NVIDIA Data Center revenue on the last print is $75.2 billion, up 92 percent. The AI compute co minimum still passes as a set.

[SC-L2-00 · AI compute stall register / co minimum]

Campus and power intention sit one layer upstream. Dominion Virginia holds 70 GW of large load requests and 51 GW under contract against a 24.7 GW system peak. Behind the meter, GE Vernova prints about 100 GW under contract against 3 to 5 GW shipped per quarter; Siemens Energy Gas Services prints 24 GW data center related inside 87 GW of commitments.

[MH-L1-01 · Campus demand anchors vs system peak]

[PG-L1-01 · Queue intention and heavy duty gas books

Facility co minimum on the campus and grid boards remains large load interconnection, large power transformers near 128 weeks, GSUs above 160 weeks, sole US GOES at Cleveland Cliffs Butler with the DLA IDIQ through 2030, and heavy duty gas turbine slots. Semiconductor Circularity Ratio about 0.18 to 0.22 on roughly $42 billion of disclosed NVIDIA equity into OpenAI, Anthropic, and CoreWeave. Campus circularity concentrated around Oracle’s $638 billion FY2026 RPO.

Still binding: CoWoS, ABF, T glass, HBM, NVIDIA allocation; interconnect, LPT, GSU, GOES, gas slots.

Not binding on our stall test: switchgear in the mid 40 week range, utility scale batteries, multi vendor campus Ethernet and optics, UPS and diesel as schedule stretch, water after dry cooling and closed loop, multi supplier WFE.

Kimi K3 is tested against the first list.

Backlog through 2027 is not a cancellation shield

The edition findings say the AI driven book is real, intention heavy, and sold out into 2027 on packaging, HBM, gas slots, and transformer leads. That is a delivery statement. It is not a claim every announced campus, reserved turbine, or non priority package buyer takes delivery.

Backlog can cancel. Take or pay and prepaid structures make some of it sticky. Circularity and vendor equity make some of it fragile. Historical interconnection completion near 13 percent already shows intention dying in process. Goldman’s campus path carries about 60 percent on time materialization on scheduled additions.

If the 27 July technical report, or the market reaction to open weights, hits AI spend hard enough:

  1. Watch conversion and cancellations first — Dominion dated load, GE Vernova path to ≥110 GW, CoWoS and HBM sold out prints, NVIDIA Data Center sequential revenue.

  2. Watch lead times second — LPT, GSU, CoWoS weeks ease when factories and second sources ship, not when a blog posts 2.5×.

  3. Keep financing separate — open weight price pressure hits the circular buyer channel before it hits GOES tons.

On 19 July evidence, do not underwrite a reduction in physical co minima from K3’s launch blog. Do underwrite that Edition 3 may cut intention and capture persistence if 27 July plus buyer prints show the spend path breaking.

What K3 does and does not move

Training efficiency versus sold out books. If the K2 proxy and the 2.5× claim hold in spirit, Moonshot gets more capability per training FLOP than last generation. That does not, by itself, cancel CoWoS through 2026, HBM into 2027, NVIDIA’s 50 to 60 percent CoWoS claim, or an ABF plant in 2032. Those clear when leads collapse, unsold capacity is disclosed, or orders are cancelled or deferred in primary prints.

Call 1. Through year end 2026, K3’s 16 July disclosures do not falsify the AI compute co minimum. Revisit after 27 July if absolute training compute and linked order relief print together.

Serving and open weights. Sixteen of 896 experts cut active parameters per token versus dense 2.8T. Moonshot still wants ≥64 accelerator supernodes and still sells accelerator backed tokens. Open weights on 27 July can widen who hosts. Hosting still burns accelerators, HBM, power, and cooling.

Call 2. K3 serving guidance is consistent with continued accelerator and power demand. It is not evidence of campus load cancellation today.

Contracts and lead times. Dominion ESAs, GE Vernova reservations, Siemens data center related gas, Oracle RPO, and hyperscaler capex guides are not contingent on Moonshot’s calendar. They can still be renegotiated, deferred, or cancelled. Physical conversion remains interconnect, transformers, GOES, gas slots, and package allocation.

Call 3. Through 2027, facility co minima remain interconnect + LPT + GSU + GOES + gas slots, with NVIDIA allocation on IT fill, unless Edition 3 shows those books breaking.

Financing. If open weights compress closed model pricing, watch Circularity and buyer funding first. That is not the same board as Cliffs GOES or 52 to 78 week CoWoS.

Call 4. Price software pressure and physical clearance on separate layers.

If 27 July confirms the efficiency story

Branch A — Lower absolute training compute than the K2 proxy implies

Token count and FLOPs that net well below a size adjusted K2 baseline after the 2.5× claim, or energy and cluster size that make the training job clearly smaller in absolute terms.

Then: mark future training intensity down; do not automatically clear CoWoS, HBM, LPT, or GOES; raise the probability of softer incremental training cluster orders in 2027 and 2028; confirm in buyer prints before cutting capture scores.

Branch B — Still enormous absolute job

Size, long context, always on reasoning, and multimodal keep absolute compute huge even if conversion versus K2 improved.

Then: efficiency watchlist stays on; clearance call unchanged; open weights matter more for inference proliferation than training relief; physical co minima stay the spine through 2027 on current falsifiers.

Branch C — Open weights plus buyer behavior show spend cancellation or deferral

Needs counterparty evidence, not only Moonshot’s PDF: cancelled or deferred campus phases, GE Vernova or Siemens book cuts, CoWoS or HBM allocation opening, hyperscaler capex guide cuts, Dominion contracted GW stalling.

Then: cut intention and persistence before lead time PASS stalls; recut Circularity; physical leads may stay long even as new orders slow.

Branch D — Report is thin and absolute compute stays blank

Edition 3 keeps today’s provisional call, adds the CN Kimi intention flag on the Terminal, and leaves the efficiency clock armed.

Public calls and how we lose

Terminal evidence: Stratum Atlas editions dated 18 July 2026. Primary model sources: Moonshot AI, “Kimi K3: Open Frontier Intelligence,” 16 July 2026, https://www.kimi.com/blog/kimi-k3 ; Kimi API K3 docs; Kimi Team, “Attention Residuals,” arXiv:2603.15031; Kimi Team, “Kimi K2: Open Agentic Intelligence,” arXiv:2507.20534 (proxy baseline).

Atlas Memo: https://www.thestratumatlas.com/research/memo

Request access

Atlas terminal access is opening in stages.

Atlas is opening in stages to allocators, lenders, and development teams actively underwriting power infrastructure. Tell us what you are working on and we will show you the research that answers it.