LatticeLog

  • Governance
  • Infrastructure
  • Markets
  • Analysis
  • Notes
  • Signals
  • Cycles
  • Learnings
  • Commentaries
  • Glossary
  • Home

Written by Nithinraj Kooneri

in Bifrost Systems
Cooling & Thermal Management — Fenrir Research
Bifrost Systems/Build/Cooling & Thermal Management
Fenrir Research · Bifrost Systems · Build / 04

Cooling & Thermal Management: The Other Half of the Build

Every watt of power that enters an AI data centre leaves as heat. Getting it out has quietly become a $6 billion sub-sector on its way to $27 billion — and, in the American Southwest, the thing that decides whether a data centre gets built at all.
Fenrir Research  ·  Jul 2026  ·  Yggdrasil Ledger / latticelog.in

Any apprentice could raise a fire. It was the master smith who understood the quenching — that the blade was not made in the heat, but in how carefully the heat was drawn back out of it. The forge that could not cool its work made nothing but slag.

Original epigraph, in the register of Tolkien’s forge- and quenching-verses
Section 01

The Heat Problem

The previous pieces were about getting power into the data centre. This one is about the law of physics that follows immediately: nearly every watt that goes in comes back out as heat, and it has to go somewhere. For most of computing history that was a trivial afterthought — a few fans, some cold air. AI broke that assumption in a single hardware generation.

The driver is rack density. A traditional server rack draws 5 to 15 kilowatts, and air cooling handles it comfortably. But an NVIDIA Blackwell GB200 rack draws around 140 kW — roughly ten times as much heat, in the same physical footprint. Air simply cannot carry that much energy away fast enough; at those densities, chips throttle their own clock speeds to avoid cooking, and the expensive GPUs you paid for stop delivering the performance you bought. Next-generation designs push toward 900 kW per rack and single chips near 2,000 watts. The industry’s own thermal standards body now recommends liquid cooling above 20 kW per rack — a line virtually every serious AI deployment crossed in 2025.

Why This Is Infrastructure, Not Plumbing

At modern AI densities, cooling is no longer a feature of the building. It is the building.

Once a rack crosses roughly 50–140 kW, the cooling architecture dictates the physical design of the entire facility — the floor loading, the piping, the power draw, the water supply, even the site selection. You do not build a data centre and then cool it; you design the cooling and wrap a data centre around it. That inversion is what turned thermal management from a maintenance line item into a distinct, investable layer of the AI build-out.

Section 02

The Technologies — and How They Compare

The shift underway is from moving air to moving liquid, because liquid carries heat up to a thousand times more effectively. There are three broad approaches, in ascending order of density and complexity:

  • Direct-to-chip (cold plate). A metal plate sits directly on the GPU, and coolant flows through it, carrying heat away at the source. It retrofits into existing racks with minimal disruption, which is why it dominates today — more than half the market.
  • Immersion cooling. Entire servers are submerged in a non-conductive dielectric fluid. It handles the highest densities and slashes water use, but requires purpose-built tanks and specialised fluids — a bigger commitment.
  • Air (the baseline). Still the incumbent for lower-density workloads, increasingly supplemented by rear-door heat exchangers as a transitional step.

The efficiency gap is stark, and it is measured in PUE — power usage effectiveness, the ratio of total facility power to the power actually reaching the computers. A PUE of 1.0 is the theoretical ideal (no overhead); everything above it is energy spent on cooling and losses.

Cooling Efficiency by Method (PUE — Lower Is Better)
Approximate power usage effectiveness for GPU-dense AI clusters. Sources: industry TCO analyses (2026), ASHRAE guidance. Two-phase immersion approaches ~1.03; traditional air containment runs 1.5–1.8 — meaning air can waste 50–80% as much power again on top of the compute itself.
MethodRack densityTypical PUEWater useBest fit
Air (containment)Up to ~15 kW1.5–1.8High (evaporative)Legacy / low-density
Rear-door exchanger~20–40 kW1.35–1.55ModerateTransitional retrofits
Direct-to-chip~40–140 kW1.15–1.30Varies by heat-rejectMainstream AI (today)
Single-phase immersion100 kW+1.03–1.0890–98% lessMax density / water-scarce

The likely future is not one winner but a dual track: direct-to-chip serving the mainstream because it retrofits easily, immersion taking the ultra-dense and water-constrained deployments. Both are liquid; the air era is ending for anything running AI.

Section 03

The Market — and the Business Model Underneath It

Cooling has crossed its inflection point. Liquid cooling penetration was around 3% in 2021; by 2026 it is roughly 37% — a more-than-tenfold jump in five years, driven by hardware that leaves no choice.

Liquid Cooling Market, 2026
~$6 bn
Up from ~$4.8bn in 2025
Forecast by 2035
~$27 bn
~18% CAGR — a near-6x expansion
Liquid Penetration
3% → 37%
2021 to 2026 — the inflection
Rack Density, 2026
+69% YoY
Average jumped to ~27 kW; Blackwell racks hit ~140 kW
Data-Centre Liquid Cooling Market ($bn)
Global data-centre liquid cooling market size, US$bn. Source: industry market research (GMInsights and others, 2026). ~18% CAGR to 2035; cold-plate solutions are the largest segment, immersion the fastest-growing at the top end.
Analyst Read — The Annuity Is in the Services, Not the Boxes

Today the market is roughly 70–80% hardware (the cooling systems themselves) and 20–30% services. Industry forecasts expect that mix to invert over the next four to five years, toward services-led revenue — monitoring, maintenance, fluid management, thermal-as-a-service. That shift matters more than the headline growth rate: hardware sales are cyclical and competitive, but the recurring service contract attached to a mission-critical cooling loop is an annuity. In the primer’s language, it is the difference between a one-off transaction and fee-bearing, recurring revenue — and the latter is what re-rates a business.

Section 04

The Water Tradeoff — Where Build Meets Strain

Here is the catch that turns a growth story into a constraint. The cheapest way to reject heat is evaporative cooling — letting water evaporate to carry heat away — and it is thirsty. A conventional evaporatively-cooled data centre consumes on the order of 2 to 5 million gallons of water per megawatt, per year. Scale that across a gigawatt-class AI campus and the number becomes a genuine claim on a regional water supply.

In the American Southwest — Phoenix, Las Vegas, much of Texas — water rights are finite and increasingly contested, and local permitting reviews now scrutinise a data centre’s water draw as closely as its power draw. Water, in other words, has joined the interconnection queue as a gate that can stop a project before it starts. This is where the cooling choice becomes a siting decision, and where the Build thread runs straight into the Strain thread.

The escape is a genuine three-way tradeoff, with no free option:

  • Evaporative cooling: lowest energy, highest water — fine where water is cheap, disqualifying where it isn’t.
  • Dry / adiabatic coolers: minimal water, but higher energy use — you trade the water bill for the power bill.
  • Immersion: cuts water use 90–98% and improves efficiency — but demands specialised fluid and purpose-built design.
Connects to: The Power-Compute Nexus (the power that becomes this heat) · Resource Adequacy: Water (the scarcity this collides with) · Colocation & the Bypass Economy (where water joins power as a siting gate).
Section 05

Reading It Through the Frameworks

Where is the moat? Not in the commodity hardware — cold plates and pumps will commoditise. It sits in three places: proprietary thermal IP and the coolant-distribution systems that are hard to replicate; the retrofit lock-in of a cooling loop that, once installed, is expensive to swap; and the service annuity attached to it. Own the recurring relationship, not the box.

Where does policy become the cash flow? In two ways. Water-permitting rules increasingly mandate low-water cooling in scarce regions, effectively legislating demand for immersion and dry cooling. And PFAS regulation cuts the other way — the phase-out of certain fluorochemical fluids used in two-phase immersion creates a real transition risk for that specific technology, and a cost advantage for single-phase and fluid-free approaches.

Thermal-Systems Specialists
Riding the inflection
Makers of CDUs, cold plates and immersion systems sit directly in the 3%→37% penetration wave — the clearest picks-and-shovels exposure.
Power & Cooling Integrators
Services annuity
Firms that bundle power distribution with thermal management capture the recurring service revenue as the mix shifts toward services.
Immersion & Dielectric Fluids
Growth with PFAS risk
Highest density and lowest water, but two-phase fluids face regulatory phase-out — single-phase and next-gen fluids are the safer exposure.
Water-Efficient Heat Rejection
Permitting tailwind
Dry and adiabatic coolers benefit directly as water-scarce regions legislate against evaporative cooling.
Retrofit & Services
The annuity layer
Monitoring, maintenance and fluid management on mission-critical loops — the recurring revenue that outlasts any hardware cycle.
Legacy Air-Only Vendors
On the wrong side
Suppliers without a liquid pathway face structural decline as AI densities make air cooling unviable.
The Bull Case
Adoption is mandatory, not optional — Blackwell-class hardware cannot be air-cooled
Penetration inflected (3%→37%) with a long runway; market near-6x by 2035
Revenue mix shifting toward recurring services — annuity economics
Water-permitting rules legislate demand for low-water cooling
The Risks
Hardware commoditisation compresses margins on the boxes themselves
PFAS phase-out is a specific, live risk to two-phase immersion fluids
Adoption is tethered to the AI capex cycle — a build slowdown hits cooling too
Retrofitting the vast installed air-cooled base is slow and costly
Bottom Line

Cooling is the half of the AI build-out that the power headlines skip, and it has quietly become mandatory infrastructure: at Blackwell densities, air cooling simply doesn’t work, so liquid is not a choice but a requirement. That has turned a maintenance line item into a sub-sector growing toward $27 billion, with the most durable value in the recurring service loop rather than the hardware.

But cooling is also where the build meets its limits. Every megawatt of AI compute is a claim on water as well as power, and in the places the data centres most want to be, water is exactly what’s scarce. The winners will be the ones who master the quenching — who reject heat without draining a river — and sell that capability as an annuity, not a box.

The great forges were not built beside the richest ore, nor the strongest fire. They were built beside cold, running water — for the masters knew that what a forge could make was limited, in the end, only by how well it could be cooled.

Original epigraph, in the register of Tolkien’s forge-verses
Bifrost Systems · Build Thread
← Previous
The Interconnection Queue
The grid bottleneck the whole thread returns to
Next →
Second-Life Infrastructure
Retired plants reborn — where connection rights outvalue the asset
Sources & Notes
Market sizing & penetration: GMInsights, Persistence Market Research, MarketsandMarkets data-centre liquid cooling reports (2026). Technology & PUE comparisons: industry TCO analyses (Adam Silva Consulting, energy-solutions.co, 2026); ASHRAE TC 9.9 2026 Thermal Guidelines (Class H1, direct-to-chip recommendation above 20 kW/rack). Hardware densities: NVIDIA GB200 / Blackwell specifications; industry reporting on rack-density trends. Water consumption: industry facility analyses (evaporative cooling ~2–5 million gallons/MW/year) and US Southwest permitting coverage. Figures are the most recent available as of publication and, for market forecasts, are third-party estimates that vary between sources. All framing and conclusions are Fenrir Research’s own.
This analysis is for informational purposes only. Not investment advice. Company and product references are illustrative of sector dynamics, not recommendations. Fenrir Research is a division of Yggdrasil Ledger (latticelog.in).
←The Interconnection Queue
Second-Life Infrastructure→

Comments

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

More posts

  • The Health Case That Closes

    July 28, 2026
  • The Border Adjustment Problem

    July 28, 2026
  • The Young Fleet

    July 28, 2026
  • Committed Emissions

    July 28, 2026

LatticeLog

Structural research across markets, infrastructure, climate, and the systems that connect them. Published under Fenrir Research, a division of Yggdrasil Ledger.

  • Blog
  • About
  • FAQs
  • Authors

Twenty Twenty-Five

Designed with WordPress