Every major operator in dense metro markets has now announced liquid-cooled capacity. The differentiator is not the loop — it is whether the operator has a failure-mode library for it.
Coolant Distribution Units, manifolds, rear-door heat exchangers and direct-to-chip cold plates are new to most ops teams. Their failure signatures do not map cleanly onto legacy CRAH-based playbooks. A leak on a secondary loop is not the same event as a CRAH fan failure; the blast radius, detection latency and remediation path are different.
A defensible answer looks like this: a refreshed technical hierarchy that includes the new equipment; criticality re-rated because a CDU sits upstream of tens of racks; and FMEA-derived strategies that spell out inspection frequency, leak-response, and hot-standby posture.
Instrumentation matters as much as topology. Leak detection at every quick-disconnect, conductivity trending on the secondary loop, and differential-pressure alarms tuned to the rack manifold — not only the CDU — are the minimum we now insist on before GPU energisation.
Get this wrong and the first coolant incident becomes both an availability event and a compliance event. Boards and bank tenants will not accept 'we followed the vendor commissioning pack' as a post-incident narrative.
Treat liquid cooling as a reliability programme, not a facilities upgrade. The capital is the easy part. The failure library is the hard part — and the part that protects uptime.
- Refresh hierarchy and criticality when CDUs and manifolds enter the hall.
- Build FMEA for leak, isolation, and ride-through — not just nameplate MTBF.
- Instrument manifolds and dielectric quality, not only CDU supply temperature.
- Commissioning packs from GCs rarely include the reliability envelope you need.
