HKT's AI DCI Superhighway linking Lok Ma Chau / Hetao with Tseung Kwan O has quietly changed the reliability picture in Hong Kong. Two sites that used to be independent failure domains now behave as one fabric for AI training and inference workloads.
When GPU jobs span halls, a cooling or power event at one endpoint can strand utilisation at the other. The interconnect itself becomes a critical asset — with its own hierarchy, criticality and FMEA — yet most facility models still stop at the meet-me room.
The implication for reliability engineering: dependency maps must cross sites. FMEA on the DCI transport layer belongs in the same asset strategy as FMEA on chilled water at either endpoint.
For bank tenants under HKMA OR-2, this is not academic. Severe-but-plausible scenarios now include DCI segment loss, and the tolerance evidence has to reflect that. A board pack that treats Site A and Site B as independent when the workload does not will fail scrutiny.
Practical starting point: extend the technical hierarchy across both sites for the shared fabric; run a joint ACR on interconnect and endpoint plant; write scenarios that kill one site while the other is under maintenance.
Cluster fabric is a competitive advantage for AI capacity. It is also a shared failure domain. Model it as one — or discover it as one during an incident.
- Treat DCI as a critical asset class, not cabling.
- Cross-site dependency maps are mandatory for AI fabric.
- OR-2 scenarios must include interconnect loss where relevant.
- Joint ACR across endpoints beats two independent hall models.
Want this analysis applied to your own facility?
Book a fixed-scope reliability assessment — one hall or one campus, four to eight weeks.
Book an assessment →