MAINTENANCE · P1 RESCUE

Live-load UPS failure at a Tier III colocation — 41-minute recovery

A parallel-redundant UPS system tripped one leg mid-week under a busy AI-training load. The customer's own night crew hit the escalation ceiling; a redundancy loss on the surviving leg would have taken down 6 MW.

Back to all case studies

It's 02:47 on a Wednesday. A 3 MVA parallel-redundant UPS system running an AI training workload trips one leg. The customer's night crew is qualified but has never worked live on this specific frame. They page us at 03:04. By 03:45 our two engineers are on the floor with pre-staged capacitors and a factory-verified transfer procedure. Here is exactly what happened over the next 96 minutes — and why the load never blinked.

01 · INITIAL ASSESSMENT

12 minutes to confirm — no premature action

The first thing we did on arrival was not touch the equipment. We spent 12 minutes reading the event log, cross-referencing the transfer switch health, and confirming the surviving leg was carrying full load within its rated margin. On a redundant system, the biggest risk in a rescue is a well-intentioned move that drops the second leg.

The failed module had a capacitor bank fault — not the more dangerous DC-bus failure the initial alarm suggested. That distinction changed the whole procedure.

  • Event log read: 3 warnings preceding the trip
  • Surviving leg load: 68% of rated — safe for isolation work
  • Transfer switch: healthy, tested in AUTO mode
  • Battery string voltage on failed leg: nominal
02 · ISOLATION AND SWAP

Live-load isolation without an ATS transfer

Because the surviving leg had headroom, we did NOT force a bypass to raw utility. Instead we electrically isolated the failed module on the DC bus, mechanically locked it out, and replaced the capacitor bank on the bench with the customer's own service engineer watching every step.

This kept the IT load on double conversion the entire time. Any bypass to raw utility would have been a short but real dip in power quality — for an AI training run mid-batch, that means a checkpoint rollback and potentially hours of lost compute.

03 · COMMISSIONING BACK IN

Slow, deliberate resync

The temptation on a rescue is to close things out fast. We didn't. Once the module was ready to reintegrate, we ramped it in at 10% load steps over 24 minutes, verifying phase sync, harmonic distortion, and battery charge current at every step. Sign-off happened at 04:23 with the customer's own facility manager on the line.

“Their engineer walked in reading the event log before he even said hello. That's when I knew we were in good hands.”
— Facility Operations Manager, cloud provider (EMEA)
TAKEAWAY

The playbook here isn't heroism — it's the discipline to read the system before touching it, and the equipment relationships to have the right spare on the truck at 03:04 in the morning. That's what a real 24/7 maintenance contract buys you.

Bring us the site, the schedule, and the problem.

Start the conversation
OTHER CASES
PRIVACY

We keep this site simple.

We use one first-party cookie to remember your language choice and light analytics (pageviews only, no cross-site tracking) to improve the site. No third-party ad networks, ever.

Read the full privacy note →