It's 02:47 on a Wednesday. A 3 MVA parallel-redundant UPS system running an AI training workload trips one leg. The customer's night crew is qualified but has never worked live on this specific frame. They page us at 03:04. By 03:45 our two engineers are on the floor with pre-staged capacitors and a factory-verified transfer procedure. Here is exactly what happened over the next 96 minutes — and why the load never blinked.
12 minutes to confirm — no premature action
The first thing we did on arrival was not touch the equipment. We spent 12 minutes reading the event log, cross-referencing the transfer switch health, and confirming the surviving leg was carrying full load within its rated margin. On a redundant system, the biggest risk in a rescue is a well-intentioned move that drops the second leg.
The failed module had a capacitor bank fault — not the more dangerous DC-bus failure the initial alarm suggested. That distinction changed the whole procedure.
- Event log read: 3 warnings preceding the trip
- Surviving leg load: 68% of rated — safe for isolation work
- Transfer switch: healthy, tested in AUTO mode
- Battery string voltage on failed leg: nominal
Live-load isolation without an ATS transfer
Because the surviving leg had headroom, we did NOT force a bypass to raw utility. Instead we electrically isolated the failed module on the DC bus, mechanically locked it out, and replaced the capacitor bank on the bench with the customer's own service engineer watching every step.
This kept the IT load on double conversion the entire time. Any bypass to raw utility would have been a short but real dip in power quality — for an AI training run mid-batch, that means a checkpoint rollback and potentially hours of lost compute.
Slow, deliberate resync
The temptation on a rescue is to close things out fast. We didn't. Once the module was ready to reintegrate, we ramped it in at 10% load steps over 24 minutes, verifying phase sync, harmonic distortion, and battery charge current at every step. Sign-off happened at 04:23 with the customer's own facility manager on the line.
“Their engineer walked in reading the event log before he even said hello. That's when I knew we were in good hands.”
The playbook here isn't heroism — it's the discipline to read the system before touching it, and the equipment relationships to have the right spare on the truck at 03:04 in the morning. That's what a real 24/7 maintenance contract buys you.
