OPERATIONS · SLA

How to Write a Data Centre Maintenance SLA That Actually Covers You

Published September 20269-minute readBack to insights

Most data center maintenance SLAs are vague exactly where it matters most. They promise "rapid response" and "24/7 support" — and the first time a CRAH unit dies on a Sunday night, you find out what those words actually meant. Usually: an acknowledgement email, an engineer "being assigned", and a resolution clock that never started.

A vague SLA is not a neutral document. Every undefined term moves risk from the contractor to you. When the contract says "best efforts", the effort you get is whatever is cheapest that night. When response time means acknowledgement, a text message resets the clock while your hall gets warmer.

This guide is written from the contractor side of the table. We have signed good SLAs and we have been asked to sign bad ones. Here is how to write a data center maintenance SLA that actually covers you — the six elements it must define, the red flags that protect the contractor at your expense, and the questions to ask before you sign.

The six things a data center maintenance SLA must define clearly

Strip a data centre maintenance contract down to its working parts and six definitions decide whether it protects you or the contractor. These are your baseline data center SLA requirements:

1. Response time by priority — in minutes and hours, not adjectives. 2. Resolution time targets — when the system is back to its normal state, not when someone starts looking at it. 3. The preventive maintenance schedule — frequency, scope per visit, and the documentation standard for every task. 4. The escalation path — who is woken up, in what order, after how many minutes without progress. 5. An RCA requirement — a written root cause analysis within 48 hours for every P1. 6. Documentation — every work order, PM record, and incident report filed and accessible to you.

If any of the six is missing or soft, you do not have a maintenance agreement. You have a brochure with a signature page.

Response time vs. resolution time

These two numbers are not interchangeable, and contractors blur them on purpose. Response time is when a qualified engineer engages with the problem — remote diagnosis counts, an automated ticket reply does not. Resolution time is when the affected system is back to its normal, documented state. Both belong in the SLA, per priority level, in minutes and hours.

For a P1 — a critical system failure or imminent risk to live load — the defensible baseline is under 30 minutes on-call response and under two hours on-site. Slower than that and you are carrying the risk window yourself. Faster is available, at a price; the point is that the number is written down.

Here is what the difference costs. A cooling failure at 14:00. The SLA says "four-hour response". The contractor sends an acknowledgement at 14:05 and calls it response met. The engineer arrives at 17:30. Resolution — an undefined term — happens sometime after. Your SLA was technically honoured while your inlet temperatures crossed the ASHRAE limits.

A data center maintenance SLA that promises response without resolution is a paging service, not a maintenance service.

Priority levels: P1, P2, P3 in a data center context

Priorities only work when they are defined in facility terms, not IT-service-desk terms. P1 is a critical system failure or imminent risk: power, cooling, or connectivity affecting live load. P2 is lost redundancy — the system runs, but one more failure becomes a P1. P3 is stable: fully functional infrastructure with minor or preventive scope.

Each priority gets its own response and resolution targets. That sounds obvious. Yet most weak SLAs carry a single response time for everything — which in practice means everything is treated as P3 until someone senior enough starts shouting. Write the matrix: three priorities, two clocks each, no shared row.

Define the priority matrix for your site, not a generic one. A financial-services hall and a development lab do not share a P1. If the contractor's template does not map to your risk register, make them rewrite it — before signature it costs a paragraph; after an incident it costs you the argument.

Scheduled preventive maintenance requirements

Preventive maintenance that is not scheduled — with a defined frequency and scope in the SLA — is optional maintenance. And optional maintenance does not happen on time. It slips a week, then a quarter, and then you find it in the RCA of an incident that should never have occurred.

Your SLA should name the frequency per asset class — UPS quarterly, batteries semi-annual, CRAH filters monthly, whatever your site demands — the scope of each visit, and the documentation standard: what gets recorded, where it is filed, and how you access it. The gap between the preventive maintenance data center operators contract for and what actually lands on the floor is almost always this clause.

Also pin the rescheduling rule: how far a PM may slip, who approves the slip, and how it is reported. Without that, "quarterly" quietly becomes "eventually".

Escalation paths and engineer access

A maintenance SLA that does not define escalation paths is not an SLA. It is a document that protects the contractor, not you.

The escalation path is the ladder an incident climbs when the first response is not enough: named roles, timeboxes per rung, and a duty manager who answers at 03:00. Test it before you sign — call the out-of-hours number on a weekend and see what picks up.

Just as important: engineer access. You want a direct line to the assigned engineer who knows your one-line diagram — not a support desk reading a script. At 2am on a Sunday, the difference between those two is the difference between a ten-minute fix and a P1 that makes the board deck.

Documentation and RCA requirements

Every intervention on your floor should leave a paper trail you own. Work orders for corrective maintenance, PM records for every scheduled visit, incident reports for every alarm — filed to a standard, stored where you can reach them, not locked inside the contractor's ticket system.

For P1 incidents, add the RCA requirement: a written root cause analysis within 48 hours. Not a phone call. Not "we restarted it and monitoring looks green." A document that states what failed, why it failed, what was done, and what changes so it does not fail again. The 48-hour clock matters — memory fades and logs rotate.

If it is not documented, it did not happen. Write that into the SLA and audits become boring — which is exactly what you want.

Red flags: data center maintenance SLA clauses that protect the contractor

When you are evaluating a data center contractor SLA, five phrases tell you who the document was written to protect. Response time defined as acknowledgement — a robot can acknowledge. Exclusions for after-hours emergency callouts — failures do not keep office hours, so your coverage should not either. Resolution time framed as "best efforts" — which is not a target, it is an excuse pre-loaded into the contract. No RCA requirement — meaning every P1 stays a mystery you pay to repeat. And no documentation standard — meaning the only record of work on your floor lives in someone else's system.

One red flag is a negotiation point. Three or more, and you are not buying maintenance — you are buying the appearance of maintenance, at maintenance prices.

FREE DOWNLOAD · PDF GUIDE

Take this guide with you.

The full article as a clean, printable PDF — save it, share it, bring it to your next planning meeting. Drop your work email and it's yours.

One email, one PDF. No newsletter unless you ask for it.

INTEGRATOR SERVICES

The Integrator Services Facility Infrastructure SLA

Integrator Services maintains data centers across EMEA and the Dutch Caribbean under a documented Facility Infrastructure SLA: response and resolution targets per priority (P1: under 30 minutes on-call, under two hours on-site), a preventive maintenance cadence with a defined documentation standard, a written RCA within 48 hours of every P1, and 24/7/365 coverage with direct WhatsApp access to the engineer on your account — no support desk, transparent pricing. Send us your current SLA and we will mark it up line by line — or start from ours.

Review your SLA with us
PRIVACY

We keep this site simple.

We use one first-party cookie to remember your language choice and light analytics (pageviews only, no cross-site tracking) to improve the site. No third-party ad networks, ever.

Read the full privacy note →