How Data Center Teams Catch Cooling Faults Before They Cause Downtime
September 30, 2026
Adapted from Darren Moran’s Clockworks webinar on data center cooling faults.
In a data center, cooling keeps everything else running. When it fails, servers overheat and services go down. A cooling failure at an AWS data center in Northern Virginia caused overheating that disrupted services, including Coinbase.
That kind of outage keeps getting more expensive. A study by Oxford Economics for Splunk put the cost of unplanned downtime for Global 2000 companies at $600 billion a year, up 50% in two years.
Most data center teams already collect the data they need to spot cooling faults early. It sits in the building management system (BMS). In a Clockworks webinar, enterprise account executive Darren Moran showed how fault detection and diagnostics (FDD) checks every cooling system around the clock and tells the team what to fix first.
Key takeaways
- No PM schedule can physically check every system. FDD pulls data every five minutes, so it can run far more checks than a technician can.
- A BMS alarm says something is wrong. A diagnostic says why, down to the sensor or valve.
- Clockworks labels every finding a fault, an opportunity, or a capital project, and scores it for energy, comfort, and maintenance.
- Clockworks connects through the data broker a site has already approved, using MQTT or OPC, over an outbound-only connection.
Darren Moran’s full webinar, “How Data Center Teams Use FDD to Catch Cooling Faults Before They Cause Downtime.” The “watch this moment” links below jump this player to the point being quoted.
Why can’t a PM schedule catch every cooling fault?
As data centers add denser racks and new kinds of cooling, the maintenance work changes too.
“As infrastructure changes, so do maintenance schedules, checks, backup equipment, and specialized technicians.”
Teams still do most of those checks by hand, on a schedule. Darren put it plainly:
“We all have extremely structured PMs, but in reality, physical checks across every single system is not possible in today’s environment based on equipment complexity and the number of checks that are required per system.”
Watch this moment in the webinar — 16:58
When something urgent comes up, the schedule slips. The team’s focus moves to keeping conditions stable so servers don’t shut down.
Clockworks reads BMS data about every five minutes.
“Since we’re pulling data every five minutes, we can carry out far more checks than the traditional PM.”
The checks that pass are useful too. They show which manual checks can be automated or dropped, so technicians go where there’s a real problem.
What’s the difference between a BMS alarm and a diagnostic?
A BMS alarm tells you a system has a problem right now. On a busy site, that means a lot of alarms, “or as we like to call alarm fatigue,” Darren said.
Watch this moment in the webinar — 09:33
Fault detection adds logic and trends on top of the alarm. Diagnostics find the cause. Clockworks runs each building through its analysis engine, which is built on expert systems, “a form of AI that uses engineering logic to determine the diagnostic result.”
Each diagnostic describes the issue and lists the fixes from most likely to least likely, “right down to the sensor and valve level.”
What data center cooling faults does FDD find?
In data centers, Darren said, Clockworks most often finds issues in chiller plants, cooling systems, CRAC units (computer room air conditioners), and air handling units. Every finding is one of three types:
- Fault. Equipment that isn’t running the way it was designed to, like a failed valve, actuator, or sensor.
- Opportunity. A better way to run the equipment. For example, a data hall that isn’t using free cooling from outside air when the temperature and humidity would allow it.
- Capital project. A sensor or piece of equipment worth adding to improve how the system runs.
One chilled water loop in the demo had all three. A pump ran around the clock, even when the chillers were off. That’s a fault. The loop had a low delta T, meaning the supply and return water temperatures were too close together. That’s an opportunity. And the pump had no variable frequency drive (VFD), a capital project Darren said “potentially might lead to all the other opportunities and problems being triggered.” Each finding came with charts the team could check for themselves.
Watch this moment in the webinar — 18:24
How does a finding turn into a fix?
Each diagnostic comes with a short summary written by generative AI. It draws only on Clockworks’ own results. “It’s not looking for external information,” Darren said. The summary names the issue, the likely fix, and the team that should own it. In the demo, that was the controls team.
Watch this moment in the webinar — 20:03
One click turns the diagnostic into a task, with the description and fix steps filled in. The team assigns it to a contractor and syncs it with their work order system.
The sync goes both ways. Close the work order and the task closes in Clockworks, “unless the underlying issue has not been rectified.” If the problem is still there, the diagnostic comes back.
How does Clockworks connect to data center systems?
An attendee asked how Clockworks handles concerns about data connections and hosting.
Clockworks works with any BMS vendor. It reads raw data from the building controllers and sends it over an outbound-only connection to the Azure cloud, where the software runs. In data centers, the usual setup is cloud to cloud, through a data broker the site’s cybersecurity team has already approved. Clockworks reads that data using MQTT or OPC, which most brokers already support.
“All the data centers we are currently operating do have that broker approved on-site before we integrate.”
Watch this moment in the webinar — 29:26
What can data center teams build on top of FDD?
Clockworks engineers help each team build programs around its own KPIs. Darren walked through several, including:
- Condition-based maintenance. Map the checks to the maintenance standard you already use, like SFG20 in Europe, and automate the repetitive ones. It’s the same idea behind predictive maintenance.
- SLA compliance. Set temperature and humidity limits for each data hall, see which halls are drifting, and calculate penalties when a limit is breached.
- PUE. Use power usage effectiveness to find the zones and data halls that are performing poorly, then improve them.
- Continuous commissioning. Bring Clockworks in before handover, so problems found against the design intent get fixed during the defects and liability period.
- BMS alarm fatigue. Pair BMS alarms with FDD, so P1 critical alerts go straight to a work order and P2 and P3 alerts arrive with more context.
Watch this moment in the webinar — 26:09
Why does the data foundation matter for data center teams?
Every finding in this post depends on the data underneath it. Clockworks maps every point to a standardized model of equipment types, relationships, and engineering units. That is how the same engineering checks can run on every chiller, CRAC unit, and air handler, at every site.
Darren explained where that started. Clockworks’ founders came from MIT and “built a global analysis engine on the foundation of expert systems that continues to improve daily with each new portfolio we bring on board.” Today that engine runs 50+ expert systems across more than 650,000 connected assets.
Watch this moment in the webinar — 09:01
Watch the full webinar above. To talk through your own cooling plant, reach out to Darren Moran on LinkedIn or request a demo. Tell him what equipment you have, and he’ll send you the diagnostics that usually trigger on it.
Frequently asked questions
What is FDD for data centers?
Fault detection and diagnostics (FDD) is software that reads the data your building management system already collects and finds equipment problems, along with their causes. In a data center, it runs checks around the clock on chillers, cooling systems, CRAC units, and air handlers. Each finding comes with steps to fix it, ranked from the most likely cause to the least likely.
What data center cooling faults can FDD catch?
FDD finds equipment that isn’t running as designed, like failed valves, actuators, and sensors, or a pump running when there’s no demand. It also finds better ways to run the plant, like using free cooling when outside conditions allow, or fixing a low delta T on a chilled water loop. And it flags equipment worth adding, like a variable frequency drive on a pump.
Can FDD replace preventive maintenance checks?
Not all of them, but it can take many repetitive checks off the schedule. Because FDD reads data every five minutes, it runs far more checks than a technician can do by hand. Teams use the checks that pass to automate or drop manual checks, and send technicians where there’s a real problem.
How does Clockworks connect to data center systems?
Clockworks works with any BMS vendor. In data centers, it most often connects cloud to cloud through a data broker the site has already approved, using MQTT or OPC. Data moves over an outbound-only connection to the Azure cloud, where the software runs.