Data centre monitoring

The method and system for monitoring cooling units in data centers address the complexity of fault identification by outputting specific alerts based on anomaly flags, enhancing operational efficiency by directing skilled personnel to the correct issues.

GB2701835APending Publication Date: 2026-05-13EKKOSENSE LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
GB · GB
Patent Type
Applications
Current Assignee / Owner
EKKOSENSE LTD
Filing Date
2024-10-02
Publication Date
2026-05-13

AI Technical Summary

Technical Problem

The complexity of cooling systems in data centers leads to difficulties in identifying the type of fault causing temperature set point deviations, necessitating different skilled operatives for various failures, resulting in inefficiencies and unnecessary costs.

Method used

A method and system for monitoring cooling units by receiving measured fluid flow temperatures, comparing to set points, and setting anomaly flags to output alerts indicating specific faults, such as machine, fluid flow, or cooling anomalies, based on combined flag statuses.

Benefits of technology

Enables efficient identification and correction of cooling unit anomalies, ensuring resources with the correct skills address the issues promptly, optimizing data center operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The invention relates to a method of monitoring a data centre 100 to detect cooling system anomalies. Example embodiments include a computer-implemented method of monitoring a data centre 100 comprisi
Need to check novelty before this filing date? Find Prior Art

Description

Field of the Invention The invention relates to monitoring a data centre to detect cooling system anomalies. Background Modern data centres perform a central function in enabling operation of services connected over the internet. A key requirement is that a data centre must have a high degree of availability. To maintain operation, it is important to keep each item of equipment in a data centre within its operational parameters. A key operational parameter is temperature. Cooling is therefore a central feature of data centre design. A typical design will involve a cooling system comprising one or more air handling units (AHUs), one or more liquid cooling distribution units (CDUs) or a combination of both, the cooling system operating to remove heat from multiple rows of equipment racks typically arranged along aisles. For CDUs, a cooling liquid such as water or oil is distributed through pipework, providing a definite connection between each individual CDU and one or more computer racks to which the cooling liquid is directed to remove heat from heat exchangers, which are directly coupled to the equipment in the computer racks. For CDUs the temperature and flow rate of the cooling liquid can be varied by reducing or increasing the flow rates to a particular item of equipment. Additionally, each individual CDU transfers its recovered heat typically to a central cooling system, and consequently can also be affected by the performance and control set points at the central plant level where flow rates and supply and return temperatures can be adjusted. For AHUs, cooling air is typically forced downwards into an underfloor plenum chamber where it is distributed via floor-mounted grilles to computer racks where the cooling air is directed to the front of the rack to remove heat from the internal equipment with fans. For AHUs, the temperature and flow rate of the cooling air can be varied by reducing or increasing the fan speed and control set-point of each AHU. Additionally, the placement and adjustment (openness) of the floor-mounted grilles can also be used to increase or decrease cooling to a particular rack or item of equipment and consequently affect its temperature. Additionally, each individual AHU transfers its recovered heat typically to a central cooling system, and consequently can also be affected by the performance and control set points at the central plant level where flow rates and supply and return temperatures can be adjusted. A general problem with cooling systems for data centres is of increasing complexity, due to the use of many cooling systems serving a common area where computer equipment racks are located. Each cooling system will have a temperature set point defining a desired temperature output. Problems can arise when cooling systems are not operating correctly, which can result in set points not being correctly maintained. Due to different types of machinery involved, it may not be clear from simple monitoring what type of fault might cause a set point not to be maintained. A failure may for example result from an electrical equipment failure, such as a fan, or alternatively may result from a cooling system failure, either at an individual cooling unit level or at a more general central cooling system level. Different types of failures will require differently skilled and equipped operatives to be called to address the problem. Not knowing what type of failure might be involved can therefore cause complications and unnecessary costs. Summary of the Invention In accordance with the invention there is provided a computer-implemented method of monitoring a data centre comprising a plurality of cooling units, the method comprising for each of the plurality of cooling units: i) receiving a measured fluid flow temperature from the cooling unit; ii) comparing a temperature set point to the measured fluid flow temperature to determine a temperature difference, AT; iii) setting a first cooling anomaly flag if a AT is greater than a positive first threshold; iv) setting a second cooling anomaly flag if a AT is greater than a negative second threshold; and v) outputting an alert indicating an anomaly in the cooling unit if either of the first or second cooling anomaly flags is set. The method may further comprise for each cooling unit: iia) receiving a measured fluid flow through the cooling unit and a cooling load of the cooling unit; iib) setting a fluid flow flag dependent on comparing the measured fluid flow to a fluid flow threshold; and iic) setting a cooling load flag dependent on comparing the cooling load to a cooling load threshold, wherein outputting the alert indicates a type of fault in the cooling unit dependent on a combination of a status of the first and second cooling anomaly flags, a status of the fluid flow flag and a status of the cooling load flag. The alert may indicate: a machine fault in the cooling unit if the fluid flow flag and cooling load flag are both set; a fluid flow anomaly in the cooling unit if the fluid flow flag is set and the cooling load flag is not set; or a cooling fault in the cooling unit if the cooling load flag is set and the fluid flow flag is not set. The method may further comprise, for each cooling unit having a chilled water supply: iid) receiving a measure of chilled water flow through the cooling unit; and iie) setting a chilled water flag dependent on comparing the measure of chilled water flow to a chilled water threshold, wherein the alert indicating the type of fault in the cooling unit is further dependent on a status of the chilled water flag. The plurality of cooling units may comprise one or more air handling units configured to provide a cooling air flow and / or one or more liquid cooling distribution units configured to provide a cooling liquid flow. The liquid may comprise water or oil, i.e. be either water- or oil-based. Outputting the alert may comprise providing a visual indication on the cooling unit in a computer model representation of the data centre. The alert may indicate that the cooling unit is overloaded if the first cooling anomaly flag is set or is underloaded if the second cooling anomaly flag is set. According to a second aspect there is provided a computer system for monitoring a data centre comprising a plurality of cooling units, the computer system configured, for each of the plurality of cooling units, to: i) receive a measured fluid flow temperature from the cooling unit; ii) compare a temperature set point to the measured fluid flow temperature to determine a temperature difference, AT; iii) set a first cooling anomaly flag if a AT is greater than a positive first threshold; iv) set a second cooling anomaly flag if a AT is greater than a negative second threshold; and v) output an alert indicating an anomaly in the cooling unit if either of the first or second cooling anomaly flags is set. The computer system may be further configured, for each cooling unit, to: iia) receive a measured fluid flow through the cooling unit and a cooling load of the cooling unit; iib) set an fluid flow flag dependent on comparing the measured fluid flow to a fluid flow threshold; and iic) set a cooling load flag dependent on comparing the cooling load to a cooling load threshold, wherein the output alert indicates a type of fault in the cooling unit dependent on a combination of a status of the first and second cooling anomaly flags, a status of the fluid flow flag and a status of the cooling load flag. The alert may indicate: a machine fault in the cooling unit if the fluid flow flag and cooling load flag are both set; a fluid flow anomaly in the cooling unit if the fluid flow flag is set and the cooling load flag is not set; or a cooling fault in the cooling unit if the cooling load flag is set and the fluid flow flag is not set. The computer system may be further configured, for each cooling unit having a chilled water supply, to: iid) receive a measure of chilled water flow through the cooling unit; and iie) set a chilled water flag dependent on comparing the measure of chilled water flow to a chilled water threshold, wherein the output alert indicating the type of fault in the cooling unit is further dependent on a status of the chilled water flag. The plurality of cooling units may comprise one or more air handling units configured to provide a cooling air flow and / or one or more liquid cooling distribution units configured to provide a cooling liquid flow. The liquid may comprise water or oil. The output alert may comprise providing a visual indication on the cooling unit in a computer model representation of the data centre. The output alert may indicate that the cooling unit is overloaded if the first cooling anomaly flag is set or is underloaded if the second cooling anomaly flag is set. According to a third aspect there is provided a data centre comprising a plurality of cooling units and a computer system according to the second aspect. According to a fourth aspect there is provided a computer program comprising instructions to cause a computer system configured to monitor a data centre to perform the method according to the first aspect. The computer program may be embodied on a non-transitory computer-readable storage medium, which may be a physical computer readable medium, such as a disc or a memory device, or may be embodied as a transient signal. Such a transient signal may be a network download, including an internet download. These and other aspects of the invention will be apparent from, and elucidated with reference to, the embodiments described hereinafter. Detailed Description The invention is described in further detail below by way of example and with reference to the accompanying drawings, in which: Figure 1 is a simplified schematic plan view drawing of an example data centre having a cooling system with multiple systems serving multiple equipment racks; Figure 2 is a schematic plot of temperature over time, illustrating different types of temperature set point breaches; Figure 3 is a screenshot output from a computer model presentation view of an example data centre; Figure 4 is an example cooling system summary table presented by the computer model of Figure 3; Figures 5a and 5b illustrate a flow diagram of an example algorithm for monitoring for anomalies in a cooling system; Figures 6a to 6c are screenshot extracts from a computer model presentation view of a cooling system in an example data centre; and Figure 7 is a three-dimensional computer model representation of an example data centre having a hybrid cooling system with multiple air and water cooling units. Figure 1 illustrates schematically a plan view of an example simplified layout of a data centre 100, in which equipment racks are arranged in rows 101-103 within a ventilated room 105, with ventilation supplied in this example by AHUs 104a-g and liquid cooling provided to one row 103 supplied by a CDU 104h. Each AHU provides a cooled air supply to the room 105 via an underfloor plenum, with air outlets 106 provided adjacent the equipment racks 101-103. The air outlets 106 direct cooling air from the underfloor plenum to locations where cooling is required. The location of the air outlets 106 may be selected and modified according to cooling requirements in the room 105. The CDU 104h provides a cooled liquid supply to the equipment racks in row 103 via a chilled water manifold 109. To support the required cooling, each of the cooling units 104a-h are operated according to specified cooling requirements, which may vary according to the position of each cooling unit. In general, each cooling unit will be provided with a set point temperature that determines the required outlet temperature of cooling fluid (i.e. air or water) that is provided by the unit. A monitoring unit 107a-h is installed on each of the cooling units 104a-h. Each monitoring unit 107a-g is arranged to monitor and / or calculate various operating parameters of the respective cooling unit. These parameters include: return and supply fluid flow temperatures; a measure of fluid flow provided by each cooling unit; and a calculation of cooling load for the cooling unit. The airflow for each AHU may for example be calculated by each monitoring unit 107a-g based on a measured fan speed and a nominal airflow for the AHU. Each monitoring unit 107a-h is in communication with a local computer 108 via a wired or wireless connection, for example via an Ethernet or WiFi connection. The local computer 108 comprises an input / output interface 110 that is configured to communicate with each of the monitoring units 107a-h to receive measurements, for example wirelessly via an antenna 113 and / or via a wired connection. The computer 108 further comprises a processor 111 for processing the received measurements and a memory 112 for storing the received measurements, together with instructions for performing the operations required. The computer 108 may transmit the measurements and other data to a remote computer system, for example a computer in another data centre, for processing, analysis and presentation of measurement data from the data centre 100. Alternatively, the local computer 108 may itself process, analyse and present the measurement data using the processor 111 and memory 112. The local computer 108 may be part of the data centre 100, i.e. a computer in one of the equipment racks. Table 1 below provides a summary of variables that may be used in determining cooling anomalies for an example AHU, together with an example value or type. The type of variable indicates whether the variable is either read from the AHU, calculated from a reading or is a property of the AHU. The examples provide typical values and their units of measurement or the possible options for the type, for example whether control for the AHU is on the return or supply airflow, the type of cooling the AHU operates according to, and the type of fan control in the AHU. Possible cooling types include the following: • DX - Direct expansion refrigeration (refrigerant-based cooling); • CW - Chilled water (in which a central chiller does the cooling and its cooling is typically chilled water circulating in pipes driven by pumps); • CWDX - Chilled water DX (refrigerant based cooling where the condenser from the refrigerant circuit is cooled with circulating water driven by pumps); • FADX - Free Air DX (in which an outside air cooling system switches across to refrigerant based cooling in hot weather); • FACW - Free Air Chilled Water (as for FADX but fed from the central chiller during hot weather); and • FA - Free Air (using outside air for cooling). Types of fans include: • Fixed (fans that can only operate at one speed and have no ability to be varied); • Fixed Dynamic (fans that can be adjusted but are fixed at a specific speed and will not adjust to operating conditions); and • Fixed Variable (fans that can ramp up or down according to a control algorithm). The reference setpoint for the AHU is the setpoint that is used to determine any cooling anomalies, i.e. the value from which a deviation will tend to indicate an anomaly. The expected setpoint, or xSP, is a value offered to an AHU and may be updated if the reference setpoint is updated. The ATe threshold is an error value for the unit or room and may be calculated or a default value provided. This threshold may be a global variable, i.e. the same for each AHU throughout the data centre, or alternatively each 5 AHU may have its own threshold. Table 2 below indicates corresponding variables for an example liquid cooling distribution units (CDU). Table 1 - Variables used for an example air handling unit (AHU) cooling system. Input Type Example Return Temperature Reading 24.5°C Supply Temperature Reading 18.2°C Airflow Calculated 3.24 m3 / s Cooling Uoad Calculated 24.7 kWc Nominal Cooling Uoad Property 60 kWc Nominal Airflow Property 4.0 m3 / s Nominal CW flow rate Property 1.2 1 / s Control On Property Return / Supply Reference Setpoint Property 24.0 Dynamic Setpoint (xSP) Property 24.5 ATe Threshold Property / Calculated 1.4 Cooling Type Property DX, CW, CWDX, FADX, FACW, FA Fan Control Property Fixed, Fixed Dynamic, Fixed VariableW Year Installed Property 2008 10 Table 2 - Variables used for an example liquid cooling distribution unit (CDU) cooling system. Input Type Example Return CW Temperature Reading 30.1°C Supply CW Temperature Reading 15.2°C Liquid flow Calculated 3.73 litres / s Cooling Load Calculated 202.0 kWc Nominal Cooling Load Property 250.0 kWc Nominal liquid flow rate Property 4.8 litres / s Control On Property Return / Supply Reference Setpoint Property 15.0 Dynamic Setpoint (xSP) Property 15.5 ATe Threshold Property / Calculated 2.0 Pump Control Property Fixed, Fixed Dynamic, Fixed Variable Year Installed Property 2023 The fault-finding algorithms described herein enable the data centre to be monitored for 5 set points of multiple cooling units. In a general aspect, if a measured fluid flow temperature from one or more cooling unit deviates from a set point by more than a defined threshold, a check can be made as to whether the temperature is above or below the set point and an alert provided accordingly. As an example, in general terms a fan alert for an AHU would require an electrician while a cooling system alert would require 10 a plumber, due to the different hardware involved. If a measured temperature deviates from a set point for multiple cooling units but by less than would trigger an alert for a single unit, a further type of alert may be issued that requires action on a common cooling system, which may require a different type of problem to be addressed than with the individual cooling units. The methods described herein thereby enable different 15 types of faults to be detected and alerts issued accordingly, which enables more efficient operation of the data centre due to being able to identify and correct anomalies more quickly so that resources with the correct skills and tools can focus on correcting the anomalies. Figure 2 illustrates schematically an example temperature plot over time for an cooling unit outlet, showing a set point and various defined temperature thresholds. The measured outlet temperature 201 fluctuates around a setpoint temperature Ts during normal operation in a first time period 202. First and second upper temperature thresholds Ti+, T2+ are defined above the setpoint temperature Ts and first and second lower temperature thresholds Tf, T2 are defined below the setpoint temperature Ts. A positive first temperature difference threshold ATi is defined by the difference between the second upper temperature threshold T2+ and the setpoint Ts. A negative second temperature difference threshold AT2 is defined by the difference between the second lower temperature threshold T2 and the setpoint Ts. A positive third temperature difference threshold AT3 is defined by the difference between the first upper temperature threshold Ti+ and the setpoint Ts. A fourth negative temperature difference threshold AT4 is defined by the difference between the first lower temperature threshold Tf and the setpoint Ts. Figure 2 also illustrates a first breach 203 of the second upper temperature threshold T2+ as a difference between the measured temperature and the setpoint Ts becomes greater than the positive first threshold ATi and a second breach 204 of the second lower temperature threshold T2 as a difference between the measured temperature and the setpoint becomes greater than the negative second threshold AT2. These breaches 203, 204 can be used as part of an algorithm to determine a type of fault in the AHU. Figure 3 is an example screenshot from a computer model representation of a data centre 300. As with the schematic example in Figure 1, the data centre 300 comprises a plurality of cooling units (in this case AHUs) 304a-d that provide cooling to rows of equipment racks 301-303 in an enclosed ventilated room 305. A summary table 310 provided by the computer model lists each of the cooling units 304a-d in the data centre 300, together with their status, type, utilisation, expected setpoint (xSP) and cooling anomaly setpoint. An example summary table for a plurality of cooling units providing cooling to a data centre is provided in Table 3 below. Each cooling unit is identified by number, or alternatively may be provided with identifying names. The status of each unit is indicated, i.e. whether the unit is either ON or OFF. The type of unit indicates the type of cooling, which in this example is either DX or CW (cooling water) for the AHUs and CDU for the liquid cooling unit. Control On indicates whether the unit is controlled according to the return (or output) fluid flow or the supply (or input) fluid flow. The Utilization indication is a measure of the percentage utilization of the unit, which is based on the mean cooling load as a percentage of the unit’s nominal cooling load. The 5 expected setpoint, or xSP, is either the setpoint entered into the unit, if available, or is a calculated value if the actual setpoint is not available. The CA Setpoint is a setpoint that is used in the computer model for determining cooling anomalies and may be either predetermined based on the expected setpoint or may be adjusted manually in the computer model summary table 310. 10 Table 3 - Example summary table for monitored cooling units in a data centre AHU Identifier Status Type Control On Utilization xSP CA Setpoint Last Update AHU 1 ON DX Return 0-25 25.0 25.0 ll / Oct / 23 AHU 2 ON DX Return 25-50 26.0 25.0 ll / Oct / 23 AHU 3 ON DX Return OFF - 25.0 ll / Oct / 23 AHU 4 ON DX Return 75-100 26.0 25.0 (A) ll / Oct / 23 AHU 5 ON CW Supply 50-75 18.0 18.0 23 / Sep / 23 AHU 6 ON CW Supply 50-75 19.0 18.0 23 / Sep / 23 AHU 7 OFF DX Return 25-50 25.0 Never CDU 1 ON CDU Supply 25-50 15.5 15.0 (A) ll / Oct / 23 Within the computer model presentation shown in Figure 3, each cooling unit may be selectable from which a further summary table may be provided, an example of which is shown in Figure 4. This summary table, in this case for an AHU, indicates the basic 15 information for the particular selected unit, including current cooling capacity, return and supply temperature readings, airflow reading, cooling utilization and current anomaly status. In this case, an anomaly is detected indicating a machine underload, which is output based on calculations of the variables for the AHU. Figures 5a and 5b illustrate an example algorithm for determining anomalies in a 20 cooling unit. The algorithm may be run for each cooling unit in a data centre, in which each unit is monitored and any flags set during running of the algorithm used to determine a type of alert to be output. Figure 5a illustrates a first part of the algorithm in which a particular AHU is selected. The process starts at step 501 and a check is made at step 502 as to whether the AHU is on or off. If the AHU is off, the process stops at step 503 and proceeds to the next AHU, i.e. starting again at step 501 for the next AHU. At step 504 a check is made regarding the cooling load, or duty, of the AHU compared to a cooling load threshold. If the cooling load is equal to or above the cooling load threshold, a high duty flag is set at step 505. This high duty flag, if set, may then be added to an output alert. The process then proceeds to a setpoint test 506, illustrated in Figure 5b. In the setpoint test 506, the airflow temperature from the AHU is compared to a threshold defined by high and low setpoints, which are in this case at +1.5°C and -1.5°C from the temperature set point for the AHU. If the airflow temperature is between these two values, the process proceeds to step 507 and ends, optionally with an alert output indicating a high duty flag if this was previously set at step 505. If at step 506 the measured airflow temperature is below the low threshold, a low anomaly flag is set at step 508, which determines the type of alert to be output at the end of the process. One or more further tests may then be performed, which are dependent on the type of AHU being tested and the level of detail required in the output alert. The outputs from the further tests may provide further diagnostic information that can enable the reason for the low anomaly flag to be determined. At step 509 an airflow test is performed, where the measured airflow for the AHU is compared to a first airflow threshold, which in this example is equal to 60% of a nominal airflow for the AHU. If the measured airflow is greater than or equal to the first airflow threshold, an airflow anomaly flag is set at step 510. The airflow anomaly flag indicates whether an airflow fault is detected. In summary, if the airflow utilization (i.e. the fan speed equivalent) is high but the AHU is unable to maintain the temperature setpoint, the airflow should be lower and this indicates a problem. If a calculated airflow is available, the logic test in step 509 is as follows: If (calculatedAirflow / nominalAirflow) >airflowHighThreshold then cAnomalyAirflow = 1 (else 0) If a calculated airflow is not available, the fan speed can be used as a proxy for airflow, and the logic test in step 509 may instead be: If (fanSpeed / 100) >airflowHighThreshold then cAnomalyAirflow = 1 (else 0) At step 511a cooling load test is performed, where the cooling load for the AHU is compared to a first cooling load threshold, which in this example is set to be 20% of a nominal cooling load for the AHU. If the cooling load is greater than or equal to the first cooling load threshold, a cooling load flag is set at step 512. The cooling load flag indicates whether a cooling fault has been detected. In summary, if the cooling load is higher than the first cooling load threshold but the AHU is unable to maintain setpoint, then the unit should be providing more cooling. The fact that the AHU is not indicates a cooling fault. If a calculated airflow is available for the AHU, the logic test in step 511 can be expressed as: If (CoolingUoad / nominalCooling) >MAX (calculatedAirflow / nominalAirflow, coolingHighThreshold) then cAnomalyCooling = 1 (else 0) If the calculated airflow is not available, the fan speed may be used as a proxy, making the logic test in step 511: If (CoolingUoad / nominalCooling) >MAX (fanSpeed / 100, coolingHighThreshold) then cAnomalyCooling = 1 (else 0) At step 513 a chilled water test is performed, which may be applicable if the AHU is of a type that uses chilled water. A chilled water flow for the AHU is compared to a first chilled water flow threshold, which in this example is set to 30% of a nominal chilled water flow rate. If the measured chilled water flow rate is greater than or equal to the first chilled water flow threshold, a chilled water flag is set at step 514. The chilled water flag indicates whether a chilled water valve is modulating water flow rate as expected. In summary, if the chilled water flowrate is higher than the threshold but the AHU is unable to maintain the temperature setpoint, the shilled water flowrate should be lower and so is indicative of a problem. The logic test in step 513 may be expressed as: If (calculatedCWflow / nominalCWflow) >CWflowHighThreshold then cAnomalyCWflow = 1 (else 0) Following the tests in steps 509, 511 and 513, which may be carried out in any order, the process ends at step 507 and an alert is output that is dependent on a combination of a status of the low anomaly flag (i.e. whether the flag is set or not set) and a status of flags set as a result of the airflow, cooling load and optionally chilled water tests. If at step 506 the measured airflow temperature is above the high threshold, a high anomaly flag is set at step 515, which determines the type of alert to be output at the end of the process. One or more further tests may then be performed, which are again dependent on the type of AHU being tested and the level of detail required in the output alert. The outputs from the further tests may provide further diagnostic information that enables the reason for the high anomaly flag being set. At step 516 an airflow test is performed, where the measured airflow for the AHU is compared to a second airflow threshold, which in this example is equal to 80% of a nominal airflow for the AHU. If the measured airflow is less than or equal to the second airflow threshold, an airflow anomaly flag is set at step 517. In summary, if the airflow utilization (i.e. the fan speed equivalent) is low but the AHU is unable to maintain the temperature setpoint, the airflow should be higher and this indicates a problem. If a calculated airflow is available, the logic test in step 516 is as follows: If (calculatedAirflow / nominalAirflow) <airflowUowThreshold then cAnomalyAirflow = 1 (else 0) If a calculated airflow is not available, the fan speed can be used as a proxy for airflow, and the logic test in step 516 may instead be: If (fanSpeed / 100) <airflowUowThreshold then cAnomalyAirflow = 1 (else 0) At step 518 a cooling load test is performed, where the cooling load for the AHU is compared to a second cooling load threshold, which in this example is set to be 80% of a nominal cooling load for the AHU. If the cooling load is less than or equal to the second cooling load threshold, a cooling load flag is set at step 519. In summary, if the cooling load is lower than the threshold but the AHU is unable to maintain setpoint, then the unit should be providing more cooling. The fact that the AHU is not indicates a cooling fault. If a calculated airflow is available, the logic test in step 516 may be expressed as: If (CoolingLoad / nominalCooling) <MIN (calculatedAirflow / nominalAirflow, coolingLowThreshold) then cAnomalyCooling = 1 (else 0) If a calculated airflow is not available, the fan speed may be used as a proxy and the logic test in step 516 expressed as: If (CoolingLoad / nominalCooling) <MIN (fanSpeed / 100, coolingLowThreshold) then cAnomalyCooling = 1 (else 0) At step 520 a chilled water test is performed, which is applicable if the AHU is of a type that uses chilled water. A chilled water flow for the AHU is compared to a second chilled water flow threshold, which in this example is set to 70% of a nominal chilled water flow rate. If the measured chilled water flow rate is less than or equal to the chilled water flow threshold, a chilled water flag is set at step 521. In summary, if the chilled water flowrate is lower than the threshold but the AHU is unable to maintain the temperature setpoint, the shilled water flowrate should be higher and so this is indicative of a problem. The logic test in step 520 may be expressed as: If (calculatedCWflow / nominalCWflow) <CWflowLowThreshold then cAnomalyCWflow = 1 (else 0) Following the tests in steps 516, 518 and 520, which may be carried out in any order, the process ends at step 507 and an alert is output that is dependent on a combination of the high anomaly flag with the outcome (i.e. whether a flag is set or not set) of the airflow, cooling load and optionally chilled water tests. The cooling anomaly type that is dependent on the combination of flags from the various tests may be displayed to the user on the computer model representation of the AHU, as illustrated in Figures 6a-6c. An icon 601 is provided on the AHU 604 in Figure 6a indicating the status of the AHU, which in this example indicates no cooling anomaly, i.e. the AHU 604 is operating normally. If a cooling anomaly is determined for the AHU 604, as shown in Figure 6b, the icon 602 may change to indicate the cooling anomaly, for example by changing colour. The icon 603 may also represent differently to indicate that the AHU 604 is not operating, as shown in Figure 6c. The cooling anomaly status may also be indicated within the readings for a particular AHU, as shown in Figure 4. The alert output by the method may therefore take different forms and may be for example in the form of a visual representation within a computer model of the data centre. As a result of the process illustrated in Figures 5a and 5b and described above, an overall cooling anomaly status may be determined from the combination of anomalies identified, i.e. the combinations of flags that are set or not set. Table 4 below summarises the overall cooling anomaly, i.e. the type of output alert provided, in terms of the combinations of anomaly flags that are set. The cooling water anomaly may be used only if the AHU is of the CW type. Not all possible permutations are indicated, for example where both the high and low threshold anomalies are 1, since such a situation should be impossible. In the case of the high anomaly flag being set, a machine overload alert is output if the positive first threshold is exceeded and no other flags are set. This indicates that the unit is overloaded and is unable to maintain set point control. In the case of a low anomaly flag being set, a machine underload alert is output if the negative second threshold is exceeded and no other flags are set. If only one of the cooling load and airflow flags is set, this indicates a cooling issue or airflow issue respectively. If both the cooling load and airflow flags are set, this indicates a machine fault. If, in the case of a CW AHU, the chilled water flag is set, either with neither of the cooling load and airflow anomaly flags being set or with only the airflow anomaly flag being set, this indicates a cooling water issue. In a general aspect therefore, the type of alert output is determined by a combination of the temperature threshold flag with one or more flags being set from an airflow test, a cooling load test and optionally a chilled water test. Table 4- Overall cooling anomaly output dependent on flags set for different tests (l=flag set, 0=flag not set). High Threshold Anomaly Flag Low Threshold Anomaly Flag Cooling Load Anomaly Flag Airflow Anomaly Flag Chilled Water Anomaly Flag Overall Cooling Anomaly 1 0 0 0 0 Machine overloaded 1 0 0 1 0 Airflow issue 1 0 1 0 0 Cooling issue 1 0 1 1 0 Machine fault 1 0 0 0 1 CW flow issue 1 0 0 1 1 CW flow issue 1 0 1 0 1 Cooling issue 1 0 1 1 1 Machine fault 0 1 0 0 0 Machine underloaded 0 1 0 1 0 Airflow issue 0 1 1 0 0 Cooling issue 0 1 1 1 0 Machine fault 0 1 0 0 1 CW flow issue 0 1 0 1 1 CW flow issue 0 1 1 0 1 Cooling issue 0 1 1 1 1 Machine fault 0 0 0 0 0 OK Table 5 below provides a list of possible cooling anomalies, the definition of each anomaly and their corresponding possible causes. Based on the possible causes, the identification of a particular cooling anomaly can enable a notification to be provided 5 that directs service personnel to address the cause of the anomaly. Table 5 - Cooling Anomaly Definitions and Causes Overall Cooling Anomaly Definition Causes Airflow issue The airflow of the cooling unit is too low to meet the cooling demand of the unit Failed fan unit, blockage to airflow path at air supply or air return, failed air filter Airflow issue (low anomaly) The airflow of the cooling unit is too high for the demand required The fan speed is stuck at a high level, or the minimum fan speed is too high for the required cooling load. CW flow issue The chilled water flow of the unit is too low to meet the cooling demand of the unit The chilled water valve is closed or is stuck in a position that restricts water flow. CW flow issue (low anomaly) The chilled water flow of the unit is too high for the required demand The chilled water valve is stuck open or in a position that allows too much water flow. Cooling issue The unit is not providing enough cooling to meet demand The refrigeration circuit is inefficient (refrigerant gas leak, or low level), faulty compressor, heat exchanger not working efficiently Cooling issue (low anomaly) The unit is providing too much cooling to the room The refrigeration circuit is stuck in operation mode. Machine fault The unit is not providing enough flow or cooling to meet demand - special case where both an airflow and cooling issue is detected. Unit is faulty, unit has tripped, electrical issue, control temperature stat issue. Total fan failure. Machine fault (low anomaly) The unit is delivering too much cooling and airflow for the required demand Unit is stuck on in full cooling mode. This is a control system fault. Machine underloaded The unit is providing more cooling that is required The unit is recirculating cold air and may not have a lower operational mode so is recommended to be switched off or in hot standby. Machine overloaded The unit cannot meet cooling demand The unit is cooling more than it is designed to (or is capable of). Optimization required to spread cooling load within room. Figure 7 is a three-dimensional computer model representation of an example data centre 700 similar to that illustrated in Figure 1, in which equipment racks are arranged in rows 701a, 701b, 702a, 702b, 703a, 703b within a ventilated room 705. Ventilation 5 is supplied by AHUs 704a-f and liquid cooling is provided to rows 703a, 703b by CDUs 704g, 704h. Other features of the example data centre described above may also be present. Other embodiments are intentionally within the scope of the invention as defined by the appended claims. 10

Claims

1. A computer-implemented method of monitoring a data centre comprising a plurality of cooling units, the method comprising for each of the plurality of cooling units:i) receiving a measured fluid flow temperature from the cooling unit;ii) comparing a temperature set point to the measured fluid flow temperature to determine a temperature difference, AT;iia) receiving a measured fluid flow through the cooling unit and a cooling load of the cooling unit;iib) setting a fluid flow flag dependent on comparing the measured fluid flow to a fluid flow threshold; andiic) setting a cooling load flag dependent on comparing the cooling load to a cooling load threshold,iii) setting a first cooling anomaly flag if a AT is greater than a positive first threshold;iv) setting a second cooling anomaly flag if a AT is greater than a negative second threshold; andv) outputting an alert indicating an anomaly in the cooling unit if either of the first or second cooling anomaly flags is set,wherein outputting the alert indicates a type of fault in the cooling unit dependent on a combination of a status of the first and second cooling anomaly flags, a status of the fluid flow flag and a status of the cooling load flag.

2. The method of claim 1, wherein the alert indicates a machine fault in the cooling unit if the fluid flow flag and cooling load flag are both set.

3. The method of claim 1, wherein the alert indicates a fluid flow anomaly in the cooling unit if the fluid flow flag is set and the cooling load flag is not set.

4. The method of claim 1, wherein the alert indicates a cooling fault in the cooling unit if the cooling load flag is set and the fluid flow flag is not set.

5. The method of claim 1, further comprising for each cooling unit having a chilled water supply:iid) receiving a measure of chilled water flow through the cooling unit; andiie) setting a chilled water flag dependent on comparing the measure of chilled water flow to a chilled water threshold,wherein the alert indicating the type of fault in the cooling unit is further dependent on a status of the chilled water flag.

6. The method of any preceding claim, wherein the plurality of cooling units comprise one or more air handling units configured to provide a cooling air flow.

7. The method of any preceding claim, wherein the plurality of cooling units comprise one or more liquid cooling distribution units configured to provide a cooling liquid flow.

8. The method of claim 7. wherein the liquid comprises water or oil.

9. The method of any preceding claim, wherein outputting the alert comprisesproviding a visual indication on the cooling unit in a computer model representation of the data centre.

10. The method of any preceding claim, wherein the alert indicates that the cooling unit is overloaded if the first cooling anomaly flag is set or is underloaded if the second cooling anomaly flag is set.

11. A computer system for monitoring a data centre comprising a plurality of cooling units, the computer system configured, for each of the plurality of cooling units, to:i) receive a measured fluid flow temperature from the cooling unit;ii) compare a temperature set point to the measured fluid flow temperature to determine a temperature difference, AT;iia) receive a measured fluid flow through the cooling unit and a cooling load of the cooling unit;iib) set an fluid flow flag dependent on comparing the measured fluid flow to a fluid flow threshold; andiic) set a cooling load flag dependent on comparing the cooling load to a cooling load threshold;iii) set a first cooling anomaly flag if a AT is greater than a positive first threshold;iv) set a second cooling anomaly flag if a AT is greater than a negative second threshold; andv) output an alert indicating an anomaly in the cooling unit if either of the first or second cooling anomaly flags is set,wherein the output alert indicates a type of fault in the cooling unit dependent on a combination of a status of the first and second cooling anomaly flags, a status of the fluid flow flag and a status of the cooling load flag.

12. The computer system of claim 11, wherein the alert indicates a machine fault in the cooling unit if the fluid flow flag and cooling load flag are both set.

13. The computer system of claim 11, wherein the alert indicates a fluid flow anomaly in the cooling unit if the fluid flow flag is set and the cooling load flag is not set.

14. The computer system of claim 11, wherein the alert indicates a cooling fault in the cooling unit if the cooling load flag is set and the fluid flow flag is not set.

15. The computer system of claim 11, wherein the computer system is further configured, for each cooling unit having a chilled water supply, to:iid) receive a measure of chilled water flow through the cooling unit; andiie) set a chilled water flag dependent on comparing the measure of chilled water flow to a chilled water threshold,wherein the output alert indicating the type of fault in the cooling unit is further dependent on a status of the chilled water flag.

16. The computer system of any one of claims 11 to 15, wherein the plurality of cooling units comprise one or more air handling units configured to provide a cooling air flow.

17. The computer system of any one of claims 11 to 16, wherein the plurality of cooling units comprise one or more liquid cooling distribution units configured to provide a cooling liquid flow.

18. The computer system of claim 17, wherein the liquid comprises water or oil.

19. The computer system of any one of claims 11 to 18, wherein the output alert5 comprises providing a visual indication on the cooling unit in a computer model representation of the data centre.

20. The computer system of any one of claims 11 to 19, wherein the output alert indicates that the cooling unit is overloaded if the first cooling anomaly flag is set or is 10 underloaded if the second cooling anomaly flag is set.

21. A data centre comprising a plurality of cooling units and a computer system according to any one of claims 11 to 20.15 22. A computer program comprising instructions to cause a computer systemconfigured to monitor a data centre to perform the method according to any one of claims 1 to 10.s