A method for dynamically controlling a data center cooling system based on AI

By using an AI dynamic control system to monitor multi-dimensional parameters in real time, identify and regulate local abnormal areas in the data center cooling system, the problem of low cooling efficiency and high energy consumption of traditional cooling systems under non-uniform heat loads is solved, and efficient and reliable cooling management is achieved.

CN121728754BActive Publication Date: 2026-05-01CHONGQING YINGFAN TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHONGQING YINGFAN TECH CO LTD
Filing Date
2026-02-09
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Traditional data center cooling systems face challenges such as reduced cooling efficiency due to the mixing of hot and cold airflows when dealing with non-uniform heat load distribution, difficulty in quickly repairing malfunctions of refrigerant pump energy-saving units, and the impact of external temperature fluctuations on heat dissipation efficiency, leading to energy waste and equipment reliability risks.

Method used

By constructing an AI-based dynamic control system, multi-dimensional parameters are acquired in real time, and equipment with uneven local heating and cooling is identified for differentiated control. Control commands are generated using a preset AI model to balance global energy efficiency with the need to eliminate local hot spots, avoid false alarms due to fluctuations in a single parameter, and achieve precise monitoring and adjustment of equipment status.

Benefits of technology

It effectively solves the problems of non-cooperative competition of cooling equipment, short-circuiting of hot and cold airflows, and high overall energy consumption, improves the responsiveness and reliability of the cooling system, reduces energy waste, and improves the intelligent control efficiency of the cooling system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121728754B_ABST
    Figure CN121728754B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of data center cooling, and particularly relates to a method for dynamically controlling a data center cooling system based on AI, comprising: data acquisition; hot field screening; temperature deviation calculation; risk equipment identification; cooling equipment classification; first regulation amount generation; second regulation amount generation; basic regulation amount generation; and regulation instruction generation. The present application identifies the cooling equipment that is locally uneven in cooling and heat, screens potential risk equipment based on multi-parameter dynamic correlation and historical normal mode to calculate abnormal probability, and generates a global basic regulation amount in combination with a preset AI model to generate a final regulation instruction, so as to ensure intelligent control of the cooling system when coping with non-uniform distribution of loads and non-synergistic operation of single-point temperature control, and effectively solve the problems of non-synergistic competition of cooling equipment, short circuit of hot and cold air flow, coexistence of local overheating and overcooling, and high overall energy consumption caused by non-uniform distribution of load spaces and non-synergistic operation of single-point temperature control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data center cooling technology, and in particular to a method for dynamically controlling a data center cooling system based on AI. Background Technology

[0002] As data centers continue to expand, the cooling demands within the server room become increasingly unevenly distributed. The static, uniform supply model of traditional cooling systems is severely mismatched with the dynamic, non-uniform heat load distribution. This structural contradiction leads to a persistent state of overall undercooling and localized overheating, resulting in significant energy waste and threatening equipment reliability. Therefore, optimizing resource allocation and improving system responsiveness while maintaining overall cooling effectiveness has become a crucial challenge in data center design and operation.

[0003] Chinese Patent Application Publication No. CN121001312A discloses a key cooling system for a green data center. This system includes: a server rack for housing server equipment and generating heat during operation; hot and cold aisles, including a hot aisle connecting the server rack and a cold aisle connecting the cooling unit, the hot aisle being used to concentrate the heat generated by the server rack; and a cooling unit comprising: an indoor cooling unit using full DC inverter technology and environmentally friendly refrigerant; an outdoor heat dissipation unit exchanging heat with the indoor cooling unit via piping; a refrigerant pump energy-saving unit exchanging heat with the indoor cooling unit via piping; a waste heat recovery device connected to the cooling unit via piping for recovering waste heat from the cooling system; and a control system connected to the server rack, cooling unit, refrigerant pump energy-saving unit, and waste heat recovery device. The control system integrates a lightweight edge AI module, which, based on real-time collected server rack temperature, load, and environmental data, predicts the heat load curve and dynamically allocates the operating power of each unit, achieving closed-loop demand-supply regulation.

[0004] Therefore, the key cooling system of the green data center has the following problems: the system's hot and cold aisles rely on a single airflow organization method, which easily leads to the mixing of hot and cold airflows and reduces cooling efficiency; the system's refrigerant pump energy-saving unit relies on pipeline circulation, which makes it difficult to repair quickly in the event of a system failure; the system relies on changes in ambient temperature for heat dissipation, and its heat dissipation efficiency is easily affected by external temperature fluctuations. Summary of the Invention

[0005] To address this, the present invention provides a method for dynamically controlling a data center cooling system based on AI. This method overcomes the problems in the prior art caused by non-uniform load distribution and non-cooperative operation of single-point temperature control, such as non-cooperative power output competition between cooling devices, short-circuiting of hot and cold airflows, coexistence of local overheating and overcooling, and high overall energy consumption. This is achieved by constructing a multi-level identification and differentiated control system based on multi-dimensional parameters.

[0006] To achieve the above objectives, the present invention provides a method for dynamically controlling a data center cooling system based on AI, comprising:

[0007] Real-time acquisition of thermal field imbalance, actual return air temperature, rack load rate, cold aisle static pressure deviation, and intake air temperature of racks in the service area for each cooling device operating based on preset cold water valve opening within the room-level data center.

[0008] Based on the threshold screening results of the thermal field unevenness, several cooling devices of interest are identified, and the relative temperature deviation is determined based on the actual return air temperature of the cooling devices of interest.

[0009] Several risky cooling devices are identified based on the relative temperature deviation, static pressure deviation, and cabinet load rate trends of the cooling devices under concern.

[0010] The risky cooling equipment is classified based on the comparison results of the relative temperature deviation of the risky cooling equipment and the load responsiveness threshold determined by the cabinet load rate, so as to obtain overloaded and underloaded equipment.

[0011] A first control amount is generated based on the thermal field imbalance and the static pressure deviation of the cold aisle of the overload response device within a preset control time period to adjust the opening of the preset chilled water valve, and a second control amount is generated based on the actual return air temperature and the intake air temperature of the underload response device.

[0012] The actual return air temperature and the static pressure deviation of the cold aisle of all the cooling equipment under test within the same preset control time period are input into the preset AI model to obtain the basic control quantity;

[0013] Control commands are generated based on the differences between the basic control quantity and the first and second control quantities, as well as the actual return air temperature.

[0014] Furthermore, the process of identifying several cooling devices of interest based on the threshold screening results of the thermal field unevenness, and determining the relative temperature deviation based on the actual return air temperature of the cooling devices of interest, includes:

[0015] When the thermal field imbalance is greater than a preset balance threshold, the cooling device under test is determined to be the cooling device of interest.

[0016] The relative temperature deviation is determined based on the degree of difference between the actual return air temperature and the preset return air threshold of the cooling equipment in question.

[0017] Furthermore, the process of determining several risky cooling devices based on the relative temperature deviation, static pressure deviation, and cabinet load rate variation trends of the cooling devices of concern includes:

[0018] Based on the relative temperature deviation, static pressure deviation, and cabinet load rate characteristics within a preset time period, the temperature-static pressure correlation, temperature-load correlation, and static pressure-load correlation are determined respectively.

[0019] A correlation feature vector is constructed based on the temperature-static pressure correlation, the temperature-load correlation, and the static pressure-load correlation, and the anomaly probability of the correlation feature vector is determined based on the relative historical anomaly deviation of the correlation feature vector.

[0020] When the anomaly probability is greater than a preset probability threshold, the cooling device of concern is determined to be the risk cooling device.

[0021] Furthermore, the process of determining the temperature-static pressure correlation, temperature-load correlation, and static pressure-load correlation based on the relative temperature deviation, static pressure deviation, and cabinet load rate characteristics within a preset time period includes:

[0022] Calculate the correlation between the normalized relative temperature deviation and the normalized static pressure deviation to determine the temperature-static pressure correlation.

[0023] Calculate the correlation between the normalized relative temperature deviation and the normalized rack load rate to determine the temperature-load correlation.

[0024] The correlation between the normalized static pressure deviation and the normalized cabinet load rate is calculated to determine the static pressure load correlation.

[0025] Furthermore, the process of determining the anomaly probability of the associated feature vector based on its relative historical anomaly deviation includes:

[0026] The associated feature vectors of all the cooling devices under interest within the previous preset normal operating time are recorded as historical associated vectors, and several historical associated distances are determined based on the deviation distance between each historical associated vector and the overall mean vector determined based on the historical associated vectors.

[0027] The current association distance is determined based on the deviation distance between the associated feature vector and the overall mean vector, and the anomaly probability of the associated feature vector is determined based on the relative order of the current association distance among all the historical association distances arranged in ascending order.

[0028] Further, the process of classifying the risky cooling devices into over-response and under-response devices based on a comparison of the relative temperature deviation of the risky cooling devices and the threshold of load responsiveness determined by the rack load rate includes:

[0029] The preset heat transfer coefficient is determined based on the response coefficient of the cabinet load rate and the relative temperature deviation of the risk cooling equipment within a preset historical period.

[0030] The load responsiveness is determined based on the cabinet load rate, the preset heat transfer coefficient, and the relative temperature deviation.

[0031] When the load responsiveness is greater than a preset response threshold, the risk cooling device is determined to be an overloaded device, and when the load responsiveness is less than the preset response threshold, the risk cooling device is determined to be an underloaded device.

[0032] Furthermore, the process of generating a first control quantity to adjust the preset cold water valve opening based on the thermal field imbalance of the overload response device and the static pressure deviation of the cold aisle within a preset control period includes:

[0033] The positive deviation ratio of the thermal field and the positive deviation ratio of the static pressure are determined according to the proportion of the time when the thermal field imbalance and the static pressure deviation of the overloaded device are continuously positive within the preset control time.

[0034] The thermal field deviation and static pressure deviation are determined based on the average degree of continuous positive deviation of the thermal field imbalance and the cold aisle static pressure deviation of the overloaded device within the preset control time.

[0035] The comprehensive deviation degree is determined by the weighted fusion result of the coupling characteristics of the positive deviation ratio of the thermal field, the deviation amount of the thermal field, the positive deviation ratio of the static pressure, and the deviation amount of the static pressure.

[0036] The first control amount is determined based on the comprehensive deviation and the preset negative control coefficient.

[0037] Furthermore, the process of generating a second control quantity based on the actual return air temperature and the intake air temperature of the under-response device includes:

[0038] The insufficient temperature difference ratio is determined based on the proportion of times during which the real-time temperature difference between the actual return air temperature and the intake air temperature is continuously lower than the preset temperature difference threshold within the preset control time.

[0039] The return air differential sequence and the intake air differential sequence are determined based on the real-time changes in the actual return air temperature and the intake air temperature within the preset control time.

[0040] The consistency ratio is determined based on the consistency of the real-time changes in the real-time return air differential sequence and the intake air differential sequence.

[0041] The product of the insufficient temperature difference ratio, the consistent change ratio, and the preset positive control coefficient is calculated to obtain the second control amount.

[0042] Furthermore, the process of generating control commands based on the differences between the basic control quantity and the first and second control quantities, respectively, and the actual return air temperature includes:

[0043] Based on the classification of the cooling equipment under test corresponding to the basic control quantity, the special control quantity is determined to be either the first control quantity or the second control quantity.

[0044] The degree of control conflict is determined based on the temporal changes in the control directions of the basic control quantity and the dedicated control quantity within a preset adjustment period.

[0045] The degree of amplitude conflict is determined based on the basic control amount and the dedicated control amount;

[0046] The control command is determined based on the control conflict degree, the dedicated control amount, the actual return air temperature, and the amplitude conflict degree.

[0047] Furthermore, the process of determining the control command based on the control conflict degree, the dedicated control amount, and the amplitude conflict degree includes:

[0048] Based on the threshold comparison results of the regulation conflict degree and the threshold comparison results of the amplitude conflict degree, a target regulation degree is generated according to the amplitude conflict degree, the basic regulation degree, and the dedicated regulation degree.

[0049] Based on the comparison results of the threshold of the control conflict degree and the comparison results of the actual return air temperature threshold, the dedicated control amount is determined as the target control amount or the basic control amount is determined as the target control amount.

[0050] The control command is generated based on the target control amount.

[0051] Compared with existing technologies, the advantages of this invention lie in its ability to shift the management focus from the global average to local abnormal areas by identifying cooling devices with uneven heating and cooling. It calculates the probability of anomalies based on multi-parameter dynamic correlation and historical normal patterns, screening out potentially risky devices and avoiding false alarms due to single-parameter fluctuations. Risky devices are classified into over-response and under-response categories based on load responsiveness, enabling differentiated control and solving the problem of a one-size-fits-all approach under non-uniform heat loads. Simultaneously, it combines the global basic control parameters generated by a preset AI model to generate final control commands, balancing global energy efficiency with the need to eliminate local hotspots. This ensures efficient and safe intelligent control of the cooling system in dealing with non-uniform load distribution and non-cooperative operation of single-point temperature control, effectively solving problems such as non-cooperative power output competition between cooling devices, short-circuiting of hot and cold airflows, coexistence of local overheating and overcooling, and high overall energy consumption caused by non-uniform load distribution and non-cooperative operation of single-point temperature control.

[0052] Furthermore, by comparing the results of thermal field unevenness thresholds to identify cooling equipment of concern, the monitoring perspective is effectively narrowed from the overall average of the computer room to specific service areas with uneven local temperature distribution, overcoming the drawback of global monitoring being insensitive to local hotspots or airflow organization defects. Based on this, the relative temperature deviation is determined by the actual return air temperature. The return air state of different equipment, which may be affected by absolute ambient temperature, is uniformly transformed into a standardized, dimensionless, or relative value characterizing the degree to which its operation deviates from the design or expected baseline. This enables standardized assessment of the local load perception of each device, avoiding misjudgments caused by single-point temperature fluctuations.

[0053] Furthermore, by calculating the dynamic correlation between parameters, the physical coupling relationship within the system can be quantified, thereby identifying early anomalies before a single parameter exceeds the limit. By constructing associated feature vectors, the description of equipment status is upgraded from scattered indicators to comprehensive characteristics. By calculating the Mahalanobis distance between the feature vectors and the historical normal pattern set and converting it into anomaly probability, an adaptive individualized judgment benchmark is established, which can effectively distinguish between global fluctuations and real faults, significantly improving the targeting and reliability of alarms. Finally, by comparing preset probability thresholds, while allowing for random fluctuations, a statistically significant and decisive judgment can be made on abnormal states that significantly deviate from historical healthy patterns.

[0054] Furthermore, in the dynamic operation of a data center cooling system, each cooling device essentially constitutes an independent local environmental control unit. When the racks within the service area of ​​a certain cooling device experience increased load due to increased computing tasks, according to the law of conservation of energy, these racks will dissipate more heat into the air flowing through them, directly causing the actual return air temperature to rise. After the cooling device detects this temperature change signal, the controller interprets this signal as insufficient current cooling capacity to balance the increased heat load, thereby generating a calculation and outputting a command to increase the fan speed. The execution of this command triggers a physical response: the increased fan speed enhances air delivery capacity, aiming to deliver more cool air with higher dynamic pressure into the underfloor plenum, establishing a relatively higher static pressure below the corresponding service area, increasing the flow and speed of cool air entering the upper cold aisle through the perforated floor, ensuring that sufficient, low-temperature air can effectively penetrate the server racks, absorbing and carrying away the additional heat generated by the increased load. If the above adjustment is precise and effective, the server intake air temperature and the return air temperature, as the final feedback signal, will show a downward trend, thus returning the system state to the set target range, completing a successful negative feedback adjustment.

[0055] To quantitatively diagnose the health status of this closed-loop control system, multivariate time-shift correlation analysis was introduced, in which:

[0056] Temperature load correlation measures the intensity and timeliness of heat source disturbances to the temperature sensed by cooling equipment. In a healthy system, load changes should be the main driving factor of temperature changes, and the two should show a significant positive correlation. If the temperature load correlation is abnormally low, one or more of the following anomalies may exist: return air temperature sensor failure or abnormal airflow organization at the measurement point, resulting in signal distortion; severe airflow short-circuiting or bypassing in the service area, that is, the cold air delivered by the cooling equipment, after flowing out from the perforated floor, fails to flow horizontally through the server rack to absorb heat, but instead rises vertically directly or short-circuits through the rack gaps to the return air vent due to excessive suction or airflow organization defects at the return air vent, so that the heat generated by the server cannot be effectively transferred to the return air sensor.

[0057] The temperature-static pressure correlation is used to identify the effectiveness of the control link from control execution to state change. Normal regulation should manifest as follows: the static pressure adjustment action taken by the cooling equipment in response to temperature changes, increasing the static pressure deviation, helps reduce the relative temperature deviation. The two are usually negatively correlated. If this correlation is significantly low, it directly indicates that the equipment's regulation action has failed or its effect has been severely interfered with. The following anomalies may exist: There is an airflow short circuit; that is, when the cold air delivery path is short-circuited, even if the equipment increases the static pressure and airflow, this cold air does not effectively cool the server. Therefore, the change in static pressure cannot cause a change in the actual cooling effect in the server area, resulting in a very low correlation between the static pressure deviation and the relative temperature deviation. Alternatively, when this cooling equipment attempts to increase the static pressure deviation of its corresponding service area, neighboring cooling equipment reduces the overall static pressure by a larger margin, causing the regulation effect of this equipment to be offset.

[0058] Regarding the correlation between static pressure and load, changes in the load of IT equipment such as servers within a data center directly affect the demand for cooling air, thus requiring the static pressure deviation to adapt accordingly. Normally, as server load increases, the heat generated increases, requiring more cooling air to remove the heat. In this case, the cooling equipment should correspondingly increase the static pressure deviation to increase airflow; therefore, static pressure and load are usually positively correlated. If the correlation between static pressure and load is significantly abnormal, such as when the static pressure fails to increase as expected when the load increases significantly, or when the increase is far less than the load increase demand, it indicates a mismatch between the static pressure supply of the cooling equipment and the actual load. This may be because when the load in a certain area increases, the cooling equipment significantly increases the fan speed to cool down quickly. The excessively strong return air suction draws too much air from the shared static pressure box, causing the static pressure deviation below that area to decrease instead of increase. This negative correlation pattern of increased load and decreased static pressure is a clear signal of non-cooperative operation between equipment, falling into a zero-sum or even negative-sum game, and is usually a precursor to systemic instability.

[0059] Furthermore, by calculating the maximum absolute value of the time-shift cross-correlation coefficient, the dynamic coupling relationship between parameters in the cooling system caused by airflow transfer, heat exchange delay, and control response lag is effectively captured. The Mahalanobis distance is calculated based on historical normal operating duration, and the anomaly probability is determined using the relative position of the current correlation distance within the historical distance distribution, where S... dThis represents the cumulative frequency of the current sample being evaluated when the value is less than the total frequency of all S0 historical Mahalanobis distance samples. Based on all normal correlation feature vectors within a preset historical timeframe, the Mahalanobis distance between each historical correlation vector and the overall mean vector is calculated. This distance quantifies the statistical correlation and variability among temperature-load correlation, temperature-static-pressure correlation, and static-pressure-load correlation in each dimension, thus characterizing the deviation of historical correlation vectors from historical norms. After obtaining the empirical distribution of historical distances, the anomaly probability is defined by calculating the relative rank (percentile) of the current correlation distance's Mahalanobis distance within that historical empirical distribution. The behavioral pattern corresponding to the current equipment operating state is unlikely to occur in the historical normal reference group, thus quantifying the anomaly probability.

[0060] Furthermore, by analyzing the dynamic relationship between rack load rate and relative temperature deviation within a preset historical time window, the actual coupling characteristics between load changes and temperature response of each cooling device under typical operating conditions can be characterized. By evaluating the load response of the equipment, temperature anomalies can be accurately identified within a reference framework that matches the current heat load, avoiding false alarms caused by natural load fluctuations. By classifying the equipment responses, risk areas of over-cooling or under-cooling can be clearly distinguished. At the same time, historical data modeling and normalization analysis reduce the impact of instantaneous noise and short-term disturbances on judgment, improving the accuracy and stability of early warnings, and helping to optimize cooling resource allocation and reduce unnecessary energy consumption fluctuations.

[0061] Furthermore, by calculating the proportion of the duration during which the thermal field imbalance and static pressure deviation exceeded their respective preset thresholds relative to the control duration, the occurrence of abnormal states was identified, and the persistence and stubbornness of the anomalies were quantified. By calculating the average positive deviation value, the severity of the abnormal state was captured, distinguishing between minor and severe deviations. In addition, the comprehensive deviation was determined by the duration proportion and the average deviation depth, clarifying that the comprehensive index would only significantly increase when the anomaly lasted for a sufficiently long time and reached a sufficient depth. Simultaneously, by multiplying this comprehensive deviation by a preset negative control coefficient to generate the first control amount, the control direction was ensured to reduce the opening of the chilled water valve to smooth out excessive cooling output. The magnitude of the coefficient determined the controller gain, making the control amount proportional to the comprehensive severity of the anomaly, achieving a smooth, gradual rather than abrupt control output.

[0062] Furthermore, calculating the proportion of insufficient temperature difference can effectively identify subtle differences between return air temperature and intake air temperature, avoiding delays in adjustment due to failure to achieve the desired temperature difference. Forward differential analysis of return air and intake air temperatures allows for real-time capture of temperature change trends and rates, revealing the dynamic characteristics of temperature changes. The differential sequences of return air and intake air temperatures provide more granular real-time data, enhancing the accuracy of control decisions. Moreover, calculating the proportion of consistent changes between the return air differential sequence and the intake air differential sequence captures the synergy between them in the direction of change. If their changes are consistent, it indicates a more urgent need for temperature adjustment, enhancing the temperature control system's sensitivity to temperature changes while ensuring the system's timely response. By combining these factors with a preset positive control coefficient, a second control value is derived, making the adjustment more precise and in line with dynamic requirements.

[0063] Furthermore, by distinguishing between the basic control quantity and the dedicated control quantity corresponding to over-response or under-response equipment, the control quantity of each cooling device can be finely matched to its actual load response characteristics. The control conflict degree is revealed from a temporal perspective by statistically analyzing the proportion of opposite control directions between the basic and dedicated control quantities within a preset time period, showing the persistence and frequency of inconsistencies between strategies. In addition, when calculating the amplitude conflict degree, the absolute difference between the basic and dedicated control quantities is compared to their maximum value, mapping the numerical pair to a standardized scale. This ensures that the amplitude conflict degree is not affected by the absolute magnitude of the basic and dedicated control quantities, making it possible to compare conflicts under different equipment and operating conditions. Furthermore, regardless of the numerical magnitude of the basic and dedicated control quantities, as long as the relative difference between them is significant, the amplitude conflict degree will increase. Simultaneously, when the basic and dedicated control quantities have opposite signs, |B K -S K | is often greater than max(|B) K |,|S K The |) indicates that the control conflict degree is greater than 1, which clearly amplifies the severity of directional conflict. Finally, the control conflict degree, amplitude conflict degree, dedicated control quantity, and real-time return air temperature are combined to generate the final control command, so that the control behavior takes into account both the equipment's history and load characteristics, as well as the current operating status and temperature control requirements.

[0064] Furthermore, when the regulatory conflict level reaches 0.7 (high conflict) and the amplitude conflict level simultaneously exceeds the threshold, it indicates that the basic regulatory amount and the dedicated regulatory amount suggest opposite directions 70% of the time within the time window, resulting in a systemic and fundamental strategy divergence. This suggests that the basic regulatory amount generated based on AI experience judgment is severely inconsistent with the dedicated regulatory amount, and the system is on the verge of instability. Direct adoption of any single strategy may trigger violent fluctuations or vicious competition. At this point, the target regulatory amount is determined using an exponentially decaying fusion formula. When the amplitude conflict level is very low, it indicates that the basic regulatory amount and the dedicated regulatory amount are highly coordinated, and the exponential function... When the value is close to 1, the dedicated control quantity is fully adopted, thus achieving accurate and powerful correction of the global benchmark; when the magnitude conflict is very low or very high, it indicates a fundamental divergence between the basic control quantity and the dedicated control quantity, at which point the exponential function... The sharp decrease in conflict level indicates that the reliability of dedicated control quantities is low under current high-conflict scenarios, significantly weakening their contribution to the strength of recommendations. The decision weights are significantly biased towards the more robust base control quantities derived from global AI experience, ensuring that the final instructions can fully utilize the accuracy of local rules to optimize performance while effectively suppressing aggressive adjustments that might cause system oscillations during severe conflicts. When the control conflict level drops to 0.3, it indicates that the base control quantity and dedicated control quantity are converging but not yet fully coordinated. At this point, the actual return air temperature is introduced to ensure production safety. If the temperature exceeds the limit, the dedicated control quantity, which can cool down more quickly, is adopted to ensure operational safety; if the temperature does not exceed the limit, the base control quantity, which focuses more on long-term energy efficiency and global coordination, is adopted to consolidate stability. Based on this, control instructions are generated based on the target control quantity, which can dynamically balance local and global cooling needs, improve temperature control accuracy, avoid local overcooling or overheating, and reduce adjustment frequency and energy consumption fluctuations. Attached Figure Description

[0065] Figure 1 This is a flowchart of the method for dynamically controlling a data center cooling system based on AI in this embodiment;

[0066] Figure 2 This embodiment defines the logic diagram for determining the cooling equipment of interest.

[0067] Figure 3 This embodiment defines the logic diagram for determining the cooling equipment of interest.

[0068] Figure 4 This is a logic diagram for classifying risky cooling equipment in this embodiment. Detailed Implementation

[0069] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.

[0070] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.

[0071] Please see Figure 1 The diagram shown is a flowchart of the AI-based dynamic control method for a data center cooling system in this embodiment. This embodiment provides a method for AI-based dynamic control of a data center cooling system, including:

[0072] Real-time acquisition of thermal field imbalance, actual return air temperature, rack load rate, cold aisle static pressure deviation, and intake air temperature of racks in the service area for each cooling device operating based on preset cold water valve opening within the room-level data center.

[0073] Based on the threshold screening results of the thermal field unevenness, several cooling devices of interest are identified, and the relative temperature deviation is determined based on the actual return air temperature of the cooling devices of interest.

[0074] Several risky cooling devices are identified based on the relative temperature deviation, static pressure deviation, and cabinet load rate trends of the cooling devices under concern.

[0075] The risky cooling equipment is classified based on the comparison results of the relative temperature deviation of the risky cooling equipment and the load responsiveness threshold determined by the cabinet load rate, so as to obtain overloaded and underloaded equipment.

[0076] A first control amount is generated based on the thermal field imbalance and the static pressure deviation of the cold aisle of the overload response device within a preset control time period to adjust the opening of the preset chilled water valve, and a second control amount is generated based on the actual return air temperature and the intake air temperature of the underload response device.

[0077] The actual return air temperature and the static pressure deviation of the cold aisle of all the cooling equipment under test within the same preset control time period are input into the preset AI model to obtain the basic control quantity;

[0078] Control commands are generated based on the differences between the basic control quantity and the first and second control quantities, as well as the actual return air temperature.

[0079] In this embodiment, the AI-based dynamic control method for data center cooling systems is applied to a traditional open-plan data center cooling system with uneven heat loads. The main component is an open-plan data center server room. Inside the server room, IT equipment is installed in rows of standard server racks. To optimize airflow, the racks are arranged face-to-face. The area between the equipment intake sides forms a cold aisle for delivering cool air, while the corridor between the exhaust sides forms a hot aisle for collecting hot exhaust gases. The cooling equipment refers to multiple room-level air conditioning units deployed around the server room. They deliver cool air to the cold aisles through a raised floor plenum and return air from the ceiling, thus creating a circulation. Each cooling unit and its physically adjacent set of racks together constitute a logical service area. In this scenario, due to the non-uniform spatial distribution of IT workloads, significant differences in heat load arise between different service areas, leading to uncoordinated or even conflicting operating states among the cooling units as they compete for cooling capacity.

[0080] In this embodiment, a more comprehensive evaluation system for the operating status of the cooling system is constructed by acquiring multi-dimensional parameters. Among them, thermal field unevenness refers to the dispersion of air temperature distribution at different spatial locations within the service area corresponding to a single cooling device under test. It characterizes whether the heat and cold distribution within the service area is uniform. Multiple temperature sensors can be deployed at different longitudinal positions of the cold aisle, different rack air inlets, or different heights within the service area of ​​the cooling device. The thermal field unevenness is obtained by collecting air temperatures from multiple points within the service area and calculating the standard deviation of the temperatures at all sampling points. Actual return air temperature refers to the actual measured air temperature collected at the return air inlet of each computer room air conditioner or air handling unit, reflecting the local heat load level perceived by the device. This can be obtained by installing temperature sensors at the return air inlet of each cooling device to monitor the temperature of the returning air in real time. Rack load rate refers to the total real-time power consumption of the IT equipment within a group of racks belonging to the service area of ​​a certain cooling device relative to the total power consumption designed for the racks in that area. The ratio of capacitance can be obtained through the metering function of intelligent power distribution units such as intelligent PDUs or busbar monitoring systems; static pressure deviation refers to the difference between the static pressure value of the air inside the cold aisle in a data center with underfloor air supply and the target static pressure value set by the system. This can be achieved by installing a micro differential pressure transmitter at a representative location in each cold aisle, such as the middle section. The high-pressure side of the sensor is connected to the air inside the cold aisle through a pressure measuring tube to measure its absolute static pressure, while the low-pressure side can be connected to a stable static pressure reference point in the computer room, or simply setting a target reference static pressure value in the software and calculating the static pressure difference measured on the high-pressure side and the low-pressure side; inlet air temperature refers to the air temperature measured at the air inlet of each cabinet located in the service area of ​​the cooling equipment, i.e., on the front side of the equipment on the cold aisle side. This can be obtained by installing temperature sensors such as RTDs, thermocouples, or semiconductor sensors at the air inlet of the cabinet to monitor the temperature of the cold air entering the cabinet in real time.

[0081] The preset chilled water valve opening is the initial opening setting value uniformly adopted by the corresponding chilled water regulating valve of each cooling device under test. It depends on the rated cooling capacity of the computer room air conditioning unit, the design cooling load level, and the stable valve position range under historical operating conditions. It is usually set between 40% and 70%. In this embodiment, it is set to 55% to ensure that the basic cooling capacity covers the current typical load demand. The preset control duration is the fixed time window length used to generate the control quantity. It depends on the thermal inertia characteristics of the computer room air conditioning system, the response time required for cold energy transfer and airflow reconstruction, and the sensor data refresh cycle. It is usually set between 10 minutes and 30 minutes. In this embodiment, it is set to 15 minutes to balance the system's response timeliness and the stability of the control command.

[0082] In this embodiment, the preset AI model is a supervised machine learning model pre-trained and deployed in the data center control system. Essentially, it is a multi-input, single-output nonlinear regression neural network used to predict basic control parameters from the overall operating status of the current cooling system. This model is custom-trained for a specific data center topology, equipment model, airflow organization method, and historical operating data. The training data comes from at least six months of historical operating logs of the data center, covering different seasons and IT load scenarios (including typical workdays, holidays, fault recovery, etc.), with a total of no less than 100,000 samples. Each training sample consists of two parts: input features and a target label. The input features are the actual return air temperature sequence and cold aisle static pressure deviation sequence of all cooling equipment collected at a certain historical moment. The target label is the optimal global valve opening adjustment suggestion value, i.e., the optimal basic control parameter, calculated retrospectively after the global state is analyzed by domain expert strategies or offline reinforcement learning algorithms, and is considered to lead to a better future system state. Among them, domain expert strategies refer to a pre-compiled rule base based on thermodynamics and fluid mechanics principles, which outputs optimization suggestions after simulating and deducing the system response under a given state; offline reinforcement learning algorithms refer to policy networks trained in a simulation environment that can maximize long-term energy efficiency and safety rewards, and the actions output by applying them to historical states are the labels. Both are offline, non-real-time computation methods;

[0083] Specifically, the rule base in the domain expert strategy draws its rules from the analysis and summary of the data center's thermodynamic characteristics and historical operating data. These rules include, but are not limited to: threshold adjustment rules for different combinations of global average return air temperature and static pressure deviation, feedforward compensation rules based on load change trends, and conservative rules to prevent localized overheating. When generating tags offline, the system inputs historical global state data into this rule base. Through matching and calculation by the rule engine, it simulates the optimal adjustment recommendations based on expert knowledge in the current state.

[0084] The offline reinforcement learning algorithm is trained in a constructed data center thermodynamics and airflow organization simulation environment. This simulation environment can simulate the coupling relationship between cooling equipment, valves, rack loads, and airflow. Its reward function is set as a weighted sum that comprehensively reflects the overall system energy efficiency (such as the negative value of total cooling power consumption) and safety margin (such as the penalty for critical point temperature deviations from the safety threshold). The agent (policy network) learns a control policy that maximizes long-term cumulative rewards through extensive offline interaction and trial and error with this simulation environment. When generating labels, the historical states are input into this trained policy network, and its output action becomes the label value.

[0085] It should be noted that, whether it is rule-based expert strategy or simulation-based reinforcement learning, the process of generating a single optimal control variable label requires a large amount of computation, search or iterative deduction under the condition of having complete historical or simulated environment data. This process is computationally intensive and time-consuming, and cannot be directly embedded into real-time control loops that require second-level or even millisecond-level response.

[0086] By training on massive amounts of historical samples and targeting the optimal control quantity generated offline, the neural network model can learn and approximate the global optimization rules inherent in expert strategies or reinforcement learning algorithms. The resulting lightweight feedforward network only needs to input the current return air temperature and static pressure deviation sequence during the online control phase to deduce a basic control quantity within milliseconds. This transforms the complex decision-making process, which originally relied on offline, global data analysis, into a fast feedforward calculation that adapts to real-time control requirements.

[0087] The model employs a lightweight, fully connected feedforward neural network architecture. Its input layer receives a concatenated, standardized input vector, with the dimension determined by the number of cooling devices (e.g., 40 dimensions when there are 20 cooling devices). There are two hidden layers, each containing 64 neurons, using a Leaky ReLU activation function (slope coefficient set to 0.01) to mitigate the vanishing gradient problem and retain negative values. The output layer is a single neuron with no activation function, directly outputting the basic control quantity. A Dropout mechanism is introduced after the hidden layers, with a dropout rate of 0.2, and an L2 weight decay term (regularization coefficient of 10⁻⁴) is added to the loss function to improve the model's generalization ability and prevent overfitting. After sufficient training, the preset AI model shows a mean squared error (MSE) of less than 0.0025 between the predicted basic control quantity and the optimal control quantity generated by the expert policy or reinforcement learning agent under historical conditions, with a determination coefficient R² exceeding 0.96, indicating that the model can highly fit the true optimal control behavior. The model's input is limited to two sequences of length N: actual return air temperature and cold aisle static pressure deviation, both formed by the actual return air temperature and cold aisle static pressure deviation of all tested cooling equipment within the same time period. Before being input into the model, these two sequences are standardized by subtracting the mean of historical operating data from each and dividing by its standard deviation. This eliminates dimensional differences and accelerates model convergence, ultimately outputting a scalar-form basic control quantity. This control quantity represents a unified, fundamental adjustment suggestion for the valve openings of all cooling equipment in the system. These standardized parameters are obtained offline through analysis of at least six months of historical operating logs during the model training phase and remain fixed after model deployment, not dynamically updated with online data. This ensures the consistency of the input distribution and avoids model performance degradation due to operating condition drift. After standardization, the system concatenates these two N-length normalized sequences in equipment order to form a joint input sequence of length 2N, which serves as the sole input to the neural network.The sequence is then fed into a lightweight, fully connected feedforward neural network: First, the input vector is linearly multiplied by the weight matrix trained in the first hidden layer and a bias is added. Then, a nonlinear transformation is performed using the Leaky ReLU activation function. This operation weights and nonlinearly transforms the original temperature and pressure readings, initially abstracting primary state features such as the global average temperature rise trend and the static pressure imbalance between regions. Based on this, these primary feature vectors undergo the same linear weighting and nonlinear activation operations with the weight matrix of the second hidden layer, thereby further combining and abstracting higher-order state patterns that reflect the global coupling relationship of the system based on the features of the first layer. Finally, the higher-order feature vectors output by the second hidden layer are fed into a single linear neuron in the output layer. This neuron performs a final linear weighted summation. The weight vector of the output layer is essentially a set of decision factors learned by the model from massive historical data. It assigns quantified contribution or decision weights to the various complex state patterns identified by the preceding network. All these weighted contributions are summed in the output layer to directly generate the final basic control quantity. The entire forward computation process is executed in inference mode on the edge controller, without involving any parameter updates or online learning, with an inference latency of less than 100 milliseconds. The basic control quantity output by the model serves as a decision benchmark derived from historical global optimization experience. This benchmark is then fused with subsequent correction quantities based on real-time risk identification to ultimately generate a coordinated control command that combines energy efficiency and safety. This ensures that the real-time control loop can inherit the global energy efficiency intelligence from offline optimization while also adding refined and safety-based adjustments based on real-time risks.

[0088] By identifying cooling devices with uneven local heating and cooling, the management focus shifts from the global average to local anomaly areas. Based on multi-parameter dynamic correlation and historical normal patterns, the probability of anomalies is calculated to screen potentially risky devices, avoiding false alarms from single-parameter fluctuations. Risky devices are categorized into over-response and under-response types based on load responsiveness, enabling differentiated control and resolving the one-size-fits-all problem of uniform strategies under non-uniform heat loads. Simultaneously, combined with global basic control values ​​generated by a pre-set AI model, the final control command is generated, balancing global energy efficiency with the need to eliminate local hotspots. This ensures efficient and safe intelligent control of the cooling system in the face of non-uniform load distribution and uncoordinated operation of single-point temperature control, effectively solving problems such as uncoordinated power output competition between cooling devices, short-circuiting of hot and cold airflows, coexistence of local overheating and overcooling, and high overall energy consumption caused by non-uniform load distribution and uncoordinated operation of single-point temperature control.

[0089] Please see Figure 2 As shown, this is the logic diagram for determining the cooling equipment of interest in this embodiment. In this embodiment, the process of identifying several cooling equipment of interest based on the threshold screening result of the thermal field imbalance, and determining the relative temperature deviation based on the actual return air temperature of the cooling equipment of interest, includes:

[0090] When the thermal field imbalance is greater than a preset balance threshold, the cooling device under test is determined to be the cooling device of interest.

[0091] The relative temperature deviation is obtained by calculating the relative deviation between the actual return air temperature and the preset return air threshold of the cooling equipment under interest.

[0092] The preset equalization threshold is a critical value used to determine whether the temperature distribution within the service area of ​​a single cooling device is uniform. It depends on the requirements of the data center's design for local temperature uniformity, the density of sensor deployment, and the measurement error range. It is usually set between 2°C and 5°C. In this embodiment, it is set to 3°C, which can effectively filter out devices with uneven heating and cooling or potential airflow organization problems in the area, realizing a shift in focus from global monitoring to local key areas. The preset return air threshold is a reference temperature value used to calculate the relative temperature deviation. It depends on the data center's designed operating temperature range, the equipment manufacturer's reliability indicators, and energy efficiency strategies. It is usually set between 24°C and 28°C. In this embodiment, it is set to 26°C, which can provide a stable and reasonable calculation benchmark for evaluating the deviation of the actual return air temperature of each cooling device from the system's expected value.

[0093] By comparing the results of thermal field unevenness thresholds, the monitoring perspective is effectively narrowed from the overall average of the computer room to specific service areas with uneven local temperature distribution. This overcomes the drawback of global monitoring being insensitive to local hotspots or airflow organization defects. Furthermore, by determining the relative temperature deviation through actual return air temperature, the return air status of different devices, which may be affected by absolute ambient temperature, is uniformly transformed into a standardized, dimensionless, or relative value characterizing the degree to which its operation deviates from the design or expected baseline. This enables standardized assessment of the local load perception of each device, avoiding misjudgments caused by single-point temperature fluctuations.

[0094] Please see Figure 3 As shown, this is the logic diagram for determining risky cooling devices in this embodiment. In this embodiment, the process of determining several risky cooling devices based on the relative temperature deviation, static pressure deviation, and cabinet load rate change trends of the cooling devices of interest includes:

[0095] Based on the relative temperature deviation, static pressure deviation, and cabinet load rate characteristics within a preset time period, the temperature-static pressure correlation, temperature-load correlation, and static pressure-load correlation are determined respectively.

[0096] A correlation feature vector is constructed based on the temperature-static pressure correlation, the temperature-load correlation, and the static pressure-load correlation, and the anomaly probability of the correlation feature vector is determined based on the relative historical anomaly deviation of the correlation feature vector.

[0097] When the anomaly probability is greater than a preset probability threshold, the cooling device of concern is determined to be the risk cooling device.

[0098] The preset time frame is the length of the statistical analysis time window used to calculate the dynamic correlation between key parameters. It depends on the load fluctuation cycle and the response time of heat transfer and pressure balance, and is usually set between 30 and 120 minutes. In this embodiment, it is set to 60 minutes to ensure that the characteristic parameters reflect the true coupling relationship trend between the parameters. The preset probability threshold is the threshold used to map the historical deviation of the associated characteristic vector to the anomaly probability and determine whether to declare a risk. It depends on the sensitivity requirements for operational risk warning and the setting of the statistical confidence level of historical normal operating conditions. It is usually set between 0.7 and 0.9. In this embodiment, it is set to 0.8, which can effectively eliminate most random fluctuation interference and reliably alarm for states that continuously and significantly deviate from the historical normal cooperative mode.

[0099] By calculating the dynamic correlation between parameters, the physical coupling relationship within the system can be quantified, thereby identifying early anomalies before a single parameter exceeds the limit. Constructing associated feature vectors upgrades the description of equipment status from scattered indicators to comprehensive characteristics. By calculating the Mahalanobis distance between the feature vectors and the historical normal pattern set and converting it into anomaly probabilities, an adaptive, individualized judgment benchmark is established, effectively distinguishing between global fluctuations and actual faults, significantly improving the targeting and reliability of alarms. Finally, by comparing preset probability thresholds, while allowing for random fluctuations, a statistically significant and decisive judgment can be made on abnormal states that significantly deviate from historical healthy patterns.

[0100] Specifically, the process of determining the temperature-static pressure correlation, temperature-load correlation, and static pressure-load correlation based on the relative temperature deviation, static pressure deviation, and cabinet load rate characteristics within a preset time period includes:

[0101] The maximum absolute value of the time-shift cross-correlation coefficient between the normalized relative temperature deviation and the normalized static pressure deviation within the preset pressure-temperature time-shift range is calculated to determine the temperature-static pressure correlation.

[0102] The maximum absolute value of the time-shift cross-correlation coefficient between the normalized relative temperature deviation and the normalized cabinet load rate within the preset load temperature time-shift range is calculated to determine the temperature-load correlation.

[0103] The maximum absolute value of the time-shift cross-correlation coefficient between the normalized static pressure deviation and the normalized cabinet load rate within the preset load time-shift range is calculated to determine the static pressure load correlation.

[0104] The preset temperature-pressure time shift range is the maximum time shift range used to calculate the time shift cross-correlation coefficient between temperature and static pressure. It depends on the transmission delay of the effect of changes in the cooling system's supply air static pressure on the return air temperature, and is usually set between 5 and 15 minutes. In this embodiment, it is set to 10 minutes, which can cover the typical lag time from static pressure adjustment to temperature response. The preset load temperature time shift range is the maximum time shift range used to calculate the time shift correlation coefficient between temperature and load. It depends on the delay of changes in return air temperature caused by changes in cabinet load, and is usually set between 10 and 30 minutes. In this embodiment, it is set to 20 minutes, which can cover the typical lag time from load change to temperature response. The preset load pressure time shift range is the maximum time shift range used to calculate the time shift cross-correlation coefficient between static pressure and load. It depends on the delay of static pressure adjustment caused by changes in cabinet load, and is usually set between 2 and 8 minutes. In this embodiment, it is set to 5 minutes, which can cover the typical lag time from load change to static pressure adjustment.

[0105] The normalization process in this embodiment uses the max-min normalization method, which linearly transforms the original data to the [0,1] interval to eliminate the influence of different physical units on the subsequent correlation coefficient calculation. This avoids the problem that a certain parameter may dominate the correlation analysis due to different units of static pressure, temperature and load, thus ensuring the accuracy and fairness of the subsequent calculation of time-shifted cross-correlation coefficients.

[0106] In the dynamic operation of a data center cooling system, each cooling device essentially constitutes an independent local environmental control unit. When the racks within a cooling device's service area experience increased load due to increased computing tasks, according to the law of conservation of energy, these racks will dissipate more heat into the air flowing through them, directly causing an increase in the actual return air temperature. After the cooling device detects this temperature change, the controller interprets it as insufficient cooling capacity to balance the increased heat load. It then calculates and outputs a command to increase the fan speed. The execution of this command triggers a physical response: the increased fan speed enhances air delivery capacity, aiming to deliver more cool air with higher dynamic pressure into the plenum beneath the floor, establishing a relatively higher static pressure below the corresponding service area. This increases the flow and velocity of cool air through the perforated floor into the upper cooling aisle, ensuring sufficient, low-temperature air can effectively penetrate the server racks, absorbing and carrying away the additional heat generated by the increased load. If the above adjustments are precise and effective, the server intake air temperature and the return air temperature, as the final feedback signal, will show a downward trend, causing the system state to return to the set target range, completing a successful negative feedback adjustment.

[0107] To quantitatively diagnose the health status of this closed-loop control system, multivariate time-shift correlation analysis was introduced, in which:

[0108] Temperature load correlation measures the intensity and timeliness of heat source disturbances to the temperature sensed by cooling equipment. In a healthy system, load changes should be the main driver of temperature changes, and the two should show a significant positive correlation. If the temperature load correlation is abnormally low, one or more of the following anomalies may exist: return air temperature sensor failure or abnormal airflow organization at the measurement point, resulting in signal distortion; severe airflow short-circuiting or bypassing in the service area, that is, the cold air delivered by the cooling equipment, after flowing out from the perforated floor, fails to flow horizontally through the server rack to absorb heat, but instead rises vertically directly or flows back to the return air vent through the rack gap due to excessive suction or airflow organization defects, so that the heat generated by the server cannot be effectively transferred to the return air sensor.

[0109] The temperature-static pressure correlation is used to identify the effectiveness of the control link from control execution to state change. Normal regulation should manifest as follows: the static pressure adjustment action taken by the cooling equipment in response to temperature changes, increasing the static pressure deviation, helps reduce the relative temperature deviation. The two are usually negatively correlated. If this correlation is significantly low, it directly indicates that the equipment's regulation action has failed or its effect has been severely interfered with. The following anomalies may exist: There is an airflow short circuit; that is, when the cold air delivery path is short-circuited, even if the equipment increases the static pressure and airflow, this cold air does not effectively cool the server. Therefore, the change in static pressure cannot cause a change in the actual cooling effect in the server area, resulting in a very low correlation between the static pressure deviation and the relative temperature deviation. Alternatively, when this cooling equipment attempts to increase the static pressure deviation of its corresponding service area, neighboring cooling equipment reduces the overall static pressure by a larger margin, causing the regulation effect of this equipment to be offset.

[0110] Regarding the correlation between static pressure and load, changes in the load of IT equipment such as servers within a data center directly affect the demand for cooling air, thus requiring the static pressure deviation to adapt accordingly. Normally, as server load increases, the heat generated increases, requiring more cooling air to remove the heat. In this case, the cooling equipment should correspondingly increase the static pressure deviation to increase airflow; therefore, static pressure and load are usually positively correlated. If the correlation between static pressure and load is significantly abnormal, such as when the static pressure fails to increase as expected when the load increases significantly, or when the increase is far less than the load increase demand, it indicates a mismatch between the static pressure supply of the cooling equipment and the actual load. This may be because when the load in a certain area increases, the cooling equipment significantly increases the fan speed to cool down quickly. The excessively strong return air suction draws too much air from the shared static pressure box, causing the static pressure deviation below that area to decrease instead of increase. This negative correlation pattern of increased load and decreased static pressure is a clear signal of non-cooperative operation between equipment, falling into a zero-sum or even negative-sum game, and is usually a precursor to systemic instability.

[0111] Specifically, the process of determining the anomaly probability of the associated feature vector based on its relative historical anomaly deviation includes:

[0112] The associated feature vectors of all the cooling devices under interest within the previous preset normal operating time are recorded as historical associated vectors, and the Mahalanobis distance between each historical associated vector and the overall mean vector determined based on the historical associated vectors is calculated to obtain several historical associated distances.

[0113] The Mahalanobis distance between the associated feature vector and the overall mean vector is calculated to obtain the current association distance. The anomaly probability is determined based on the ratio of the number of all historical association distances less than the current association distance to the total number of all historical association distances, where Y=S d / S0, where Y is the anomaly probability, S d S0 is the total number of historical association distances that are less than the current association distance.

[0114] The preset normal operating time is the historical data duration used to establish a normal correlation feature vector benchmark. It depends on the periodic characteristics of the data center operation and the minimum time window for stable operation. It is usually set between 7 and 30 days. In this embodiment, it is set to 14 days, which can cover the complete work cycle changes and form a stable normal benchmark.

[0115] By calculating the maximum absolute value of the time-shift cross-correlation coefficient, the dynamic coupling relationship between parameters in the cooling system caused by airflow transfer, heat exchange delay, and control response lag is effectively captured. The Mahalanobis distance is calculated based on historical normal operating duration, and the anomaly probability is determined using the relative position of the current correlation distance within the historical distance distribution, where S... d This represents the cumulative frequency of the current sample being evaluated when the value is less than the total frequency of all S0 historical Mahalanobis distance samples. Based on all normal correlation feature vectors within a preset historical timeframe, the Mahalanobis distance between each historical correlation vector and the overall mean vector is calculated. This distance quantifies the statistical correlation and variability among temperature-load correlation, temperature-static-pressure correlation, and static-pressure-load correlation in each dimension, thus characterizing the deviation of historical correlation vectors from historical norms. After obtaining the empirical distribution of historical distances, the anomaly probability is defined by calculating the relative rank (percentile) of the current correlation distance's Mahalanobis distance within that historical empirical distribution. The behavioral pattern corresponding to the current equipment operating state is unlikely to occur in the historical normal reference group, thus quantifying the anomaly probability.

[0116] Please see Figure 4As shown, this is a logic diagram for classifying risky cooling equipment in this embodiment. In this embodiment, the process of classifying risky cooling equipment based on the comparison result of the relative temperature deviation of the risky cooling equipment and the threshold of load responsiveness determined by the rack load rate to obtain overloaded and underloaded equipment includes:

[0117] The least squares method is used to perform linear regression fitting on the cabinet load rate and the relative temperature deviation of the risk cooling equipment within a preset historical period to obtain a linear fitting curve. The slope of the linear fitting curve is determined as the preset heat transfer coefficient, where W = α × J + b, where W is the relative temperature deviation, α is the preset heat transfer coefficient, J is the cabinet load rate, b is the intercept, which is the theoretical relative temperature deviation benchmark value at zero load, reflecting the inherent thermal deviation of the environment and system.

[0118] The load responsiveness is determined based on the cabinet load rate, the preset heat transfer coefficient, and the relative temperature deviation, where X = W / (α×J)-1, and X is the load responsiveness.

[0119] When the load responsiveness is greater than a preset response threshold, the risk cooling device is determined to be an overloaded device, and when the load responsiveness is less than the preset response threshold, the risk cooling device is determined to be an underloaded device.

[0120] In this embodiment, in the formula X=W / (α×J)-1 for load responsiveness, the numerator is the relative temperature deviation, representing the real-time heat load intensity sensed by the equipment. It is the relative deviation between the actual return air temperature and the preset target value. However, it is a mixed signal that is easily contaminated and may include the actual load heat, thermal interference from adjacent areas, and false temperature rises caused by airflow organization problems. The rack load rate is an objective, independent, and accurate representation of the rack's heat, representing the true source of heat generation and is unaffected by any airflow issues. The preset heat transfer coefficient is a health baseline model established for the equipment. It is obtained by linear regression fitting of the rack load rate and the corresponding relative temperature deviation under normal operating conditions over a long period of time in the equipment's history. Under historical healthy operating conditions, for every unit increase in load rate in the service area corresponding to the cooling equipment, its return air temperature typically increases by α units. The intercept b reflects the system's basic thermal deviation at zero load. The denominator is calculated using the current objective rack load rate and preset heat transfer coefficient to determine the theoretical relative temperature deviation of the equipment under the current load. W / (α×J) compares the actual relative temperature deviation with the theoretical relative temperature deviation, revealing the multiple relationship between the actual thermal response and the theoretical value. Subtracting 1 centers the ratio to zero. That is, when X≈0, it indicates that the actual response matches the theoretical value, which is a normal response. When X>0 and exceeds the preset positive threshold, it indicates that the actual temperature rise significantly exceeds the theoretical value. This will cause the cooling equipment to operate faster, indicating that the equipment is in an over-response state. The reason may be: airflow short-circuit, that is, cold air is returned directly to the server without cooling it, causing the server to over-respond. Heat is generated, which in turn produces hot exhaust gas at an even higher temperature. This exhaust gas mixes with the short-circuited cold air in a distorted manner, ultimately causing the return air sensor to read a temperature value higher than that of normal effective cooling. This results in an artificially high return air temperature and ineffective regulation, or it may be suffering from severe thermal interference from nearby equipment. When X < 0 and is below the preset negative threshold, it indicates that the actual temperature rise is significantly lower than the theoretical value. The actual relative temperature deviation sensed by the cooling equipment is much lower than the theoretical relative temperature deviation, so it will reduce its operation because it believes that the cabinet does not need higher cooling at this time. Therefore, the cooling equipment exhibits under-response, which may be due to insufficient cooling capacity, load signal distortion, or benefiting from the cooling overflow of the adjacent area.

[0121] The preset history duration is the length of the statistical time window used to calculate the linear fit between the cabinet load rate and the relative temperature deviation of the risk cooling equipment to estimate the heat transfer coefficient. It depends on the thermal inertia of the equipment, the load fluctuation cycle, and the reliability requirements of the fit. It is usually set between 24 hours and 7 days. In this embodiment, it is set to 72 hours, which can obtain a robust heat transfer coefficient estimate while taking into account the sample size and response sensitivity. The preset response threshold is the critical value used to determine whether the calculated load response is over-response or under-response. It depends on the maintenance's tolerance for false alarms and false alarms and the requirements for cooling matching accuracy. It is usually set between 0.10 and 0.40. In this embodiment, it is set to 0.25, which can avoid misjudging short-term disturbances.

[0122] By analyzing the dynamic relationship between rack load rate and relative temperature deviation within a preset historical time window, the actual coupling characteristics between load changes and temperature response of each cooling device under typical operating conditions can be characterized. By evaluating the load response of the equipment, temperature anomalies can be accurately identified within a reference framework matching the current heat load, avoiding false alarms caused by natural load fluctuations. Classifying equipment responses clearly distinguishes risk areas of over-cooling or under-cooling. Simultaneously, historical data modeling and normalization analysis reduce the impact of instantaneous noise and short-term disturbances on judgment, improving the accuracy and stability of early warnings, and helping to optimize cooling resource allocation and reduce unnecessary energy consumption fluctuations.

[0123] Specifically, the process of generating a first control amount to adjust the preset cold water valve opening based on the thermal field imbalance and the cold aisle static pressure deviation of the overload response device within a preset control time includes:

[0124] The proportion of the total length of the time window during which the thermal field imbalance continuously exceeds the preset thermal field threshold to the preset control time is calculated to obtain the positive deviation ratio of the thermal field. The proportion of the total length of the time window during which the cold aisle static pressure deviation continuously exceeds the preset static pressure threshold to the preset control time is also calculated to obtain the positive deviation ratio of the static pressure.

[0125] Calculate the average of the positive relative deviations between the thermal field imbalance degree and the preset thermal field threshold at each moment within the preset control time period to obtain the thermal field deviation, where P R = (R-R0) / R0, where P R R is the thermal field deviation, R is the thermal field unevenness, R0 is the preset thermal field threshold, and the average of the positive relative deviations between the cold channel static pressure deviation and the preset static pressure threshold at each moment within the preset control time is calculated to obtain the static pressure deviation, P. L = (L-L0) / L0, where P LIt is the static pressure deviation, where L is the cold aisle static pressure deviation and L0 is the preset static pressure threshold.

[0126] The product of the positive deviation ratio of the thermal field and the deviation amount of the thermal field and the product of the positive deviation ratio of the static pressure and the deviation amount of the static pressure are weighted and summed to obtain the comprehensive deviation degree.

[0127] The product of the overall deviation and the preset negative control coefficient is calculated to obtain the first control amount.

[0128] The preset thermal field threshold is a critical value used to determine whether the thermal field imbalance is in an abnormal state. It depends on the temperature difference and thermal environment uniformity requirements of the data center's hot and cold aisle design and is usually set between 3°C and 8°C. In this embodiment, it is set to 5°C, which can effectively identify significant thermal field deviations caused by local overheating or uneven cooling distribution. The preset static pressure threshold is a critical value used to determine whether the static pressure deviation of the cold aisle is in an abnormal state. It depends on the static pressure setting accuracy of the air supply system and the stability requirements of airflow organization. It is usually set between 5Pa and 15Pa. In this embodiment, it is set to 10Pa, which can sensitively detect abnormal static pressure fluctuations caused by fan control misalignment or changes in air resistance. The preset negative control coefficient is a proportional coefficient used to convert the comprehensive deviation into a specific reduction in the opening of the chilled water valve. It depends on the adjustment characteristics of the chilled water valve and the maximum allowable adjustment range of the system. It is usually set between -0.05 and -0.15. In this embodiment, it is set to -0.1, which can convert the calculated deviation into a moderate and safe valve closing command, thereby suppressing oscillations caused by over-response equipment.

[0129] In this embodiment, during the weighted summation of the product of the positive deviation ratio of the thermal field and the deviation amount of the thermal field, and the product of the positive deviation ratio of the static pressure and the deviation amount of the static pressure, the weight corresponding to the product of the positive deviation ratio of the thermal field and the deviation amount of the thermal field is a coefficient used to measure the influence of thermal field unevenness on the control decision of over-response equipment. It depends on the airflow organization design, cooling system architecture, and different priority requirements for temperature uniformity and pressure stability of the specific data center, and is usually set between 0.5 and 0.8. In this embodiment, it is set to 0.6, which can highlight the dominant role of thermal field uniformity in the control of over-response equipment and ensure that the system prioritizes the correction of regional uneven heating and cooling. The weight corresponding to the product of the positive deviation ratio of the static pressure and the deviation amount of the static pressure is a coefficient used to measure the influence of static pressure deviation on the control decision of over-response equipment. It depends on the stability of airflow delivery and the avoidance of interfering with the temperature field due to excessive adjustment of static pressure. It is usually set between 0.2 and 0.5. In this embodiment, it is set to 0.4, which can appropriately incorporate the supporting role of static pressure stability on airflow organization.

[0130] By calculating the proportion of the duration during which thermal field imbalance and static pressure deviation exceeded their respective preset thresholds relative to the control duration, the occurrence of abnormal states was identified, and the persistence and stubbornness of the anomalies were quantified. The severity of the abnormal state was captured by calculating the average positive deviation value, distinguishing between minor and severe deviations. Furthermore, the overall deviation was determined by the duration proportion and the average deviation depth, clarifying that the overall index would only significantly increase when the anomaly lasted for a sufficiently long time and reached a sufficient depth. Simultaneously, by multiplying this overall deviation by a preset negative control coefficient to generate the first control amount, the control direction was ensured to reduce the opening of the chilled water valve to mitigate excessive cooling output. The magnitude of the coefficient determined the controller gain, making the control amount proportional to the overall severity of the anomaly, achieving a smooth, gradual rather than abrupt, control output.

[0131] Specifically, the process of generating a second control quantity based on the actual return air temperature and the intake air temperature of the under-response device includes:

[0132] Calculate the proportion of the total duration during which the difference between the actual return air temperature and the intake air temperature is less than a preset temperature difference threshold at each moment within the preset control time, so as to obtain the proportion of insufficient temperature difference.

[0133] Calculate the forward difference of the actual return air temperature within the preset control time to obtain the return air difference sequence, and calculate the forward difference of the intake air temperature within the preset control time to obtain the intake air difference sequence.

[0134] Calculate the proportion of the number of times when the forward difference of the return air temperature at each time in the return air differential sequence and the forward difference of the intake air temperature at each time in the intake air differential sequence are both positive or both negative, relative to the total number of times in the preset control duration, so as to obtain the consistency ratio of change.

[0135] The product of the insufficient temperature difference ratio, the consistent change ratio, and the preset positive control coefficient is calculated to obtain the second control amount.

[0136] The preset temperature difference threshold is a critical value used to determine whether the difference between the actual return air temperature and the intake air temperature is too small. It depends on the design temperature difference and heat exchange efficiency requirements of the data center cooling system and is usually set between 2°C and 8°C. In this embodiment, it is set to 5°C, which can effectively identify the situation of insufficient cooling effect. The preset positive control coefficient is a proportional coefficient used to calculate the second control quantity. It depends on the adjustment characteristics of the chilled water valve and the adjustment range allowed by the system. It is usually set between 0.1 and 0.5. In this embodiment, it is set to 0.2, which can convert the calculated proportion into a moderate and safe valve opening command to enhance the cooling response.

[0137] Calculating the proportion of insufficient temperature difference effectively identifies subtle differences between return air and intake air temperatures, preventing delays in adjustments due to failure to achieve the desired temperature difference. Forward differential analysis of return and intake air temperatures allows for real-time capture of temperature change trends and rates, revealing the dynamic characteristics of temperature variations. The differential sequences of return and intake air temperatures provide more granular real-time data, enhancing the accuracy of control decisions. Furthermore, calculating the consistency ratio between the return and intake differential sequences captures the synergy between them in terms of direction of change. Consistent changes indicate a more urgent need for temperature adjustment, enhancing the temperature control system's sensitivity to temperature changes while ensuring timely system response. Combining these factors with a preset positive control coefficient yields a second control value, making adjustments more precise and aligned with dynamic requirements.

[0138] Specifically, the process of generating control commands based on the differences between the basic control quantity and the first and second control quantities, as well as the actual return air temperature, includes:

[0139] When the cooling device under test corresponding to the basic control quantity is the overload response device, the dedicated control quantity is determined as the first control quantity; and when the cooling device under test corresponding to the basic control quantity is the underload response device, the dedicated control quantity is determined as the second control quantity.

[0140] The proportion of times when the basic control quantity and the dedicated control quantity have opposite control directions within the preset adjustment period is statistically analyzed to obtain the degree of control conflict.

[0141] The amplitude conflict degree is determined based on the basic control amount and the dedicated control amount, where U=|B K -S K | / max(|B K |,|S K |), where U is the magnitude conflict degree, B K It is the basic control quantity, S K It is a dedicated control quantity;

[0142] The control command is determined based on the control conflict degree, the dedicated control amount, the actual return air temperature, and the amplitude conflict degree.

[0143] The preset adjustment duration is the length of the observation time window used to statistically analyze the directional conflict between the basic control quantity and the dedicated control quantity. It depends on the response time of the cooling system, the sensor data refresh frequency, and the tolerance of the control strategy to short-term fluctuations. It is usually set between 5 minutes and 20 minutes. In this embodiment, it is set to 10 minutes, which can effectively capture the overall coordination trend of the two types of control suggestions in a short period of time and avoid misjudgment caused by differences in a single instantaneous calculation.

[0144] By distinguishing between the basic control quantity and the dedicated control quantity corresponding to over-response or under-response equipment, the control quantity of each cooling device can be finely matched to its actual load response characteristics. Control conflict degree is revealed from a temporal perspective by statistically analyzing the proportion of opposite control directions between the basic and dedicated control quantities within a preset time period, showing the persistence and frequency of inconsistencies between strategies. Furthermore, when calculating amplitude conflict degree, the absolute difference between the basic and dedicated control quantities is compared to their maximum value, mapping the numerical pair to a standardized scale. This ensures that amplitude conflict degree is not affected by the absolute magnitude of the basic and dedicated control quantities, making it possible to compare conflicts under different equipment and operating conditions. Regardless of the numerical magnitude of the basic and dedicated control quantities, as long as the relative difference between them is significant, the amplitude conflict degree will increase. Additionally, when the basic and dedicated control quantities have opposite signs, |B K -S K | is often greater than max(|B) K |,|S K The |) indicates that the control conflict degree is greater than 1, which clearly amplifies the severity of directional conflict. Finally, the control conflict degree, amplitude conflict degree, dedicated control quantity, and real-time return air temperature are combined to generate the final control command, so that the control behavior takes into account both the equipment's history and load characteristics, as well as the current operating status and temperature control requirements.

[0145] Specifically, the process of determining the control command based on the control conflict degree, the dedicated control amount, and the amplitude conflict degree includes:

[0146] When the degree of control conflict equals a preset conflict threshold and the degree of amplitude conflict is greater than a preset amplitude threshold, a target control quantity is generated based on the degree of amplitude conflict, the basic control quantity, and the dedicated control quantity, wherein...

[0147] ,

[0148] in, It is the target control quantity. It is a preset attenuation coefficient. It is the degree of amplitude conflict, and the control command is generated according to the target control amount;

[0149] Based on the fact that the degree of control conflict is equal to the preset change threshold, when the actual return air temperature is greater than the preset temperature threshold, the special control amount is determined as the target control amount, or when the actual return air temperature is less than the preset temperature threshold, the basic control amount is determined as the target control amount.

[0150] The control command is generated based on the target control amount;

[0151] Otherwise, the system will automatically trigger a manual intervention mechanism, notifying administrators via audible and visual alarms or system message pushes that the current control logic is abnormal and requires manual intervention.

[0152] The preset conflict threshold is the boundary value for determining whether the degree of control conflict has reached the point where a special arbitration mechanism needs to be activated. It depends on the system's tolerance for control direction conflicts and is usually set between 0.6 and 0.8. In this embodiment, it is set to 0.7, which can accurately identify abnormal states where directional conflicts persist. The preset amplitude threshold is the criterion for judging whether there is a significant difference in amplitude between the basic control quantity and the dedicated control quantity. It depends on the system's tolerance for control quantity differences and the need to avoid unnecessary complex processing of minor differences. It is usually set between 0.3 and 0.6. In this embodiment, it is set to 0.5, which can filter out serious conflict situations where control suggestions not only contradict each other in direction but also have significant differences in adjustment amplitude. The preset attenuation coefficient is the coefficient used to adjust the fusion amplitude of the basic control quantity and the dedicated control quantity when amplitude conflicts exist. It depends on the system's trust in the dedicated strategy, the desired fusion smoothness, and the need to suppress oscillations. The requirement is typically set between 1.0 and 3.0; in this embodiment, it is set to 2.0. This allows for a rapid reduction in the weight of the dedicated control quantity through an exponential decay function when the magnitude of the conflict increases. The preset change threshold is the boundary value used to determine whether the degree of control conflict has reached the level indicating a significant change in the control direction. It depends on the system's sensitivity to the trend of changes in the control direction and is typically set between 0.2 and 0.4. In this embodiment, it is set to 0.3, which can effectively capture the key node where the control direction changes from conflict to consistency. The preset temperature threshold is the safe critical temperature used to determine whether the actual return air temperature is too high, thereby determining the final adoption right of the control command. It depends on the server's inlet air temperature requirements, the equipment's safety red line, and the operating energy efficiency curve. It is typically set between 27°C and 30°C; in this embodiment, it is set to 28°C, which can ensure that the temperature control decision responds promptly to abnormal temperature changes while maintaining system stability.

[0153] When the regulatory conflict level reaches 0.7 (high conflict) and the amplitude conflict level also exceeds the threshold, it indicates that the basic and dedicated regulatory amounts suggest opposite directions 70% of the time within the time window, resulting in a systemic and fundamental strategy divergence. This suggests a severe mismatch between the basic and dedicated regulatory amounts generated based on AI experience-based judgment, placing the system on the verge of instability. Direct adoption of any single strategy could trigger violent fluctuations or vicious competition. In this case, the target regulatory amount is determined using an exponentially decaying fusion formula. When the amplitude conflict level is very low, it indicates a high degree of coordination between the basic and dedicated regulatory amounts, at which point the exponential function... When the value is close to 1, the dedicated control quantity is fully adopted, thus achieving accurate and powerful correction of the global benchmark; when the magnitude conflict is very low or very high, it indicates a fundamental divergence between the basic control quantity and the dedicated control quantity, at which point the exponential function... The sharp decrease in conflict level indicates that the reliability of dedicated control quantities is low under current high-conflict scenarios, significantly weakening their contribution to the strength of recommendations. The decision weights are significantly biased towards the more robust base control quantities derived from global AI experience, ensuring that the final instructions can fully utilize the accuracy of local rules to optimize performance while effectively suppressing aggressive adjustments that might cause system oscillations during severe conflicts. When the control conflict level drops to 0.3, it indicates that the base control quantity and dedicated control quantity are converging but not yet fully coordinated. At this point, the actual return air temperature is introduced to ensure production safety. If the temperature exceeds the limit, the dedicated control quantity, which can cool down more quickly, is adopted to ensure operational safety; if the temperature does not exceed the limit, the base control quantity, which focuses more on long-term energy efficiency and global coordination, is adopted to consolidate stability. Based on this, control instructions are generated based on the target control quantity, which can dynamically balance local and global cooling needs, improve temperature control accuracy, avoid local overcooling or overheating, and reduce adjustment frequency and energy consumption fluctuations.

[0154] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for dynamically controlling a data center cooling system based on AI, characterized in that, include: Real-time acquisition of thermal field imbalance, actual return air temperature, rack load rate, cold aisle static pressure deviation, and intake air temperature of racks in the service area for each cooling device operating based on preset cold water valve opening within the room-level data center. Based on the threshold screening results of the thermal field unevenness, several cooling devices of interest are identified, and the relative temperature deviation is determined based on the actual return air temperature of the cooling devices of interest. Several risky cooling devices are identified based on the relative temperature deviation, static pressure deviation, and cabinet load rate trends of the cooling devices under concern. The risky cooling equipment is classified based on the comparison results of the relative temperature deviation of the risky cooling equipment and the load responsiveness threshold determined by the cabinet load rate, so as to obtain overloaded and underloaded equipment. A first control amount is generated based on the thermal field imbalance and the static pressure deviation of the cold aisle of the overload response device within a preset control time period to adjust the opening of the preset chilled water valve, and a second control amount is generated based on the actual return air temperature and the intake air temperature of the underload response device. The actual return air temperature and the static pressure deviation of the cold aisle of all the cooling equipment under test within the same preset control time period are input into the preset AI model to obtain the basic control quantity; Control commands are generated based on the differences between the basic control quantity and the first and second control quantities, as well as the actual return air temperature.

2. The method for dynamically controlling a data center cooling system based on AI according to claim 1, characterized in that, The process of identifying several cooling devices of interest based on the threshold screening results of the thermal field unevenness, and determining the relative temperature deviation based on the actual return air temperature of the cooling devices of interest, includes: When the thermal field imbalance is greater than a preset balance threshold, the cooling device under test is determined to be the cooling device of interest. The relative temperature deviation is determined based on the degree of difference between the actual return air temperature and the preset return air threshold of the cooling equipment in question.

3. The method for dynamically controlling a data center cooling system based on AI according to claim 2, characterized in that, The process of identifying several risky cooling devices based on the relative temperature deviation, static pressure deviation, and rack load rate trends of the cooling devices under concern includes: Based on the relative temperature deviation, static pressure deviation, and cabinet load rate characteristics within a preset time period, the temperature-static pressure correlation, temperature-load correlation, and static pressure-load correlation are determined respectively. A correlation feature vector is constructed based on the temperature-static pressure correlation, the temperature-load correlation, and the static pressure-load correlation, and the anomaly probability of the correlation feature vector is determined based on the relative historical anomaly deviation of the correlation feature vector. When the anomaly probability is greater than a preset probability threshold, the cooling device of concern is determined to be the risk cooling device.

4. The method for dynamically controlling a data center cooling system based on AI according to claim 3, characterized in that, The process of determining the temperature-static pressure correlation, temperature-load correlation, and static pressure-load correlation based on the relative temperature deviation, static pressure deviation, and cabinet load rate characteristics within a preset time period includes: Calculate the correlation between the normalized relative temperature deviation and the normalized static pressure deviation to determine the temperature-static pressure correlation. Calculate the correlation between the normalized relative temperature deviation and the normalized rack load rate to determine the temperature-load correlation. The correlation between the normalized static pressure deviation and the normalized cabinet load rate is calculated to determine the static pressure load correlation.

5. The method for dynamically controlling a data center cooling system based on AI according to claim 4, characterized in that, The process of determining the anomaly probability of the associated feature vector based on its relative historical anomaly deviation includes: The associated feature vectors of all the cooling devices under interest within the previous preset normal operating time are recorded as historical associated vectors, and several historical associated distances are determined based on the deviation distance between each historical associated vector and the overall mean vector determined based on the historical associated vectors. The current association distance is determined based on the deviation distance between the associated feature vector and the overall mean vector, and the anomaly probability of the associated feature vector is determined based on the relative order of the current association distance among all the historical association distances arranged in ascending order.

6. The method for dynamically controlling a data center cooling system based on AI according to claim 5, characterized in that, The process of classifying the risky cooling equipment into over-response and under-response equipment based on a comparison of the relative temperature deviation of the risky cooling equipment and the load responsiveness threshold determined by the rack load rate includes: The preset heat transfer coefficient is determined based on the response coefficient of the cabinet load rate and the relative temperature deviation of the risk cooling equipment within a preset historical period. The load responsiveness is determined based on the cabinet load rate, the preset heat transfer coefficient, and the relative temperature deviation. When the load responsiveness is greater than a preset response threshold, the risk cooling device is determined to be an overloaded device, and when the load responsiveness is less than the preset response threshold, the risk cooling device is determined to be an underloaded device.

7. The method for dynamically controlling a data center cooling system based on AI according to claim 6, characterized in that, The process of generating a first control amount to adjust the preset cold water valve opening based on the thermal field imbalance and the cold aisle static pressure deviation of the overloaded device within a preset control time period includes: The positive deviation ratio of the thermal field and the positive deviation ratio of the static pressure are determined according to the proportion of the time when the thermal field imbalance and the static pressure deviation of the overloaded device are continuously positively deviated within the preset control time. The thermal field deviation and static pressure deviation are determined based on the average degree of continuous positive deviation of the thermal field imbalance and the cold aisle static pressure deviation of the overloaded device within the preset control time. The comprehensive deviation degree is determined by the weighted fusion result of the coupling characteristics of the positive deviation ratio of the thermal field, the deviation amount of the thermal field, the positive deviation ratio of the static pressure, and the deviation amount of the static pressure. The first control amount is determined based on the comprehensive deviation and the preset negative control coefficient.

8. The method for dynamically controlling a data center cooling system based on AI according to claim 7, characterized in that, The process of generating a second control quantity based on the actual return air temperature and the intake air temperature of the under-response load device includes: The insufficient temperature difference ratio is determined based on the proportion of times during which the real-time temperature difference between the actual return air temperature and the intake air temperature is continuously lower than the preset temperature difference threshold within the preset control time. The return air differential sequence and the intake air differential sequence are determined based on the real-time changes in the actual return air temperature and the intake air temperature within the preset control time. The consistency ratio is determined based on the consistency of the real-time changes in the real-time return air differential sequence and the intake air differential sequence. The product of the insufficient temperature difference ratio, the consistent change ratio, and the preset positive control coefficient is calculated to obtain the second control amount.

9. The method for dynamically controlling a data center cooling system based on AI according to claim 8, characterized in that, The process of generating a control command based on the difference between the basic control value and the first and second control values, respectively, and the actual return air temperature, includes: Based on the classification of the cooling equipment under test corresponding to the basic control quantity, the special control quantity is determined to be either the first control quantity or the second control quantity. The degree of control conflict is determined based on the temporal changes in the control directions of the basic control quantity and the dedicated control quantity within a preset adjustment period. The degree of amplitude conflict is determined based on the basic control amount and the dedicated control amount; The control command is determined based on the control conflict degree, the dedicated control amount, the actual return air temperature, and the amplitude conflict degree.

10. The method for dynamically controlling a data center cooling system based on AI according to claim 9, characterized in that, The process of determining the control command based on the control conflict degree, the dedicated control amount, and the amplitude conflict degree includes: Based on the threshold comparison results of the regulation conflict degree and the threshold comparison results of the amplitude conflict degree, a target regulation degree is generated according to the amplitude conflict degree, the basic regulation degree, and the dedicated regulation degree. Based on the comparison results of the threshold of the control conflict degree and the comparison results of the actual return air temperature threshold, the dedicated control amount is determined as the target control amount or the basic control amount is determined as the target control amount. The control command is generated based on the target control amount.

Citation Information

Patent Citations

  • Key cooling system of green data center

    CN121001312A

  • Water-cooling energy consumption self-adaptive regulation and control method suitable for data center

    CN121174454A

  • Systems and methods for cooling enclosure control and adaptive learning

    US20250008704A1