Deep reinforcement learning adaptive control method and system for air-liquid hybrid cooling system

By using a deep reinforcement learning adaptive control method, real-time data acquisition and processing of air-liquid hybrid cooling data are performed to construct a heat load feature vector and generate action commands. This solves the problem that air-liquid hybrid cooling systems cannot adapt to changes in data center load, and achieves efficient energy consumption management and intelligent temperature control.

CN121604369BActive Publication Date: 2026-04-03TIANJIN TIER TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-28
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing air-liquid hybrid cooling systems cannot adapt to the dynamic changes in data center loads, resulting in rigid power distribution, reliance on manual adjustments and fixed parameter settings, leading to inaccurate temperature control and energy waste.

Method used

A deep reinforcement learning adaptive control method is adopted. By collecting and preprocessing air-liquid mixed cooling data in real time, a heat load feature vector is constructed. Action commands are generated using a DRL action mapping model. Through feedback optimization and self-learning evolution control strategies, dynamic optimal allocation of air-liquid resources and intelligent collaborative control of multiple actuators are achieved.

Benefits of technology

It enables real-time, accurate judgment and dynamic perception of complex operating conditions, improves energy efficiency and temperature control accuracy, reduces the need for manual intervention, and enhances the intelligence and reliability of operation and maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121604369B_ABST
    Figure CN121604369B_ABST
Patent Text Reader

Abstract

This invention discloses a deep reinforcement learning adaptive control method and system for a hybrid air-liquid cooling system, belonging to the field of adaptive cooling control technology. It includes the following steps: S1, real-time acquisition of hybrid air-liquid cooling data, data preprocessing, assessment of equipment heat load, and preliminary adjustment; S2, determination of the proportion of air-cooling adjustment, construction of a DRL action mapping model, and generation of specific action commands; S3, execution of specific action commands, evaluation of action execution effect, model feedback optimization, and construction of a DRL interaction experience tuple set; S4, training the DRL interaction experience tuple set to drive optimized action commands, achieving self-learning and evolution of the hybrid air-liquid cooling system. This solves the problem that existing hybrid air-liquid cooling systems cannot adapt to dynamic changes in data center load, have rigid power distribution relying on manual intervention and fixed parameter settings, leading to inaccurate temperature control of multiple actuators and energy waste.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of adaptive cooling control technology, specifically to a deep reinforcement learning adaptive control method and system for air-liquid hybrid cooling systems. Background Technology

[0002] With the rapid development of next-generation data centers, high-performance computing, and cloud services, hybrid air-liquid cooling technology, with its combination of efficient heat dissipation and flexible allocation capabilities, has become one of the mainstream trends in the green and energy-saving upgrades of modern large-scale IT infrastructure. Faced with the demands of variable loads and refined energy efficiency management, single cooling modes and static strategies are no longer sufficient to meet the higher requirements of data centers for reliability, energy efficiency, and intelligence. At the same time, artificial intelligence methods such as reinforcement learning are rapidly emerging, integrating deep learning with intelligent decision-making to achieve dynamic perception, real-time optimization, and self-learning evolution driven by multi-source data.

[0003] However, existing air-liquid hybrid cooling systems cannot adapt to the dynamic changes in data center loads. Their rigid power distribution patterns, reliance on manual adjustments and fixed parameter settings, result in significant energy waste and inaccurate temperature control, hindering efficient collaboration and intelligent regulation of multiple actuators. Consequently, this not only leads to complex equipment maintenance, shortened lifespan, and high energy costs, but also limits the further upgrades and sustainable development of next-generation green and intelligent data centers.

[0004] Therefore, in order to address the above problems, there is an urgent need for a deep reinforcement learning adaptive control method and system for air-liquid hybrid cooling systems. Summary of the Invention

[0005] Technical problems to be solved

[0006] To address the shortcomings of existing technologies, this invention provides a deep reinforcement learning adaptive control method and system for air-liquid hybrid cooling systems. This solves the problem that existing air-liquid hybrid cooling systems cannot adapt to the dynamic changes in data center loads, have rigid power distribution that relies on manual intervention and fixed parameter settings, resulting in inaccurate temperature control of multiple actuators and energy waste.

[0007] Technical solution

[0008] To achieve the above objectives, the present invention provides the following technical solution: a deep reinforcement learning adaptive control method for a wind-liquid hybrid cooling system, comprising the following steps: S1, real-time acquisition of wind-liquid hybrid cooling data, preprocessing the data, evaluating the heat load of the equipment based on the data, constructing a heat load feature vector, and performing preliminary heat load risk adjustment; S2, receiving the wind-liquid hybrid cooling data and the heat load feature vector, determining the proportion of air cooling regulation in wind-liquid synergistic cooling, constructing a DRL action mapping model based on the proportion of air cooling regulation and the wind-liquid hybrid cooling data, and generating specific action commands; S3, executing the specific action commands, evaluating the execution effect after execution, optimizing the DRL action mapping model based on the execution effect, and constructing a DRL interaction experience tuple set; S4, training the DRL interaction experience tuple set to drive the optimization of action commands, thereby achieving self-learning and evolution of the control strategy for wind-liquid hybrid cooling.

[0009] Furthermore, the specific process of real-time acquisition and preprocessing of air-liquid mixed cooling data is as follows: Real-time acquisition of air-liquid mixed cooling data is achieved through temperature sensors, intelligent power distribution units (PDUs), flow sensors, pressure sensors, air volume sensors, and speed sensors. This data includes: chip temperature, cabinet temperature, coolant inlet temperature, coolant outlet temperature, coolant specific heat capacity, total power consumption, total coolant flow rate, total coolant pressure, total fan airflow, fan speed, and pump speed. Outliers in the air-liquid mixed cooling data are identified and removed using sliding window statistics and a 3x standard deviation rule. Missing points are filled using linear interpolation and historical mean regression. The data is then uniformly aligned to the main time axis and resampled. The air-liquid mixed cooling data is smoothed using window mean and exponentially weighted moving average, and then subjected to standard normalization scaling. An air-liquid mixed cooling database is established to store the original and preprocessed data, and the data is transmitted to the DRL controller.

[0010] Furthermore, the specific process for evaluating the heat load of the equipment based on air-liquid mixed cooling data is as follows: Based on a sliding time window, the air-liquid mixed cooling data at each moment within the window is acquired in real time. The coolant temperature difference is obtained by subtracting the coolant inlet temperature from the coolant outlet temperature and taking the absolute value. The instantaneous cooling capacity is obtained by multiplying the total coolant flow rate, the coolant specific heat capacity, and the coolant temperature difference. Within the entire sliding time window, the instantaneous cooling capacity at each moment is numerically integrated to obtain the total cooling capacity. The heat load risk coefficient is obtained by dividing the total power consumption of the equipment at the current moment by the total cooling capacity.

[0011] Furthermore, the specific process of constructing a heat load feature vector based on the equipment's heat load and performing preliminary heat load risk adjustment is as follows: The heat load risk coefficient is calculated in real time and compared with the heat load risk threshold. If the heat load risk coefficient is greater than the heat load risk threshold, preliminary heat load risk adjustment is performed: adjusting the fan, liquid pump, and valves to increase fan speed and coolant flow rate. If the heat load risk coefficient is less than or equal to the heat load risk threshold, continuous monitoring and regular recording are conducted to maintain the current control strategy and operating status. The coolant inlet temperature, coolant outlet temperature, total equipment power consumption, total coolant flow rate, and corresponding heat load risk coefficients are integrated to construct a heat load feature vector, which is then written into the air-liquid hybrid cooling database and proceeds to the next step.

[0012] Furthermore, the specific process of receiving air-liquid mixed cooling data and heat load feature vector to determine the proportion of air cooling regulation in air-liquid synergistic cooling is as follows: receiving air-liquid mixed cooling data and heat load feature vector in real time; adding the current total fan airflow, total coolant flow rate, and minimum constant value to obtain the total air-liquid flow rate; dividing the total fan airflow by the total air-liquid flow rate to obtain the proportion of air cooling capacity; multiplying the proportion of air cooling capacity by the heat load risk coefficient to obtain the weighted heat load pressure; dividing the weighted heat load pressure by the sum of the heat load risk coefficient and the minimum constant value to obtain the air-cooling synergistic ratio value.

[0013] Furthermore, based on the proportion of air-cooled regulation and the air-liquid hybrid cooling data, the specific process of constructing a DRL action mapping model and generating specific action commands is as follows: the air-cooled synergy ratio, thermal load risk coefficient, fan speed, pump speed, and chip temperature are used as the state mapping feature set. The state mapping feature set is trained by interacting with the simulation environment. The DRL action mapping model is constructed through deep deterministic policy gradient and DRL deep reinforcement learning algorithm. The action mapping module performs action mapping operations to generate specific action commands for each actuator of the fan, liquid pump, and valve. These commands are then output by the policy network and sent to each actuator.

[0014] Furthermore, the specific process of executing specific action instructions and evaluating the effect of action execution after the specific action instructions are executed is as follows: Specific action instructions are received from the DRL action mapping model; each actuator is synchronously adjusted according to the specific action instructions, and the adjusted air-liquid mixed cooling data is acquired in real time; the highest temperature value between the chip temperature and the rack temperature at the current moment is selected as the current overall maximum temperature; based on a sliding time window, the overall maximum temperature sequence is obtained, and the average value is calculated as the temperature control target value; the total power consumption of the device at the previous moment and the total power consumption of the device at the current moment are obtained, and the change in total power consumption is obtained by subtracting the total power consumption of the device at the previous moment from the total power consumption of the device at the current moment; simultaneously, within the sliding time window, the overall maximum temperature, total power consumption of the device, total coolant flow rate, and fan speed are monitored. The standard deviation of the total air volume is calculated separately. The standard deviations of each parameter are normalized. Principal component analysis is used to train the normalized standard deviations of each parameter. The coefficients of each parameter of the first principal component are extracted as weights. The standard deviations of the overall maximum temperature, total power consumption, total coolant flow, and total fan air volume are weighted and summed, and the negative value is taken to obtain the operational stability value. The absolute value of the difference between the current overall maximum temperature and the temperature control target value is calculated and taken as a negative value to obtain the temperature control deviation value. The change in total power consumption is multiplied by the power consumption change weight factor and the negative value is taken to obtain the power consumption adjustment value. The operational stability value is multiplied by the operational stability weight factor to obtain the operational stability value. The temperature control deviation value, power consumption adjustment value, and operational stability value are added together to obtain the model's comprehensive reward value.

[0015] Furthermore, the specific process of optimizing the DRL action mapping model based on the action execution effect and constructing the DRL interactive experience tuple set is as follows: After each specific action instruction is generated and executed, the model's comprehensive reward value is calculated and fed back to the DRL action mapping model to evaluate the merits of the specific action instruction. A random sampling optimization strategy and value network are adopted to improve the decision performance of the DRL action mapping model. The model's comprehensive reward value is continuously monitored. If the standard deviation of the model's comprehensive reward value is found to be greater than the fluctuation threshold within a certain period of time, it is determined to be an anomaly in temperature control, energy consumption, and stability. The fans and liquid pumps are switched to the maximum load operation mode to reduce the load on some non-core servers. The safety thresholds for temperature, pressure, and power consumption are increased. If the safety threshold is exceeded, load reduction and local cooling are forcibly executed. At the same time, an experience replay pool is constructed. The air-liquid mixed cooling data before execution, the action instruction, the model's comprehensive reward value, and the air-liquid mixed cooling data after execution are combined to construct the DRL interactive experience tuple set and stored in the experience replay pool and the air-liquid mixed cooling database.

[0016] Furthermore, the specific process of training the DRL interaction experience tuple set to drive the optimization of action commands and realize the self-learning and regulation strategy evolution of air-liquid hybrid cooling is as follows: continuously sampling the DRL interaction experience tuple set from the experience replay pool as training samples, inputting iterative updates of the DRL algorithm's policy network and value network through trial and error, feedback, and learning. The policy network adjusts the mapping between the state mapping feature set and specific action commands based on sample feedback to achieve dynamic optimization of policy parameters. Each round of training is based on the model's comprehensive reward value, guiding the network to continuously converge toward the target policy of achieving temperature target, low energy consumption, and stable operation. When the policy network converges and the comprehensive reward value is maximized, the current optimal policy is solidified into an online deployment model to drive the regulation decision of air-liquid hybrid cooling in real time. At the same time, the model's comprehensive reward value is evaluated periodically. If the standard deviation of the model's comprehensive reward value exceeds the fluctuation threshold or the policy fails, it reverts to the historical optimal policy.

[0017] The second aspect of this invention provides a deep reinforcement learning adaptive control system for a wind-liquid hybrid cooling system, comprising: a multi-source data acquisition and risk assessment module, used to acquire wind-liquid hybrid cooling data in real time, preprocess the wind-liquid hybrid cooling data, assess the heat load of the equipment based on the wind-liquid hybrid cooling data, construct a heat load feature vector according to the heat load of the equipment, and perform preliminary heat load risk adjustment; a DRL adaptive decision-making and action mapping module, used to receive wind-liquid hybrid cooling data and heat load feature vector, determine the proportion of air cooling regulation in wind-liquid synergistic cooling, construct a DRL action mapping model according to the proportion of air cooling regulation and wind-liquid hybrid cooling data, and generate specific action commands; an actuator linkage and real-time feedback module, used to execute specific action commands, evaluate the action execution effect after the specific action commands are executed, optimize the DRL action mapping model based on the action execution effect, and construct a DRL interactive experience tuple set; and a closed-loop self-learning and self-evolution module, used to drive the optimization of action commands by training the DRL interactive experience tuple set, thereby realizing the self-learning and control strategy evolution of wind-liquid hybrid cooling.

[0018] Beneficial effects

[0019] The present invention has the following beneficial effects:

[0020] (1) This invention achieves comprehensive dynamic perception of operating conditions by real-time acquisition and high-quality preprocessing of multi-dimensional data of chips, cabinets, coolant and fans, and constructs a sliding window integral and thermal load risk coefficient assessment mechanism to provide real-time and accurate judgment basis for whether the cooling capacity is sufficient and the risk of overheating, thereby improving the adaptability and safety of complex operating conditions.

[0021] (2) This invention utilizes the DRL algorithm to input multiple features such as air-cooling and liquid-cooling ratio, heat load risk, and actuator status into the state space, thereby realizing the dynamic optimal allocation of air and liquid resources and intelligent collaborative control of multiple actuators. This breaks through the limitations of traditional rigid allocation and manual setting, and greatly improves energy efficiency and temperature control accuracy.

[0022] (3) This invention constructs a multi-objective weighted model of temperature control deviation, energy consumption change and operational stability to comprehensively reward value, thereby driving the intelligent control strategy to continuously learn and adapt and evolve, continuously improving the temperature control compliance rate, energy consumption efficiency and operational stability, and realizing long-term, closed-loop intelligent optimization.

[0023] (4) This invention, through abnormal fluctuation identification and self-healing strategy, once abnormal temperature control, energy consumption and stability are detected, switches to maximum cooling capacity and load reduction emergency mode, and links multiple actuators to actively respond, greatly reducing the need for manual intervention and improving the intelligence and reliability of operation and maintenance.

[0024] Of course, any product implementing this invention does not necessarily need to achieve all of the advantages described above at the same time. Attached Figure Description

[0025] Figure 1 Flowchart of a deep reinforcement learning adaptive control method for a wind-liquid hybrid cooling system;

[0026] Figure 2 A module diagram of a deep reinforcement learning adaptive control system for a wind-liquid hybrid cooling system;

[0027] Figure 3 Flowchart of deep reinforcement learning adaptive control for air-liquid hybrid cooling system;

[0028] Figure 4 A trend chart of air-cooled synergistic ratio values;

[0029] Figure 5 Schematic diagram of the training principle of DRL action mapping model;

[0030] Figure 6 This is a diagram of the deep reinforcement learning adaptive control system architecture for a wind-liquid hybrid cooling system.

[0031] In the diagram, 1 is the temperature sensor; 2 is the flow sensor; 3 is the DRL controller; 4 is the motion mapping; 5 is the liquid pump; and 6 is the fan. Detailed Implementation

[0032] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. As those skilled in the art will understand, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0033] Please see Figures 1-6 This invention provides a technical solution: a deep reinforcement learning adaptive control method and system for a wind-liquid hybrid cooling system, such as... Figure 1 As shown, the process includes the following steps: S1, real-time acquisition of air-liquid hybrid cooling data, data preprocessing of the air-liquid hybrid cooling data, assessment of the equipment's heat load based on the air-liquid hybrid cooling data, construction of a heat load feature vector based on the equipment's heat load, and preliminary heat load risk adjustment; S2, receiving air-liquid hybrid cooling data and the heat load feature vector, determining the proportion of air cooling regulation in air-liquid synergistic cooling, constructing a DRL action mapping model based on the proportion of air cooling regulation and the air-liquid hybrid cooling data, and generating specific action commands; S3, executing the specific action commands, evaluating the action execution effect after the specific action commands are executed, optimizing the DRL action mapping model based on the action execution effect, and constructing a DRL interaction experience tuple set; S4, training the DRL interaction experience tuple set to drive the optimization of action commands, realizing the self-learning and control strategy evolution of air-liquid hybrid cooling.

[0034] Specifically, the real-time acquisition and preprocessing of air-liquid mixed cooling data is as follows: Feedback from temperature sensor 1, intelligent power distribution unit (PDU), flow sensor 2, pressure sensor, air volume sensor, and speed sensor (such as Hall effect sensors and encoders) allows for the real-time measurement of the actual rotational speeds of the fan 6 and liquid pump 5. This real-time acquisition of air-liquid mixed cooling data includes: chip temperature, cabinet temperature, coolant inlet temperature, coolant outlet temperature, coolant specific heat capacity, total power consumption, total coolant flow rate, total coolant pressure, total fan air volume, fan speed, and pump speed. Outliers in the air-liquid mixed cooling data are identified and removed using a sliding window statistical method and a 3x standard deviation rule. This involves using a sliding window to statistically analyze each parameter in real-time. The values ​​and standard deviations are used to identify and remove data that exceed the mean ± 3σ range, thus ensuring the robustness and reliability of the data. Linear interpolation and historical mean regression are used to fill in missing points. If a single point is missing, interpolation of preceding and following values ​​is used. For multiple points and long missing segments, historical operating data is combined for mean and regression filling. The data is then uniformly aligned to the main time axis and resampled. The air-liquid mixed cooling data is smoothed using window mean and exponentially weighted moving average. Window mean is used to eliminate instantaneous fluctuations, while exponentially weighted moving average smoothing can respond more quickly to new trends and smooth out abnormal jitter. Standard normalization scaling is performed using the minimum-maximum normalization method. An air-liquid mixed cooling database is established to store the original and pre-processed air-liquid mixed cooling data, and the air-liquid mixed cooling data is transmitted to the DRL controller 3.

[0035] like Figure 3The diagram shows the deep reinforcement learning adaptive control flowchart of the air-liquid hybrid cooling system. It illustrates the core closed-loop structure of the air-liquid hybrid cooling regulation based on DRL deep reinforcement learning. The core consists of key operational data collected in real time by temperature sensor 1 and flow sensor 2, including chip temperature, cabinet temperature, coolant inlet temperature, coolant outlet temperature, and total coolant flow rate. Simultaneously, it collects total power consumption, total fan airflow, fan speed, and pump speed, integrating these data as state inputs and transmitting them to the DRL controller 3. Upon receiving the latest sensor data, the DRL controller 3 calculates the proportion of air cooling in the air-liquid hybrid cooling system and inputs this as a state feature into the deep reinforcement learning model of the DRL action mapping model. Through the built-in intelligent policy network, the DRL controller 3 performs action mapping 4 on the current state using the action mapping module, outputting the optimal control decision, i.e., generating specific action commands for each actuator, including fan 6, liquid pump 5, and valves. The core actuators, liquid pump 5 and fan 6, adjust their operating states, such as speed and flow rate, according to the specific action commands, linking air cooling and liquid cooling capabilities to achieve dynamic and coordinated optimization of cooling capacity allocation. New sensor data is collected and fed back to DRL controller 3, forming a closed loop of sensing, decision-making, execution, feedback, and re-sensing. This ensures real-time response to dynamic changes in load and environment, and continuously optimizes temperature control accuracy and energy efficiency.

[0036] In this implementation scheme, high-precision, real-time data acquisition of all parameters across the entire air-liquid hybrid cooling process is achieved through multiple types of sensors. In the data preprocessing stage, outliers are efficiently removed using sliding window statistics and a 3x standard deviation rule, while linear interpolation and historical means are used to fill in missing data, ensuring data continuity and reliability. Unified alignment and resampling resolve the issue of asynchronous data from multiple sources. Window mean, exponentially weighted average, and smoothing normalization further enhance data robustness and algorithm adaptability. Finally, all raw and high-quality preprocessed data are uniformly stored and managed, providing an accurate, traceable, unbiased, and high-value data foundation for subsequent intelligent control models, significantly improving anomaly robustness, the effectiveness of intelligent decision-making, and the feasibility of engineering implementation.

[0037] Specifically, the process of evaluating the heat load of equipment based on air-liquid mixed cooling data is as follows: Based on a sliding time window, the air-liquid mixed cooling data at each moment within the window is acquired in real time. The coolant temperature difference is obtained by subtracting the coolant inlet temperature from the coolant outlet temperature and taking the absolute value. The absolute value of the coolant temperature difference can accurately reflect the heat exchange effect. The instantaneous cooling capacity is obtained by multiplying the total coolant flow rate, the coolant specific heat capacity, and the coolant temperature difference. This is the amount of heat that can be removed instantaneously. Within the entire sliding time window, the instantaneous cooling capacity at each moment is numerically integrated to obtain the total cooling capacity. The numerical integration calculation is the sum of the instantaneous cooling capacity at all sampling moments within the sliding window, which is equivalent to the total heat actually removed during that time period. The heat load risk coefficient is obtained by dividing the total power consumption of the equipment at the current moment by the total cooling capacity. The heat load risk coefficient reflects the degree of matching between the heat load and the cooling capacity. The larger the value, the greater the load pressure and the smaller the cooling margin.

[0038] The specific formula for the heat load risk coefficient is as follows:

[0039] ;

[0040] In the formula, This represents the heat load risk coefficient, used to dynamically evaluate the heat load risk level of a data center within a specific time period. It measures the ratio of the heat generated by the current equipment to the heat that the cooling system can actually remove during that time window, thereby determining whether the cooling capacity is sufficient and whether there is a risk of hot spots. A large value indicates a high thermal load pressure and relatively insufficient cooling capacity, which may pose a risk of overheating. A small size indicates sufficient cooling capacity and a large margin for operational safety. It represents the total power consumption of the device at the current moment, and the heat generated by all devices at the current moment, which is the direct source of heat load; It indicates the total flow rate of coolant and measures the ability of coolant to remove heat per unit time. This indicates the specific heat capacity of the coolant, which determines how much heat a unit mass of coolant can remove when absorbing a unit temperature difference. Indicates the coolant outlet temperature; Indicates the coolant inlet temperature; This indicates the coolant temperature difference, reflecting the actual amount of heat removed by the coolant. The greater the coolant temperature difference, the stronger the heat exchange. It represents the total cooling capacity, reflecting the total amount of heat that can be removed within the sliding time window.

[0041] In this implementation scheme, a sliding time window is used to dynamically calculate the coolant temperature difference, total flow rate, and specific heat capacity, accurately determining the instantaneous cooling capacity at each moment. The total cooling capacity is then obtained through integration, accurately reflecting the actual heat dissipation capacity over a period of time. Combining the current total power consumption of the equipment with the total cooling capacity yields a thermal load risk coefficient, which quantitatively assesses the balance between the current load and cooling capacity, enabling real-time dynamic evaluation of thermal safety margin. This method not only improves the accuracy and predictability of thermal risk identification but also provides scientific risk criteria for subsequent intelligent control, enhancing adaptability and operational safety under complex and fluctuating loads.

[0042] Specifically, the process of constructing a heat load feature vector based on the equipment's heat load and performing preliminary heat load risk adjustment is as follows: The heat load risk coefficient is calculated in real time and compared with the heat load risk threshold. If the heat load risk coefficient is greater than the heat load risk threshold, preliminary heat load risk adjustment is performed: adjusting fan 6, liquid pump 5, and valves to increase the fan 6 speed and coolant flow rate; simultaneously, according to the valve opening linkage rules, including: proportionally increasing the valve opening of the main cooling circuit and branch pipes based on the extent to which the heat load risk coefficient exceeds the heat load risk threshold, prioritizing increasing the fluid throughput capacity in areas with higher temperature and flow rates, and optimizing the cooling circuit. The system allocates resources to improve cooling efficiency and ensures maximum synergistic gain across all stages of the air-liquid hybrid cooling process. If the heat load risk coefficient is less than or equal to the heat load risk threshold, continuous monitoring and regular recording are performed to maintain the current control strategy and operating status. This means keeping the current parameters of fan 6, liquid pump 5, and valves unchanged, maintaining a steady-state energy-saving mode, and dynamically collecting operating data for subsequent analysis and self-learning. The system integrates the coolant inlet temperature, coolant outlet temperature, total equipment power consumption, total coolant flow rate, and corresponding heat load risk coefficients to construct a heat load feature vector. This vector serves as a high-dimensional information summary of the current operating condition of the air-liquid hybrid cooling system, is written into the air-liquid hybrid cooling database, and then proceeds to the next step.

[0043] In this implementation scheme, the high-heat-risk state is instantly identified and responded to by calculating the heat load risk coefficient in real time and intelligently comparing it with a threshold. When the risk coefficient exceeds the threshold, the parameters of key actuators such as fan 6, liquid pump 5, and valves are adjusted to improve cooling capacity and effectively prevent overheating risks. When the heat load risk coefficient is within a safe range, continuous monitoring and routine recording ensure efficient and stable operation. Simultaneously, by integrating coolant inlet and outlet temperatures, total power consumption, total flow rate, and heat load risk coefficient, a comprehensive heat load feature vector is constructed, providing a rich and accurate feature foundation for subsequent regulation and self-learning processes.

[0044] Specifically, the process of receiving air-liquid mixed cooling data and heat load feature vectors to determine the proportion of air cooling regulation in air-liquid synergistic cooling is as follows: Real-time reception of air-liquid mixed cooling data and heat load feature vectors; addition of the current total fan airflow, total coolant flow rate, and a minimum constant value to obtain the total air-liquid flow rate, where the minimum constant value is 0.001 to avoid calculation anomalies caused by a zero denominator; division of the total fan airflow by the total air-liquid flow rate to obtain the air cooling capacity proportion, reflecting the contribution of air cooling to the air-liquid mixed cooling resources; multiplication of the air cooling capacity proportion by the heat load risk coefficient to obtain the weighted heat load pressure, which reflects the strength of the air cooling capacity's role in the actual heat load pressure; division of the weighted heat load pressure by the sum of the heat load risk coefficient and the minimum constant value to obtain the air cooling synergistic allocation value, representing the synergistic regulation proportion that air cooling resources should undertake, for subsequent intelligent decision-making and actuator linkage.

[0045] The specific formula for the air-cooled synergistic ratio is as follows:

[0046] ;

[0047] In the formula, This represents the air-cooling synergy ratio, used to dynamically determine the optimal proportion of air cooling in air-liquid hybrid cooling. Based on real-time air-cooling and liquid cooling capacities, as well as the heat load risk level, the air-cooling and liquid cooling ratio is adaptively adjusted to ensure optimal safety and energy efficiency. The larger the value, the higher the proportion of air cooling. This indicates the total airflow of the fan, which measures the ability of the air cooler to remove heat at this time. This indicates the total flow rate of the coolant, measuring the ability of the liquid coolant to remove heat at this point. This represents the heat load risk coefficient, which measures the pressure ratio between the current equipment's heat load and its cooling capacity. The higher the risk, the larger the value. This represents a very small constant value to prevent the denominator from being zero and to ensure the stability of the algorithm; its value is 0.001. This indicates the proportion of air cooling capacity, reflecting the weight of air cooling in the entire cooling system at this moment. The larger the value, the more prominent the air cooling capacity. This indicates the weighted thermal load pressure, allowing the air-cooling ratio and risk to be dynamically linked, balancing capacity and safety. When the air-cooling capacity is stronger and the thermal risk is higher, the priority of air-cooling allocation is increased.

[0048] In this embodiment, Table 1 is a data table of air-cooling synergistic ratio values. It details the total fan airflow, total coolant flow rate, heat load risk coefficient, and air-cooling synergistic ratio values ​​at five different times. The minimum constant value is 0.001. At time 1, the total fan airflow is 3200, the total coolant flow rate is 2400, the heat load risk coefficient is 0.70, and the air-cooling synergistic ratio is 0.5706. At time 2, the total fan airflow is 2900, the total coolant flow rate is 2200, the heat load risk coefficient is 0.95, and the air-cooling synergistic ratio is 0.5680. At time 3, the total fan airflow is 4000, the total coolant flow is 2100, the heat load risk factor is 1.20, and the air-cooling synergy ratio is 0.6552; at time 4, the total fan airflow is 2100, the total coolant flow is 3100, the heat load risk factor is 0.55, and the air-cooling synergy ratio is 0.4031; at time 5, the total fan airflow is 2500, the total coolant flow is 2000, the heat load risk factor is 0.88, and the air-cooling synergy ratio is 0.5549.

[0049] Table 1. Data on the synergistic ratio of air cooling systems

[0050]

[0051] like Figure 4 The graph shown is a trend chart of the air-cooling synergistic ratio. The horizontal axis represents the time number, and the vertical axis represents the air-cooling synergistic ratio at each time point; it illustrates the specific values ​​and trends of the proportion of air cooling in the hybrid cooling system as the operating conditions change. (Based on Table 1 and...) Figure 3 As can be seen, the overall trend initially fluctuates slowly, then reaches a peak at time 3, followed by a significant decline at time 4, and finally rebounds. This reflects the dynamic adjustment process of the proportion of air cooling in the total hybrid cooling capacity under different load and distribution conditions. The curve reaches its highest value at time 3, indicating that air cooling accounts for the largest proportion at this time. The significant decline at time 4 indicates that under certain conditions such as high pump flow and low air volume, the proportion of air cooling will be reduced in time, and the participation of liquid cooling will be increased, achieving intelligent and coordinated control of hybrid cooling. Through the dynamic change of the coordinated ratio, energy efficiency and temperature control accuracy can be balanced under various complex loads, avoiding insufficient temperature control and energy waste caused by single-mode operation.

[0052] In this implementation scheme, by real-time fusion of total fan airflow, total coolant flow, and heat load characteristic vectors, and using mathematical weighting and normalization, the actual capacity ratio of air cooling regulation in air-cooled hybrid cooling is accurately calculated. Simultaneously, the air cooling capacity ratio is organically combined with the heat load risk coefficient to quantify an air-cooling synergistic ratio value reflecting current demand and air cooling contribution. This effectively avoids the rigidity of single control and parameter settings, enhancing the flexible allocation capability of cooling resources. This method improves the dynamic adaptability and intelligent synergy of air-cooled hybrid cooling under varying loads and complex operating conditions, providing a scientific and quantifiable criterion for subsequent intelligent control decisions and energy-saving optimization.

[0053] Specifically, based on the proportion of air-cooled regulation and air-liquid mixed cooling data, the process of constructing a DRL action mapping model and generating specific action commands is as follows: The air-cooled synergy ratio, thermal load risk coefficient, fan speed, pump speed, and chip temperature are used as a state mapping feature set, comprehensively reflecting the heat and cold distribution, load pressure, and operating status of the main actuators. This serves as a key input for subsequent intelligent control. The state mapping feature set is trained through interaction with the simulation environment. The simulation model is identified based on measured data, utilizing an ARX autoregressive external model and an LSTM long short-term memory neural network data-driven method, combined with a first-order RC thermal network mechanism model. The model coupling is established to dynamically and accurately reflect the thermal response and dynamic behavior of air-liquid mixing cooling under different actions, providing real and effective environmental interaction feedback for deep reinforcement learning algorithms. A DRL action mapping model is constructed through deep deterministic policy gradient and DRL deep reinforcement learning algorithm, which can handle continuous action space and has end-to-end mapping capability. The action mapping module performs action mapping operation 4 to generate specific action commands for each actuator of fan 6, liquid pump 5 and valve, such as fan PWM speed regulation signal, pump frequency conversion set value, valve opening command, etc., and output by policy network and sent to each actuator.

[0054] In this implementation plan, a high-dimensional, real-time-aware state feature set is constructed by integrating air-cooling coordination ratio values, heat load risk coefficients, and key equipment status parameters. This set serves as input to drive a deep reinforcement learning algorithm model centered on DDPG. Through continuous interactive training with the simulation environment, the system can learn and optimize the control strategies of multiple actuators, including fan 6, liquid pump 5, and valves, effectively achieving intelligent mapping from complex states to optimal action commands. This not only improves the targeting and real-time performance of control commands but also realizes end-to-end intelligent control of multi-actuator coordinated adjustment, enhancing the adaptability to dynamic operating conditions and operational efficiency. This provides solid support for the efficient, safe, and intelligent upgrade of data center cooling systems.

[0055] Specifically, the process of executing specific action instructions and evaluating the effect of action execution after the specific action instructions are executed is as follows: Receive the specific action instructions output by the DRL action mapping model; each actuator synchronously adjusts according to the specific action instructions and acquires the adjusted air-liquid mixed cooling data in real time to reflect the actual operating condition response of the current control action; select the highest temperature value between the chip temperature and the cabinet temperature at the current moment as the current overall maximum temperature; based on a sliding time window, obtain the overall maximum temperature sequence and calculate the average value as the temperature control target value to dynamically adapt to the optimal temperature control range under different loads; obtain the total power consumption of the equipment at the previous moment and the total power consumption of the equipment at the current moment; subtract the total power consumption of the equipment at the previous moment from the total power consumption of the equipment at the current moment to obtain the change in total power consumption, which is used to capture short-term fluctuations in load and energy consumption, providing a quantitative basis for energy-saving control and energy consumption evaluation; simultaneously, within the sliding time window, calculate the standard deviation of the overall maximum temperature, total power consumption of the equipment, total coolant flow rate, and total fan airflow, respectively. The standard deviation measures the volatility of each parameter within the time period, reflecting the dynamic stability of operation. Normalize the standard deviation of each parameter to facilitate subsequent weighting and comparison, using principal component analysis... The algorithm is trained on the normalized standard deviations of each parameter, and the coefficients of each parameter in the first principal component are extracted as weights. PCA can extract the weights that reflect the most sensitive direction of overall fluctuations. By scientifically integrating multi-parameter features, the accuracy of the evaluation is enhanced. Based on the weighted weights of each parameter, the standard deviations of the overall maximum temperature, total equipment power consumption, total coolant flow rate, and total fan airflow are weighted, summed, and the negative value is taken to obtain the operational stability value. Taking a negative value ensures that the smaller the fluctuation, the higher the reward, which is conducive to guiding the cooling towards adaptive optimization towards low fluctuation and stable operation. The difference between the current overall maximum temperature and the temperature control target value is calculated and taken. The absolute value, taken as negative, yields the temperature control deviation value, reflecting whether the temperature meets the standard. The smaller the absolute value, the more precise the control; a negative value also indicates a larger reward. The total power consumption change is multiplied by the power consumption change weighting factor, and the negative value is taken to obtain the power consumption adjustment value, reinforcing the energy-saving orientation. The operating stability value is multiplied by the operating stability weighting factor to obtain the operating stability value. The temperature control deviation value, power consumption adjustment value, and operating stability value are added together to obtain the model's comprehensive reward value, which reflects the actual effect of the current control action under multiple objectives of temperature control compliance, energy consumption optimization, and fluctuation stabilization, providing direct feedback for subsequent DRL self-learning and strategy optimization.

[0056] The specific formula for the model's overall reward value is as follows:

[0057] ;

[0058] In the formula, The comprehensive reward value represents the overall reward function of the DRL action mapping model in air-liquid hybrid cooling. It is used to evaluate the temperature control accuracy, energy consumption optimization, and operational stability under each decision, and to drive the DRL action mapping model to learn the comprehensive optimal strategy that is temperature stable, energy-efficient, and stable in operation. It indicates the current overall highest temperature, reflecting the point with the highest temperature in real time. It is the most critical indicator for temperature control safety and is used to assess whether there is a risk of overheating. This represents the target temperature value, indicating the desired temperature threshold to be reached and maintained, serving as a reference benchmark for judging whether the temperature control meets the standard. It represents the change in total power consumption, reflects the real-time change in total energy consumption, and evaluates the effectiveness of the control strategy in energy saving. A decrease in energy consumption is a positive signal, and an increase in energy consumption is a negative signal. The value represents the operational stability, reflecting the short-term fluctuations of key parameters including the overall maximum temperature, total power consumption of the equipment, total coolant flow rate, and total fan airflow. The smaller the fluctuation and the more stable the operation, the higher the operational stability value and the greater the model's overall reward value. This represents the temperature control deviation value, reflecting the degree of deviation between the current highest temperature and the target temperature. The smaller the deviation, the better the temperature control effect. Taking a negative value makes the temperature control deviation value smaller and the reward greater. This represents the power consumption adjustment value, quantifying the impact of each action on energy consumption. As energy consumption increases, the power consumption adjustment value decreases, which reduces the reward. This represents the stable operating value, reflecting the degree of operational stability. The smaller the fluctuation, the more stable the operation, and the greater the reward. The power consumption change weight factor is represented by a historical sample of the operation. The initial value of the power consumption change weight factor is set to 1. The temperature control deviation value, power consumption adjustment value and stable operation value are normalized respectively. The power consumption change weight factor and the stable operation weight factor are adjusted so that the average contribution of the temperature control deviation value, power consumption adjustment value and stable operation value in the model's comprehensive reward value is consistent, ensuring that the temperature control, energy saving and stability objectives are optimally balanced. The current power consumption adjustment value is taken as the optimal power consumption change weight factor, with a value range between 0.1 and 5. This represents the operational stability weight factor. Based on a historical sample of operation, the initial value of the operational stability weight factor is set to 1. The temperature control deviation value, power consumption adjustment value, and operational stability value are normalized respectively. The power consumption change weight factor and the operational stability weight factor are adjusted to ensure that the average contribution of the temperature control deviation value, power consumption adjustment value, and operational stability value in the model's comprehensive reward value is consistent, ensuring that the balance between temperature control, energy saving, and stability is optimal. The weight of the current operational stability value is taken as the optimal operational stability weight factor, with a value range between 0.1 and 5.

[0059] This implementation scheme achieves multi-dimensional and dynamic quantitative feedback of control results through a full-process evaluation of the actuator's action effects. It not only collects operational data after adjustment in real time, but also extracts the weights of each parameter based on sliding window and principal component analysis algorithms, enabling automated and objective evaluation of three key indicators: temperature control target, energy consumption change, and operational stability. Through normalization and weighted processing, the sensitivity and scientific rigor to complex state fluctuations and synergistic effects are effectively improved. Ultimately, temperature control accuracy, energy consumption optimization, and operational stability are integrated into a single comprehensive reward value, providing high-value, quantifiable performance feedback for subsequent self-learning and strategy optimization, thereby enhancing the adaptive optimization capability and operational reliability of the air-liquid hybrid cooling system.

[0060] Specifically, the process of optimizing the DRL action mapping model based on the action execution effect and constructing a DRL interaction experience tuple set is as follows: After each specific action instruction is generated and executed, the model's comprehensive reward value is calculated and fed back to the DRL action mapping model to evaluate the merits of the specific action instruction. That is, the policy network and value network are updated with the current comprehensive reward value, enabling the model to adjust its future decision-making tendency according to different behavioral effects. Random sampling is used to optimize the policy and value network, improving the decision-making performance of the DRL action mapping model. Random sampling refers to randomly extracting historical experience samples to break data correlation. The value network is a neural network used to evaluate the long-term benefits of the current state and action combination, and is the core component of the DRL deep reinforcement learning algorithm. Through continuous trial and error, feedback, learning, and optimization, the DRL action mapping model can continuously improve its autonomous decision-making and coping capabilities under complex conditions. The comprehensive reward value of the model is continuously monitored. If the comprehensive reward value of the model is found to be within a certain time window, the model will be evaluated. If the standard deviation of the value exceeds the fluctuation threshold, it is judged as an anomaly in temperature control, energy consumption, and stability. The standard deviation serves as a volatility criterion, reflecting whether the model's overall reward value is stable. Large fluctuations may indicate operational anomalies, strategy failures, and sudden events. Fan 6 and liquid pump 5 are switched to maximum load operation mode, i.e., fans at full speed and pumps at maximum flow rate, to cope with high heat risks and ensure equipment safety. The load on some non-core servers is reduced. The safety thresholds for temperature, pressure, and power consumption are increased. If the safety thresholds are exceeded, load reduction and local cooling are forcibly executed to achieve rapid self-healing. At the same time, an experience replay pool is constructed. The pre-execution air-liquid hybrid cooling data, action instructions, model overall reward value, and post-execution air-liquid hybrid cooling data are combined to construct a DRL interactive experience tuple set, i.e., state, action, reward, and new state, for model training and optimization. The tuples are stored in the experience replay pool and the air-liquid hybrid cooling database. The experience replay pool is the core storage structure of the reinforcement learning algorithm, which can randomly sample and repeatedly utilize historical experience to improve self-learning efficiency and generalization ability.

[0061] In this implementation plan, by promptly feeding back the model's overall reward value after each action execution to the DRL action mapping model, continuous self-optimization and dynamic improvement of the model's decision-making capabilities are achieved. An experience replay pool mechanism is employed to continuously accumulate and utilize rich experience tuples of states, actions, rewards, and new states, enhancing the model's generalization ability and self-learning efficiency under varying operating conditions. Simultaneously, by monitoring fluctuations in the model's overall reward value in real time, anomalies in temperature control, energy consumption, and stability can be identified, triggering maximum load operation, load degradation, and local cooling emergency measures, and dynamically raising safety thresholds, effectively ensuring safety and robustness. This process enhances adaptive adjustment and fault self-healing capabilities under extreme operating conditions, providing a solid guarantee for long-term efficient, safe, and intelligent operation.

[0062] Specifically, the process of training the DRL interaction experience set to drive the optimization of action commands and achieve self-learning and evolution of the control strategy for air-liquid hybrid cooling is as follows: The DRL interaction experience set is continuously sampled from the experience replay pool, i.e., samples are randomly drawn from historical data to break data correlation and improve the stability and generalization ability of self-learning. These samples are then used as training samples and input into the policy network and value network of the DRL algorithm for iterative updates through trial and error, feedback, and learning. The policy network is responsible for outputting the optimal action command based on the current state, while the value network evaluates the comprehensive reward of the action in the long term. Both are alternately optimized through the Bellman equation in reinforcement learning to continuously improve the control strategy. Based on sample feedback, the policy network adjusts the mapping between the state mapping feature set and the specific action command, achieving dynamic optimization of the policy parameters, thus improving the DRL action mapping model. It can learn to select the optimal actuator control action under different operating conditions, thereby balancing temperature control, energy saving, and stable operation. Each training round is based on the model's comprehensive reward value, guiding the network to continuously converge towards the target strategy of achieving the target temperature, low energy consumption, and stable operation. When the strategy network converges and the comprehensive reward value is maximized, the current optimal strategy is solidified into an online deployment model. That is, the learned optimal parameter network is used for real-time control decisions, improving operational efficiency and driving the control decisions of air-liquid mixed cooling in real time. At the same time, the model's comprehensive reward value is evaluated periodically. If the standard deviation of the model's comprehensive reward value exceeds the fluctuation threshold or the strategy fails, it reverts to the historical optimal strategy. If the standard deviation of the model's comprehensive reward value exceeds the fluctuation threshold, it indicates that the DRL action mapping model is unstable. Strategy failure may cause abnormal temperature control and energy consumption. It then switches back to the model parameters with the best recent historical performance to ensure operational safety and robustness.

[0063] like Figure 5The diagram illustrates the training principle of the DRL action mapping model. First, various real-time operating conditions are used as the current state S. The DRL controller 3 receives state S and inputs it into the DRL deep reinforcement learning algorithm for training. Based on the current state, a reinforcement learning model for DRL action mapping is constructed, generating the optimal control action A, which is the specific control command for the fan 6, liquid pump 5, and valve actuators. The specific control commands are issued to each actuator by the action mapping module, achieving dynamic allocation of air-liquid hybrid cooling capacity. After executing action A, the model's comprehensive reward value R is calculated based on a multi-objective evaluation standard of temperature control deviation, energy consumption change, and operational fluctuation, used to quantify the quality of the current control effect. The model's comprehensive reward value R is fed back to the DRL deep reinforcement learning model for trial and error, feedback, and learning loops, achieving continuous optimization and self-learning of the strategy. This forms a closed-loop control system of state perception, intelligent decision-making, action execution, reward feedback, and self-learning optimization, continuously improving the adaptability to dynamic loads in the data center and overall energy efficiency, achieving efficient, intelligent, and self-evolving operation of the air-liquid hybrid cooling system. The specific meanings of each parameter in the process shown in the diagram are as follows: This represents multi-dimensional operating data of the air-liquid hybrid cooling system in its initial stage. This represents the regulatory action in the original state. Indicates by The overall reward value of the model after execution. Indicate execution The new status is then reported back by the system. This indicates the system state during the first round of interactive trial and error. According to The specific action instructions generated, for The overall reward value of the model after execution. Indicates the execution of an action The new state that is subsequently obtained; , , , This refers to the status, actions, rewards, and new status in subsequent rounds. (The above...) , , , The quadruplets accumulate continuously, driving the DRL reinforcement learning model to achieve continuous trial and error, feedback, learning and optimization, thereby achieving the goal of accurate decision-making and self-evolution.

[0064] like Figure 6The diagram shows the architecture of a deep reinforcement learning adaptive control system for a hybrid air-liquid cooling system. First, the perception layer collects real-time operating data through a sensor network, forming a multi-dimensional state S that fully reflects the data center's operational status. State S is input to the DRL controller 3 in the decision layer for data processing, feature extraction, and calculation of the air-cooling adjustment ratio. Based on the current state S, it generates the optimal control action A, such as the target speed of fan 6 and the target flow rate of liquid pump 5. These action commands are then sent to various actuators in the hybrid air-liquid cooling system. The actuators adjust their operating parameters in real-time according to the action commands, driving responses to changes in the environment and load, thereby affecting the cooling effect and energy consumption. After action A is executed, the control effect is quantitatively evaluated using the model's comprehensive reward value R, based on indicators such as current temperature control deviation, energy consumption changes, and operational stability. A higher reward value indicates more precise temperature control, lower energy consumption, and more stable operation. The model's comprehensive reward value R serves as feedback for the policy network training, guiding the DRL reinforcement learning model to continuously adjust parameters and optimize strategies. State S, action A, reward R, and new state S' are stored in the experience replay pool. By training on the samples in the experience replay pool, an experience sample loop of state, action, reward, and new state is formed, realizing a continuous trial-and-error, feedback, learning, and optimization self-learning closed loop. This enables the strategy to dynamically adapt to changing loads and environments, achieving long-term intelligent self-evolution and optimal regulation.

[0065] In this implementation scheme, by continuously sampling and training the DRL interaction experience tuple set, the parameters of the policy network and value network are continuously optimized, achieving self-learning and self-evolution of the control strategy. With the model's comprehensive reward value as the core, the strategy converges to a comprehensive goal of achieving temperature control targets, optimal energy consumption, and stable operation, ensuring that each strategy update is dynamically optimized based on actual operating results and multi-objective constraints. When the policy network performs optimally, it can be solidified into a real-time online control model, improving the decision-making efficiency and intelligence level of the air-liquid hybrid cooling system. Simultaneously, the periodic dynamic evaluation of the comprehensive reward value and the anomaly rollback mechanism provide robust protection, ensuring safe, efficient, and adaptive operation under load changes and extreme conditions.

[0066] Reference Figure 2As shown, the second aspect of the present invention provides a deep reinforcement learning adaptive control system for a wind-liquid hybrid cooling system, applied to the aforementioned deep reinforcement learning adaptive control method for the wind-liquid hybrid cooling system. The system includes: a multi-source data acquisition and risk assessment module, used to acquire wind-liquid hybrid cooling data in real time, preprocess the wind-liquid hybrid cooling data, assess the heat load of the equipment based on the wind-liquid hybrid cooling data, construct a heat load feature vector based on the heat load of the equipment, and perform preliminary heat load risk adjustment; a DRL adaptive decision-making and action mapping module, used to receive wind-liquid hybrid cooling data and the heat load feature vector, determine the proportion of air cooling regulation in wind-liquid synergistic cooling, construct a DRL action mapping model based on the proportion of air cooling regulation and the wind-liquid hybrid cooling data, and generate specific action commands; an actuator linkage and real-time feedback module, used to execute specific action commands, evaluate the action execution effect after the specific action commands are executed, optimize the DRL action mapping model based on the action execution effect, and construct a DRL interactive experience tuple set; and a closed-loop self-learning and self-evolution module, used to drive the optimization of action commands by training the DRL interactive experience tuple set, thereby realizing the self-learning and control strategy evolution of wind-liquid hybrid cooling.

[0067] In this implementation plan, comprehensive perception and dynamic analysis of heat load under air-liquid hybrid cooling conditions are achieved through real-time acquisition of multi-source data and risk assessment. An adaptive decision-making and DRL action mapping model based on deep reinforcement learning is constructed, enabling the air-cooling and liquid-cooling adjustment ratio to be automatically optimized according to the actual load and dynamically generating optimal control commands. Relying on actuator linkage and real-time feedback mechanisms, the action execution effect can be continuously evaluated and optimized, and the actual operation results are fed back to the model for self-evolution. Finally, through a closed-loop self-learning process, the control strategy is continuously iterated and improved, realizing the adaptive, intelligent, high-efficiency, energy-saving, and long-term reliable operation of air-liquid hybrid cooling, effectively solving the problems of traditional cooling systems such as difficulty in dynamically responding to complex loads, rigid control, and energy waste.

[0068] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0069] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. As those skilled in the art will understand, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.

Claims

1. A deep reinforcement learning adaptive control method for a wind-liquid hybrid cooling system, characterized in that, Includes the following steps: S1: Real-time acquisition of air-liquid mixed cooling data, data preprocessing of air-liquid mixed cooling data, assessment of equipment heat load based on air-liquid mixed cooling data, construction of heat load feature vector based on equipment heat load, and preliminary heat load risk adjustment. S2, receives air-liquid mixed cooling data and heat load feature vector, determines the proportion of air cooling regulation in air-liquid synergistic cooling, constructs DRL action mapping model based on the proportion of air cooling regulation and air-liquid mixed cooling data, and generates specific action commands. The specific process of receiving the air-liquid mixed cooling data and the heat load feature vector to determine the proportion of air-cooling regulation in air-liquid synergistic cooling is as follows: Real-time reception of air-liquid mixing cooling data and heat load feature vector; addition of the current total fan airflow, total coolant flow rate, and minimum constant value to obtain the total air-liquid flow rate; division of the total fan airflow by the total air-liquid flow rate to obtain the air-cooling capacity ratio; multiplication of the air-cooling capacity ratio by the heat load risk coefficient to obtain the weighted heat load pressure; division of the weighted heat load pressure by the sum of the heat load risk coefficient and the minimum constant value to obtain the air-cooling synergistic ratio. S3 executes specific action instructions, evaluates the action execution effect after the specific action instructions are executed, optimizes the DRL action mapping model based on the action execution effect, and constructs a DRL interaction experience tuple set; S4, through training on the DRL interactive experience tuple set, drives the optimization of action commands, realizing the self-learning and control strategy evolution of air-liquid hybrid cooling.

2. The deep reinforcement learning adaptive control method for the air-liquid hybrid cooling system according to claim 1, characterized in that, The specific process of real-time acquisition of air-liquid mixing cooling data and data preprocessing of the air-liquid mixing cooling data is as follows: Real-time data collection of air-liquid mixed cooling is achieved through temperature sensor (1), intelligent power distribution unit (PDU), flow sensor (2), pressure sensor, air volume sensor and speed sensor. The air-liquid mixed cooling data includes: chip temperature, cabinet temperature, coolant inlet temperature, coolant outlet temperature, coolant specific heat capacity, total power consumption of equipment, total coolant flow rate, total coolant pressure, total fan air volume, fan speed and pump speed. Outliers in the air-liquid mixed cooling data were identified and removed by sliding window statistics and the rule of three times standard deviation. Missing points were filled in by linear interpolation and historical mean regression. The data were uniformly aligned to the main time axis and resampled. The air-liquid mixed cooling data were smoothed by window mean and exponential weighted moving average, and then normalized and scaled. An air-liquid mixed cooling database was established to store the original and pre-processed air-liquid mixed cooling data, and the air-liquid mixed cooling data was transmitted to the DRL controller (3).

3. The deep reinforcement learning adaptive control method for the air-liquid hybrid cooling system according to claim 1, characterized in that, The specific process for evaluating the heat load of the equipment based on air-liquid mixed cooling data is as follows: Based on a sliding time window, the air-liquid mixing cooling data at each moment within the window is acquired in real time. The coolant temperature difference is obtained by subtracting the coolant inlet temperature from the coolant outlet temperature and taking the absolute value. The instantaneous cooling capacity is obtained by multiplying the total coolant flow rate, the coolant specific heat capacity, and the coolant temperature difference. The total cooling capacity is obtained by numerically integrating the instantaneous cooling capacity at each moment within the entire sliding time window. The heat load risk coefficient is obtained by dividing the total power consumption of the equipment at the current moment by the total cooling capacity.

4. The deep reinforcement learning adaptive control method for the air-liquid hybrid cooling system according to claim 1, characterized in that, The specific process of constructing a heat load feature vector based on the equipment's heat load and performing preliminary heat load risk adjustment is as follows: Calculate the heat load risk coefficient in real time and compare it with the heat load risk threshold. If the heat load risk coefficient is greater than the heat load risk threshold, perform preliminary heat load risk adjustment: adjust the fan (6), liquid pump (5) and valve to increase the fan (6) speed and coolant flow rate. If the heat load risk coefficient is less than or equal to the heat load risk threshold, continuous monitoring and routine recording should be carried out to maintain the current control strategy and operating status. The system integrates the coolant inlet temperature, coolant outlet temperature, total power consumption of the equipment, total coolant flow rate, and corresponding heat load risk coefficient to construct a heat load feature vector, writes it into the air-liquid hybrid cooling database, and proceeds to the next step.

5. The deep reinforcement learning adaptive control method for the air-liquid hybrid cooling system according to claim 1, characterized in that, The specific process of constructing a DRL action mapping model and generating specific action commands based on the proportion of air-cooled regulation and air-liquid mixed cooling data is as follows: The air-cooling synergy ratio, heat load risk coefficient, fan speed, pump speed and chip temperature are used as the state mapping feature set. The state mapping feature set is trained by interacting with the simulation environment. The DRL action mapping model is constructed by deep deterministic policy gradient and DRL deep reinforcement learning algorithm. The action mapping module performs action mapping (4) operation to generate specific action instructions for each actuator of fan (6), liquid pump (5) and valve. The policy network outputs the instructions and sends them to each actuator.

6. The deep reinforcement learning adaptive control method for the air-liquid hybrid cooling system according to claim 1, characterized in that, The specific process of executing specific action instructions and evaluating the effect of action execution after the specific action instructions are executed is as follows: Receive specific action instructions output by the DRL action mapping model, and each actuator will be synchronously adjusted according to the specific action instructions, and the adjusted air-liquid mixing cooling data will be obtained in real time. The highest temperature value between the chip temperature and the cabinet temperature at the current moment is selected as the current overall maximum temperature. Based on the sliding time window, the overall maximum temperature sequence is obtained, and the average value is calculated as the temperature control target value. The total power consumption of the device at the previous moment and the total power consumption of the device at the current moment are obtained. The total power consumption of the device at the current moment is subtracted from the total power consumption of the device at the previous moment to obtain the change in total power consumption. Meanwhile, within the sliding time window, the standard deviations of the overall maximum temperature, total power consumption of the equipment, total coolant flow rate, and total fan air volume are calculated separately. The standard deviations of each parameter are normalized. The normalized values ​​of the standard deviations of each parameter are trained using the principal component analysis algorithm. The coefficients of each parameter of the first principal component are extracted as weights. Based on the weights of each parameter, the standard deviations of the overall maximum temperature, total power consumption of the equipment, total coolant flow rate, and total fan air volume are weighted and summed, and the negative value is taken to obtain the operating stability value. The absolute value of the difference between the current overall maximum temperature and the temperature control target value is calculated and taken as negative to obtain the temperature control deviation value; the total power consumption change is multiplied by the power consumption change weight factor and taken as negative to obtain the power consumption adjustment value; the operating stability value is multiplied by the operating stability weight factor to obtain the operating stability value; the temperature control deviation value, the power consumption adjustment value, and the operating stability value are added together to obtain the model comprehensive reward value.

7. The deep reinforcement learning adaptive control method for the air-liquid hybrid cooling system according to claim 1, characterized in that, The specific process of optimizing the DRL action mapping model based on the action execution effect and constructing the DRL interaction experience tuple set is as follows: After each specific action instruction is generated and executed, the model's comprehensive reward value is calculated and fed back to the DRL action mapping model to evaluate the merits of the specific action instruction. A random sampling optimization strategy and a value network are used to improve the decision performance of the DRL action mapping model. Continuously monitor the model's comprehensive reward value. If the standard deviation of the model's comprehensive reward value is found to be greater than the fluctuation threshold within a certain period of time, it is determined to be abnormal in temperature control, energy consumption, and stability. Switch the fan (6) and liquid pump (5) to the maximum load operation mode to reduce the load of some non-core servers. Increase the safety thresholds for temperature, pressure, and power consumption. If the safety threshold is exceeded, the load reduction and local cooling will be forcibly executed. Simultaneously, an experience replay pool is constructed, which combines the pre-execution air-liquid mixing cooling data, action commands, model comprehensive reward value, and post-execution air-liquid mixing cooling data to construct a DRL interactive experience tuple set, and stores it in the experience replay pool and the air-liquid mixing cooling database.

8. The deep reinforcement learning adaptive control method for the air-liquid hybrid cooling system according to claim 1, characterized in that, The specific process of training the DRL interaction experience tuple set to drive the optimization of action commands and realize the self-learning and control strategy evolution of air-liquid hybrid cooling is as follows: The DRL interaction experience tuple set is continuously sampled from the experience replay pool as training samples and input into the policy network and value network of the DRL algorithm for iterative updates of trial and error, feedback, and learning. The policy network adjusts the mapping between the state mapping feature set and specific action instructions based on the sample feedback to achieve dynamic optimization of policy parameters. Each round of training is based on the model's comprehensive reward value to guide the network to continuously converge toward the target policy of achieving temperature target, low energy consumption, and stable operation. When the policy network converges and the comprehensive reward value is maximized, the current optimal policy is solidified into an online deployment model to drive the control decision of air-liquid hybrid cooling in real time. At the same time, the model's comprehensive reward value is evaluated periodically. If the standard deviation of the model's comprehensive reward value exceeds the fluctuation threshold or the policy fails, it will revert to the historical optimal policy.

9. A deep reinforcement learning adaptive control system for a wind-liquid hybrid cooling system, employing the deep reinforcement learning adaptive control method for a wind-liquid hybrid cooling system as described in any one of claims 1-8, characterized in that, include: The multi-source data acquisition and risk assessment module is used to acquire air-liquid mixed cooling data in real time, preprocess the air-liquid mixed cooling data, assess the heat load of the equipment based on the air-liquid mixed cooling data, construct a heat load feature vector based on the heat load of the equipment, and perform preliminary heat load risk adjustment. The DRL adaptive decision and action mapping module is used to receive air-liquid hybrid cooling data and heat load feature vector, determine the proportion of air cooling regulation in air-liquid synergistic cooling, construct a DRL action mapping model based on the proportion of air cooling regulation and air-liquid hybrid cooling data, and generate specific action commands. The actuator linkage and real-time feedback module is used to execute specific action instructions, evaluate the action execution effect after the specific action instructions are executed, optimize the DRL action mapping model based on the action execution effect, and construct a DRL interactive experience tuple set. The closed-loop self-learning and self-evolution module is used to drive the optimization of action commands by training on the DRL interaction experience tuple set, so as to realize the self-learning and regulation strategy evolution of air-liquid hybrid cooling.

Citation Information

Patent Citations

  • Data center energy efficiency optimization method and system based on reinforcement learning

    CN114970358A

  • Cold and hot channel air volume optimization energy-saving method based on AI prediction

    CN120354716A