Liquid cooling system control method based on Actor-Critic algorithm

By introducing the Actor-Critic algorithm into the liquid-cooled cooling system, intelligent dynamic control is realized, solving the cooling efficiency and energy consumption problems of the liquid-cooled system when the load changes in the data center, and improving the stability and energy efficiency of the system.

CN120276248APending Publication Date: 2025-07-08INSPUR TIANYUAN COMM INFORMATION SYST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510353968.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The static control method of existing liquid cooling systems is difficult to adapt to changes in data center load requirements, resulting in reduced cooling efficiency and increased energy consumption.

Method used

Using reinforcement learning method based on Actor-Critic algorithm, by establishing state space, action space and reward functions, dynamically adjusting the solenoid valve opening and coolant temperature of the cooling system to achieve intelligent control.

Benefits of technology

Improves the dynamic adaptability and efficiency of the cooling system, reduces energy consumption, ensures optimal performance under different load conditions, reduces energy waste and improves system stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120276248A_ABST
    Figure CN120276248A_ABST
Patent Text Reader

Abstract

The invention discloses a liquid cooling system control method based on an Actor-Critic algorithm, and relates to the technical field of intelligent control. Comprising the following steps: 1, establishing a model according to cooling system environment data of a data center, defining a state space, an action space and a reward function through the model, describing the state of the cooling system environment by using a state vector in the state space, and providing an action vector executed by a cooling system by using the action space, a reward function is used as a standard for evaluating whether the selection strategy is good or not; and 2, an Actor-Critic algorithm is used for controlling the cooling system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention discloses a control method for a liquid cooling system based on the Actor-Critic algorithm, which relates to the technical field of intelligent control. Background Art

[0002] With the continuous increase in the scale and power density of data centers, the energy efficiency of the cooling system has become a key issue. Most existing liquid cooling systems use fixed frequency or simple feedback regulation methods for control and adjustment. This static control method is difficult to adapt to the changing load requirements, resulting in reduced cooling efficiency and increased energy consumption. Summary of the Invention

[0003] In view of the problems of the prior art, the present invention provides a control method for a liquid cooling system based on the Actor-Critic algorithm, introducing a reinforcement learning algorithm to achieve intelligent and automatic control of the liquid cooling system.

[0004] The specific solution proposed by the present invention is as follows:

[0005] The present invention provides a control method for a liquid cooling system based on the Actor-Critic algorithm, including: Step 1: Establish a model according to the environmental data of the cooling system in the data center, define the state space, action space and reward function through the model, use the state vector in the state space to describe the state of the cooling system environment, use the action space to provide the action vector executed by the cooling system, and use the reward function as the standard for evaluating the quality of the selected strategy.

[0006] Step 2: Use the Actor-Critic algorithm to control the cooling system:

[0007] Initialize the environmental data of the cooling system, the experience replay pool, and the Actor network parameters and Critic network parameters of the Actor-Critic algorithm.

[0008] Obtain the current state vector s of the cooling system, input it into the Actor network, and select the action vector a according to the ε-greedy strategy.

[0009] Execute the action vector a, and obtain the next state vector s of the cooling system environment feedback. next and the reward r, and store (s, a, r, s next ) in the experience replay pool.

[0010] Randomly sample N samples (s, a, r, s next ) from the experience replay pool. For i = 1 to N, loop and execute the following steps:

[0011] Use the Critic network to calculate the current Q value: Q current= Critic(s, a); Calculate the Q value of the next state vector using the Critic network, Q next = Critic(s next , Actor(s next ));

[0012] Calculate the target Q value: Q target = r + γ × Q next ;

[0013] Calculate the loss of the Critic network: L critic = (Q target - Q current ) 2 ;

[0014] Update the Critic network parameters using gradient descent to minimize L critic ;

[0015] Calculate the loss of the Actor network: L actor = -Critic(s, Actor(s));

[0016] Update the Actor network parameters using gradient ascent to minimize L actor ;

[0017] Transfer the state vector to: s = s next ;

[0018] Loop: The process from obtaining the state vector s and inputting it into the Actor network to transferring the state vector is carried out to maintain real-time control of the cooling system.

[0019] Furthermore, in the sampling period within a fixed window time in step 1 of the method for controlling a liquid cooling system based on the Actor-Critic algorithm, the state space of the cooling system is represented as S i = (t1, t2, …, t n , p i ), where t i represents the average temperature of each rack sensor responsible for by the liquid cooling system within the period, and p i is the power consumption of the liquid cooling system within this period.

[0020] Furthermore, in step 1 of the method for controlling a liquid cooling system based on the Actor-Critic algorithm, the action space of the cooling system is A i , and the action vector executed is represented as a i = (vo i , lt i ), a i ∈ A i , where voi <lt> i respectively represent the solenoid valve opening degree and the coolant temperature of the refrigeration system, and both are adjusted within the safe adjustment range.

[0021] Furthermore, the reward function set in step 1 of the liquid cooling system control method based on the Actor-Critic algorithm includes two parts, namely the temperature penalty cost and the power consumption of the cooling system. The temperature penalty cost t cost The calculation formula is as follows:

[0022] t cost = cost1 + cost2 +... + cost n

[0023]

[0024] t min and t max respectively represent the minimum temperature and the maximum temperature of the safety threshold of the equipment cooled by the cooling system.

[0025] The calculation formula of the reward function is expressed as:

[0026] R i = -(αt cost + (1 - α)p i × w)

[0027]

[0028] α represents the weight of the power consumption of the cooling system in the reward function. p min and p max respectively represent the minimum power consumption and the maximum power consumption within the safety threshold of the cooling system. w is a normalization coefficient to ensure that the temperature penalty cost and the power consumption of the cooling system are in the same order of magnitude.

[0029] The present invention also provides a liquid cooling system control device based on the Actor-Critic algorithm, including a model management module and a control module.

[0030] The model management module establishes a model according to the cooling system environment data of the data center, defines the state space, action space and reward function through the model, describes the state of the cooling system environment by using the state vector in the state space, provides the action vector executed by the cooling system by using the action space, and uses the reward function as the standard for evaluating the quality of the selection strategy.

[0031] The control module controls the cooling system by using the Actor-Critic algorithm:

[0032] Initialize the environmental data of the cooling system, the experience replay pool, as well as the Actor network parameters and Critic network parameters of the Actor-Critic algorithm.

[0033] Obtain the current state vector s of the cooling system, input it into the Actor network, and select the action vector a according to the ε-greedy strategy.

[0034] Execute the action vector a, and obtain the next state vector s and the reward r feedback from the cooling system environment. Store (s, a, r, s next ) into the experience replay pool. next )

[0035] Randomly sample N samples (s, a, r, s next ) from the experience replay pool. For i = 1 to N, loop and execute the following steps:

[0036] Use the Critic network to calculate the current Q value: Q current = Critic(s, a); Use the Critic network to calculate the Q value of the next state vector, Q next = Critic(s next , Actor(s next ));

[0037] Calculate the target Q value: Q target = r + γ × Q next ;

[0038] Calculate the loss of the Critic network: L critic = (Q target - Q current ) 2 ;

[0039] Use the gradient descent method to update the Critic network parameters to minimize L critic ;

[0040] Calculate the loss of the Actor network: L actor = -Critic(s, Actor(s));

[0041] Use the gradient ascent method to update the Actor network parameters to minimize L actor ;

[0042] Transfer the state vector to: s = s next ;

[0043] Loop: The process from obtaining the state vector s and inputting it into the Actor network to transferring the state vector, and maintain real-time control of the cooling system.

[0044] Furthermore, in the sampling period within a fixed window time of the model management module of the liquid cooling system control device based on the Actor-Critic algorithm, the state space of the cooling system is represented as S i =(t1, t2, …, t n , p i ), where t i represents the average temperature of each rack sensor responsible for by the liquid cooling system within the period, and p i is the power consumption of the liquid cooling system within this period.

[0045] Furthermore, the model management module of the liquid cooling system control device based on the Actor-Critic algorithm represents the action space of the cooling system as A i , and the executed action vector is represented as a i =(vo i , lt i ), a i ∈A i , where vo i and lt i represent the solenoid valve opening degree and the coolant temperature of the refrigeration system respectively, and both are adjusted within the safe adjustment range.

[0046] Furthermore, the reward function set by the model management module of the liquid cooling system control device based on the Actor-Critic algorithm includes two parts, namely the temperature penalty cost and the power consumption of the cooling system. The calculation formula of the temperature penalty cost t cost is as follows:

[0047] t cost =cost1 + cost2 +... + cost n

[0048]

[0049] t min and t max represent the minimum temperature and the maximum temperature of the safety threshold of the equipment cooled by the cooling system respectively,

[0050] The calculation formula of the reward function is expressed as:

[0051] R i =-(αt cost +(1 - α)p i ×w)

[0052]

[0053] α represents the weight of the power consumption of the cooling system in the reward in the reward function, p min and pmax respectively represent the minimum power consumption and the maximum power consumption within the safety threshold of the cooling system, and w is a normalization coefficient to ensure that the temperature penalty cost and the power consumption of the cooling system are in the same order of magnitude.

[0054] The advantages of the present invention are as follows:

[0055] Dynamic adaptability: By integrating real-time load and cooling demand data, it can autonomously and dynamically adjust the parameters of the cooling system. This not only ensures the efficient operation of the liquid cooling system, but also significantly improves the cooling efficiency, and can maintain optimal performance under different load conditions, thereby effectively reducing energy consumption.

[0056] Intelligent strategy optimization: Using reinforcement learning algorithms, it can continuously learn and adapt to environmental changes. This intelligent method enables the liquid cooling system to update control strategies in real time in complex and changing environments and always maintain the best operating state. Through this continuous optimization process, it can better handle various complex situations and improve the overall operating efficiency.

[0057] Efficient energy consumption management: Through advanced intelligent adjustment technologies, the present invention can avoid overcooling and the resulting energy waste. By precisely controlling the operating state of the liquid cooling system, it can be adjusted according to the actual cooling demand, thereby optimizing the energy use efficiency. This energy-saving strategy not only reduces energy consumption, but also helps to reduce the operating costs of the data center.

[0058] Prediction and early response: Through the analysis of historical data of reinforcement learning, it can predict future load trends and changes in cooling demand. The system can make necessary adjustments in advance to avoid a decrease in cooling efficiency caused by sudden load changes or environmental fluctuations, and ensure that the system is always in an efficient operating state. This prediction ability greatly enhances the stability and reliability of the system.

[0059] Enhanced system stability: Through intelligent control strategies and dynamic adjustment mechanisms, the present invention significantly enhances the stability of the liquid cooling system. It can maintain balance under various operating conditions and avoid operating instability caused by frequent load fluctuations or environmental changes. This enhanced stability not only increases the service life of the equipment, but also further optimizes the efficiency and energy consumption management of the overall cooling system. Description of the Drawings

[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0061] Figure 1It is a schematic diagram of the cooling system modeling of the Actor-Critic framework of the present invention.

[0062] Figure 2 It is a schematic diagram of the instantaneous reward broken line.

[0063] Figure 3 It is a schematic diagram of the PUE broken line. Specific embodiments

[0064] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the embodiments given are not intended to limit the present invention.

[0065] Embodiment 1

[0066] The present invention provides a liquid cooling system control method based on the Actor-Critic algorithm, including: Step 1: Establish a model according to the cooling system environment data of the data center, define the state space, action space and reward function through the model, use the state vector in the state space to describe the state of the cooling system environment, use the action space to provide the action vector executed by the cooling system, and use the reward function as the standard for evaluating the quality of the selection strategy.

[0067] Among them, in the sampling period within a fixed window time, the state space of the cooling system is represented as S i =(t1,t2,…,t n ,p i ), where t i represents the average temperature of each rack sensor responsible for the liquid cooling system within the period, and p i is the power consumption of the liquid cooling system within this period.

[0068] The action space of the cooling system is A i , and the executed action vector is represented as a i =(vo i ,lt i ), a i ∈A i , where vo i and lt i respectively represent the solenoid valve opening degree and the coolant temperature of the refrigeration system, and both are adjusted within the safe adjustment range.

[0069] The set reward function includes two parts, namely the temperature penalty cost and the power consumption of the cooling system. After executing the action a i , the energy consumption is the difference between the power meter acquisition data of the start node and the end node within the sliding window period. The temperature penalty cost is related to the equipment temperature exceeding the set range. The calculation formula of the temperature penalty cost t cost is as follows:

[0070] t cost = cost1 + cost2 +... + cost n

[0071]

[0072] t min and t max respectively represent the minimum temperature and the maximum temperature of the safety threshold of the equipment cooled by the cooling system.

[0073] The calculation formula of the reward function is expressed as:

[0074] R i = -(αt cost + (1 - α)p i × w)

[0075]

[0076] α represents the weight of the power consumption of the cooling system in the reward function. If α increases, it means a greater temperature penalty. If α decreases, it means a greater penalty for excessive power consumption of the cooling system. p min and p max respectively represent the minimum power consumption and the maximum power consumption within the safety threshold of the cooling system. w is a normalization coefficient to ensure that the temperature penalty cost and the power consumption of the cooling system are in the same order of magnitude.

[0077] Step 2: Control the cooling system using the Actor-Critic algorithm:

[0078] Initialize the environmental data of the cooling system, the experience replay pool, and the Actor network parameters and Critic network parameters of the Actor-Critic algorithm.

[0079] Obtain the current state vector s of the cooling system, input it into the Actor network, and select the action vector a according to the ε-greedy strategy.

[0080] Execute the action vector a, and obtain the next state vector s next and the reward r, and store (s, a, r, s next ) in the experience replay pool.

[0081] Randomly sample N samples (s, a, r, s next ) from the experience replay pool. For i = 1 to N, loop and execute the following steps:

[0082] Use the Critic network to calculate the current Q value: Q current = Critic(s, a); Use the Critic network to calculate the Q value of the next state vector, Qnext = Critic(s next , Actor(s next ));

[0083] Calculate the target Q value: Q target = r + γ × Q next ;

[0084] Calculate the loss of the Critic network: L critic = (Q target - Q current ) 2 ;

[0085] Update the parameters of the Critic network using gradient descent to minimize L critic ;

[0086] Calculate the loss of the Actor network: L actor = -Critic(s, Actor(s));

[0087] Update the parameters of the Actor network using gradient ascent to minimize L actor ;

[0088] Transfer the state vector to: s = s next ;

[0089] Loop: The process of obtaining the state vector s, inputting it into the Actor network until the state vector is transferred, and maintaining real-time control of the cooling system.

[0090] By introducing the Actor-Critic algorithm, the cooling system can perceive load changes in real time, dynamically adjust the operating parameters of the circulation pump, and achieve a more precise cooling effect. In addition, this reinforcement learning method can optimize the control strategy through long-term data, reduce energy consumption, and improve system stability. By predicting future load trends, the system can adjust the cooling strategy in advance to avoid efficiency degradation caused by sudden load changes.

[0091] Where Figure 2 Show the instantaneous reward line chart of the Actor-Critic algorithm and the PID control algorithm of the present invention. Due to its fixed control strategy, the PID control algorithm, although stable in the steady state, fails to achieve significant optimization in terms of instantaneous reward. In contrast, the Actor-Critic algorithm explores actions with a high probability in the initial stage, ensuring a wide and deep search of the action space. After a period of training, the instantaneous reward of the Actor-Critic algorithm gradually stabilizes and is significantly higher than that of the PID control algorithm, achieving a better control effect.

[0092] PUE value optimization: Figure 3Show the PUE line chart under steady state. PUE is a key indicator for energy efficiency management in computer rooms, defined as the ratio of the total energy consumption of the computer room to the energy consumption of IT equipment. Ideally, the closer the PUE value is to 1, the higher the energy efficiency ratio of the computer room. The experimental results show that the PUE value of the PID control algorithm is stable at 1.362, while the PUE value of the Actor-Critic algorithm is stable at 1.173 under steady state. This indicates that the application of the Actor-Critic algorithm has increased the energy-saving efficiency of the computer room cooling system by 13.88%, significantly reducing the overall energy consumption of the computer room.

[0093] Example 2

[0094] The present invention also provides a liquid cooling system control device based on the Actor-Critic algorithm, including a model management module and a control module.

[0095] The model management module establishes a model based on the cooling system environment data of the data center, defines the state space, action space, and reward function through the model, uses the state vector in the state space to describe the state of the cooling system environment, uses the action space to provide the action vector executed by the cooling system, and uses the reward function as the standard for evaluating the quality of the selection strategy.

[0096] The control module uses the Actor-Critic algorithm to control the cooling system:

[0097] Initialize the cooling system environment data, experience replay pool, and the Actor network parameters and Critic network parameters of the Actor-Critic algorithm.

[0098] Obtain the current state vector s of the cooling system, input it into the Actor network, and select the action vector a according to the ε-greedy strategy.

[0099] Execute the action vector a, and obtain the next state vector s feedback by the cooling system environment next and the reward r, and store (s, a, r, s next ) in the experience replay pool.

[0100] Randomly sample N samples (s, a, r, s next ) from the experience replay pool. For i = 1 to N, loop and execute the following steps:

[0101] Use the Critic network to calculate the current Q value: Q current = Critic(s,a); Use the Critic network to calculate the Q value of the next state vector, Q next = Critic(s next ,Actor(s next ));

[0102] Calculate the target Q value: Q target = r + γ × Q next ;

[0103] Calculate the loss of the Critic network: L critic = (Q target - Q current ) 2 ;

[0104] Update the Critic network parameters using gradient descent to minimize L critic ;

[0105] Calculate the loss of the Actor network: L actor = -Critic(s, Actor(s));

[0106] Update the Actor network parameters using gradient ascent to minimize L actor ;

[0107] Transfer the state vector to: s = s next ;

[0108] Loop: The process of obtaining the state vector s, inputting it into the Actor network until the state vector is transferred, and maintaining real-time control of the cooling system.

[0109] Regarding the information interaction and execution process among the modules in the above device, since they are based on the same concept as the method embodiment of the present invention, the specific content can be referred to the description in the method embodiment of the present invention, and will not be elaborated here.

[0110] Similarly, the advantages of the device of the present invention are:

[0111] Dynamic adaptability: By integrating real-time load and cooling demand data, it can autonomously and dynamically adjust the cooling system parameters. This not only ensures the efficient operation of the liquid cooling system but also significantly improves the cooling efficiency. At the same time, it can maintain the optimal performance under different load conditions, thereby effectively reducing energy consumption.

[0112] Intelligent strategy optimization: Using reinforcement learning algorithms, it can continuously learn and adapt to environmental changes. This intelligent method enables the liquid cooling system to update control strategies in real time in a complex and changing environment and always maintain the best operating state. Through this continuous optimization process, it can better handle various complex situations and improve the overall operating efficiency.

[0113] Efficient Energy Consumption Management: Through advanced intelligent adjustment technologies, the present invention can avoid overcooling and the resulting energy waste. By precisely controlling the operating state of the liquid cooling system, it can be adjusted according to the actual cooling demand, thereby optimizing the energy usage efficiency. This energy-saving strategy not only reduces energy consumption but also helps to lower the operating costs of the data center.

[0114] Prediction and Early Response: Through the analysis of historical data using reinforcement learning, the present invention can predict future load trends and changes in cooling demand. The system can make necessary adjustments in advance to avoid a decrease in cooling efficiency caused by sudden load changes or environmental fluctuations, ensuring that the system is always in an efficient operating state. This predictive ability greatly enhances the stability and reliability of the system.

[0115] Enhanced System Stability: Through intelligent control strategies and dynamic adjustment mechanisms, the present invention significantly enhances the stability of the liquid cooling system. It can maintain balance under various operating conditions and avoid operating instability caused by frequent load fluctuations or environmental changes. This enhanced stability not only increases the service life of the equipment but also further optimizes the efficiency and energy consumption management of the overall cooling system.

[0116] It should be noted that not all steps and modules in the above-mentioned processes and device structures are necessary, and some steps or modules can be ignored according to actual needs. The execution order of each step is not fixed and can be adjusted as required. The system structures described in the above embodiments can be physical structures or logical structures, that is, some modules may be implemented by the same physical entity, or some modules may be implemented by multiple physical entities, or some components in multiple independent devices may be jointly implemented.

[0117] The above-described embodiments are only preferred embodiments given to fully illustrate the present invention, and the protection scope of the present invention is not limited thereto. Equivalent substitutions or transformations made by those skilled in the art on the basis of the present invention are all within the protection scope of the present invention. The protection scope of the present invention is subject to the claims.

Claims

1. A control method for a liquid cooling system based on the Actor-Critic algorithm, characterized in that Including: Step 1: Establish a model based on the cooling system environment data of the data center. Define the state space, action space, and reward function through the model. Use the state vector in the state space to describe the state of the cooling system environment, use the action space to provide the action vector executed by the cooling system, and use the reward function as the criterion for evaluating the quality of the selection strategy. Step 2: Use the Actor-Critic algorithm to control the cooling system: Initialize the cooling system environment data, experience replay pool, as well as the Actor network parameters and Critic network parameters of the Actor-Critic algorithm. Obtain the current state vector s of the cooling system, input it into the Actor network, and select the action vector a according to the ε-greedy strategy. Execute the action vector a and obtain the next state vector s of the cooling system environment feedback next and the reward r, and store (s, a, r, s next ) in the experience replay pool Randomly sample N samples (s, a, r, s next ) from the experience replay pool. For i = 1 to N, loop through and execute the following steps: Calculate the current Q value using the Critic network: Q current = Critic(s, a); Calculate the Q value of the next state vector using the Critic network, Q next = Critic(s next , Actor(s next )); Calculate the target Q value: Q target = r + γ × Q next ; Calculate the loss of the Critic network: L critic = (Q target - Q current ) 2 ; Update the Critic network parameters using gradient descent to minimize L critic ; Calculate the loss of the Actor network: L actor = -Critic(s, Actor(s)); Update the Actor network parameters using the gradient ascent method to minimize L actor ; Transfer the state vector to: s = s next ; Loop: The process of obtaining the state vector s and inputting it into the Actor network until the state vector is transferred, maintaining real-time control of the cooling system.

2. The control method of a liquid cooling system based on the Actor-Critic algorithm according to claim 1, characterized in that in the sampling period within a fixed window time in step 1, the state space of the cooling system is represented as S i =(t1,t2,…,t n ,p i ), where t i represents the average temperature of each rack sensor responsible for by the liquid cooling system within the period, and p i is the power consumption of the liquid cooling system within this period.

3. A control method for a liquid cooling system based on the Actor-Critic algorithm according to claim 1, characterized in that the action space of the cooling system in step 1 is A i , and the executed action vector is represented as a i =(vo i , lt i ), a i ∈A i , where vo i and lt i respectively represent the solenoid valve opening degree of the refrigeration system and the coolant temperature, and both are adjusted within the safe adjustment range.

4. The control method of a liquid cooling system based on the Actor-Critic algorithm according to claim 2, characterized in that The reward function set in Step 1 consists of two parts, namely the temperature penalty cost and the power consumption of the cooling system. The temperature penalty cost t cost is calculated as follows: t cost = cost1 + cost2 +... + cost n t min and t max respectively represent the minimum temperature and the maximum temperature of the safety threshold of the equipment cooled by the cooling system, The calculation formula of the reward function is expressed as: R i = -(αt cost + (1 - α)p i × w) α represents the weight of the power consumption of the cooling system in the reward function in the reward, p min and p max respectively represent the minimum power consumption and the maximum power consumption within the safety threshold of the cooling system. w is a normalization coefficient to ensure that the temperature penalty cost and the power consumption of the cooling system are in the same order of magnitude.

5. A control device for a liquid cooling system based on the Actor-Critic algorithm, characterized in that Including a model management module and a control module. The model management module establishes a model based on the cooling system environment data of the data center. Define the state space, action space, and reward function through the model. Use the state vector in the state space to describe the state of the cooling system environment, use the action space to provide the action vector executed by the cooling system, and use the reward function as the criterion for evaluating the quality of the selection strategy. The control module uses the Actor-Critic algorithm to control the cooling system: Initialize the cooling system environment data, experience replay pool, as well as the Actor network parameters and Critic network parameters of the Actor-Critic algorithm. Obtain the current state vector s of the cooling system, input it into the Actor network, and select the action vector a according to the ε-greedy strategy. Execute the action vector a and obtain the next state vector s of the cooling system environment feedback next and the reward r, and store (s, a, r, s next ) in the experience replay pool Randomly sample N samples (s, a, r, s next ) from the experience replay pool, and loop through the following steps for i = 1 to N: Calculate the current Q value using the Critic network: Q current = Critic(s,a); Calculate the Q value of the next state vector using the Critic network, Q next = Critic(s next , Actor(s next )); Calculate the target Q value: Q target = r + γ × Q next ; Calculate the loss of the Critic network: L critic =(Q target -Q current ) 2 ; Update the Critic network parameters using gradient descent to minimize L critic ; Calculate the loss of the Actor network: L actor = -Critic(s, Actor(s)); Update the Actor network parameters using the gradient ascent method to minimize L actor ; Transfer the state vector to: s = s next ; Loop: The process of obtaining the state vector s and inputting it into the Actor network until the state vector is transferred, maintaining real-time control of the cooling system.

6. The control device of the liquid cooling system based on the Actor-Critic algorithm according to claim 5, wherein the model During the sampling period within the fixed window time, the management module represents the state space of the cooling system as S i =(t1,t2,…,t n ,p i ), where t i represents the average temperature of each rack sensor responsible for the liquid cooling system within the period, and p i is the power consumption of the liquid cooling system within this period.

7. The control device of the liquid cooling system based on the Actor-Critic algorithm according to claim 5, wherein the model The management module represents the action space of the cooling system as A i , and the executed action vector is represented as a i =(vo i ,lt i ), a i ∈A i , where vo i and lt i respectively represent the solenoid valve opening degree and the coolant temperature of the refrigeration system, and both are adjusted within the safe adjustment range.

8. The control device of the liquid cooling system based on the Actor-Critic algorithm according to claim 6, characterized in that The reward function set by the model management module consists of two parts, namely the temperature penalty cost and the power consumption of the cooling system. The temperature penalty cost t cost is calculated as follows: t cost = cost1 + cost2 +... + cost n t min and t max respectively represent the minimum temperature and the maximum temperature of the safety threshold of the equipment cooled by the cooling system, The calculation formula of the reward function is expressed as: R i = -(αt cost + (1 - α)p i × w) α represents the weight of the power consumption of the cooling system in the reward function, p min and p max respectively represent the minimum power consumption and the maximum power consumption within the safety threshold of the cooling system. w is a normalization coefficient to ensure that the temperature penalty cost and the power consumption of the cooling system are in the same order of magnitude.