Liquid cooling heat dissipation power intelligent control method fusing multiple parameters of machine room

By constructing a liquid cooling heat dissipation control model based on reinforcement learning algorithms and optimizing the heat dissipation power of the liquid cooling system using multi-parameter data, the problems of high energy consumption and low cooling efficiency of existing liquid cooling systems are solved, achieving high efficiency, energy saving and safe operation.

CN121843094APending Publication Date: 2026-04-10DONGGUAN QIQIN PRECISION THERMAL CONDUCTIVITY TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-03
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing liquid cooling systems are unable to achieve high efficiency and energy saving while ensuring safe operation, cannot meet the heat dissipation control requirements of high-density computing rooms, and lack synergistic optimization of cooling efficiency, energy efficiency and safety.

Method used

A liquid cooling heat dissipation control model based on reinforcement learning algorithm is constructed. An environmental state matrix is ​​built using multi-parameter data. The reward function is optimized based on the uniformity of heat load distribution, the trend of heat load change, and the system energy efficiency ratio. Dynamic adjustment is carried out in combination with variable frequency cooling pump, regulating valve and cooling fan.

Benefits of technology

It improves the liquid cooling effect, achieves synergistic optimization of cooling efficiency and energy efficiency, and enhances the system's safety and energy efficiency ratio.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121843094A_ABST
    Figure CN121843094A_ABST
Patent Text Reader

Abstract

The invention provides a liquid cooling heat dissipation power intelligent control method fusing multiple parameters of a machine room, and relates to the technical field of data centers, and the method comprises the steps: building an environment state matrix through obtaining multi-parameter data representing the environment thermal degree of the machine room in real time; a liquid cooling heat dissipation control model based on a reinforcement learning algorithm is constructed, a strategy function takes the environment state matrix as input and takes the heat dissipation power adjusting quantity of the liquid cooling system as action output, and a reward function taking a thermal load distribution uniformity index, a thermal load change trend index and a system energy efficiency ratio index as optimization targets is constructed; parameters of the strategy function are iteratively updated through a reinforcement learning algorithm, so that the strategy function outputs a heat dissipation power adjusting quantity of a maximized reward function output value under a given environment state matrix; and inputting the environment state matrix constructed in real time into the strategy function, and outputting an optimal liquid cooling heat dissipation power control instruction. The problem that in the prior art, the cooling efficiency, the energy efficiency and the safety are difficult to collaboratively optimize is solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data centers, in particular to a liquid cooling heat dissipation power intelligent control method fusing multiple parameters of a computer room. BACKGROUND

[0002] With the continuous growth of data center computing power demand, the power consumption of single cabinet continues to rise, and liquid cooling has become the mainstream cooling method for high-density computer rooms. However, most existing liquid cooling systems rely on single-point temperature or fixed thresholds for control, making it difficult to adapt to complex and variable load conditions. Traditional control methods only aim to meet temperature standards, without considering system energy efficiency in a unified optimization framework. Liquid cooling system energy consumption accounts for a high proportion, and energy utilization efficiency is low. Due to the mutual restraint relationship between thermal distribution uniformity, temperature control response speed and system energy efficiency, the existing scheme lacks a collaborative optimization mechanism, making it difficult to achieve efficient energy saving while ensuring safe operation, and unable to meet the cooling control needs of new-generation high-density computing power computer rooms. SUMMARY

[0003] Embodiments of the present application provide a liquid cooling heat dissipation power intelligent control method fusing multiple parameters of a computer room, aiming to solve the problem of simultaneous optimization of cooling efficiency, energy efficiency and safety in the prior art.

[0004] In order to achieve the above-mentioned purpose, the present application provides a liquid cooling heat dissipation power intelligent control method fusing multiple parameters of a computer room, comprising the following steps: real-time acquisition of multiple parameter data representing the thermal degree of the computer room environment, construction of an environment state matrix; construction of a liquid cooling heat dissipation control model based on a reinforcement learning algorithm, the liquid cooling heat dissipation control model including a policy function, the policy function taking the environment state matrix as input and the heat dissipation power adjustment amount of the liquid cooling system as action output, and constructing a reward function with the thermal load distribution uniformity index, the thermal load change trend index and the system energy efficiency ratio index as optimization targets; training the liquid cooling heat dissipation control model using historical operation data, iteratively updating the parameters of the policy function through the reinforcement learning algorithm, so that the policy function outputs the heat dissipation power adjustment amount that maximizes the reward function output value under the given environment state matrix, and obtains the trained policy function; inputting the real-time constructed environment state matrix into the trained policy function, outputting the optimal liquid cooling heat dissipation power control instruction, and issuing it to the liquid cooling system actuator; The reward function is calculated by the following formula: R=-(a1U + a2T + a3E+P safe ), In the formula, R is the output value of the reward function; U is the heat load distribution uniformity index, calculated based on the variance or standard deviation of the temperature values ​​at each monitoring point in the computer room, and a1 is the weighting coefficient of the corresponding heat load distribution uniformity index; T is the heat load change trend index, calculated based on the rate of change of the average temperature of the computer room over time, and a2 is the weighting coefficient of the corresponding heat load change trend index; E is the system energy efficiency ratio index, calculated based on the ratio of the real-time power consumption of the liquid cooling system to the real-time heat load of the IT equipment, and a3 is the weighting coefficient of the corresponding system energy efficiency ratio index; P safe As a safety penalty item, when the heat load distribution uniformity index, heat load change trend index, and system energy efficiency ratio index all do not exceed their respective preset safety thresholds, P safe It is zero.

[0005] Furthermore, the multi-parameter data includes: server inlet and outlet temperatures, CPU utilization, GPU utilization, memory usage, liquid cooling medium flow rate, liquid cooling medium inlet and outlet temperatures, cold plate surface temperature, and / or pump speed.

[0006] Furthermore, the liquid cooling system actuator includes a variable frequency cooling pump, a regulating valve, and a cooling fan, and the cooling power adjustment amount corresponds to adjusting the operating frequency or opening degree of the liquid cooling system actuator.

[0007] Furthermore, the weighting coefficients a1, a2, and a3 are dynamically adjusted according to the operating conditions of the computer room, including one or more of the following: During the preset peak business hours, increase the weighting coefficients a1 and a2 of the heat load distribution uniformity index and the heat load change trend index. During the pre-defined off-peak business periods, increase the value of the weighting coefficient a3 of the system energy efficiency ratio indicator; When the temperature of a local hot spot is detected to exceed the preset hot spot threshold, the value of the weighting coefficient a1 of the heat load distribution uniformity index is increased.

[0008] Furthermore, the policy function is fitted using a deterministic policy network, which takes the environment state matrix as input and outputs the heat dissipation power adjustment amount; the reinforcement learning algorithm adopts a deep deterministic policy gradient algorithm and constructs a value network to evaluate the expected cumulative reward of the state-action pair.

[0009] Furthermore, before the optimal liquid cooling power control command is issued, it undergoes command smoothing processing, specifically: a weighted average of the output command at the current moment and the actual execution command at the previous moment is performed.

[0010] Furthermore, the environment state matrix is ​​constructed by rack partitioning, with each partition corresponding to an independent sub-state matrix.

[0011] Furthermore, the construction by rack partitioning specifically includes: Based on the physical layout of the server racks and cooling zones within the server room, the server room is divided into multiple control zones, each containing one or more adjacent server racks. Within each control area, the multi-parameter data of the monitoring points in that area are constructed into a sub-state matrix according to the spatial location distribution of that area; The sub-state matrix of each region is input in parallel to the corresponding local liquid cooling heat dissipation control model, and each local liquid cooling heat dissipation control model independently outputs the heat dissipation power adjustment amount of that region. The adjustment values ​​output by each local liquid cooling heat dissipation control model are collaboratively synthesized to generate a global liquid cooling heat dissipation power control command.

[0012] The above technical solution has the following technical effects: By acquiring multi-parameter data characterizing the thermal state of the computer room environment in real time, an environmental state matrix is ​​constructed. A liquid cooling heat dissipation control model based on a reinforcement learning algorithm is built. The policy function of the liquid cooling heat dissipation control model takes the environmental state matrix as input and the heat dissipation power adjustment of the liquid cooling system as the action output. A reward function is constructed with heat load distribution uniformity index, heat load change trend index, and system energy efficiency ratio index as optimization objectives. The liquid cooling heat dissipation control model is trained using historical operating data, and the parameters of the policy function are iteratively updated through the reinforcement learning algorithm so that the policy function outputs the heat dissipation power adjustment that maximizes the output value of the reward function under a given environmental state matrix, resulting in a trained policy function. The real-time constructed environmental state matrix is ​​input into the trained policy function to output the optimal liquid cooling heat dissipation power control command, which is then sent to the liquid cooling system actuator. This invention solves the problem of the difficulty in coordinating the optimization of cooling efficiency, energy efficiency, and safety in existing technologies.

[0013] In summary, this invention utilizes the heat load distribution uniformity index and the heat load change trend index to characterize the cooling efficiency of the system, and the system energy efficiency ratio index to characterize the system energy efficiency. By incorporating the heat load distribution uniformity index, the heat load change trend index, and the system energy efficiency ratio index as optimization objectives into the reinforcement learning policy function, the recommended heat dissipation power adjustment action takes into account both cooling efficiency and energy efficiency, thereby improving the liquid cooling heat dissipation effect. Attached Figure Description

[0014] Figure 1 This is a flowchart illustrating an embodiment of the intelligent control method for multi-parameter liquid cooling power in a converged data center according to the present invention. Detailed Implementation

[0015] To further illustrate the various embodiments, the present invention provides accompanying drawings. These drawings are part of the disclosure of the present invention, primarily used to illustrate the embodiments and to explain the operating principles of the embodiments in conjunction with the relevant descriptions in the specification. With reference to these drawings, those skilled in the art should be able to understand other possible implementations and the advantages of the present invention. Components in the drawings are not drawn to scale, and similar component symbols are generally used to represent similar components.

[0016] The present invention will now be further described in conjunction with the accompanying drawings and specific embodiments.

[0017] Example 1 Figure 1 This is a flowchart illustrating an embodiment of the intelligent control method for multi-parameter liquid cooling power in a converged data center according to the present invention. Figure 1 As shown, the method includes the following steps: Real-time acquisition of multi-parameter data characterizing the thermal state of the computer room environment; construction of an environmental state matrix. First, various types of sensors are deployed within the data center server room. Temperature sensors are placed at the server's air inlets, outlets, and cold plate surfaces to collect server thermal status data; utilization data, including CPU utilization, GPU utilization, and memory usage, is obtained through the server baseboard management controller; flow sensors are placed at the liquid cooling pipe supply and return ports to collect the flow rate and temperature of the liquid cooling medium; and pump speed sensors are connected to the liquid cooling pump control cabinet to collect the pump's operating status.

[0018] The data acquisition cycle is set to 10 seconds. The acquired multi-parameter data includes: server inlet temperature, server outlet temperature, CPU utilization, GPU utilization, memory usage, liquid cooling medium flow rate, liquid cooling medium supply temperature, liquid cooling medium return temperature, cold plate surface temperature, and pump speed.

[0019] The collected raw data undergoes preprocessing. First, outlier detection is performed by calculating the mean and standard deviation of historical data for each monitoring point. The current collected value is compared with the historical mean; if the deviation exceeds a preset threshold, it is identified as an outlier. For identified outliers, the previous normal value for that monitoring point is used as a replacement. Then, the data is normalized to eliminate differences between different units of measurement.

[0020] The preprocessed data is arranged according to the physical location of the monitoring points in the computer room to construct a two-dimensional environmental state matrix. The rows of the matrix correspond to the column numbers of the server racks, and the columns correspond to the row numbers of the server racks. Each matrix cell contains multi-parameter data of the monitoring point at that location. To reflect the temporal evolution of the thermal environment, the two-dimensional matrix at the current moment is stacked with the two-dimensional matrices at five consecutive historical moments along the time axis to form a three-dimensional state tensor, which serves as the input for subsequent models.

[0021] A liquid cooling heat dissipation control model based on reinforcement learning algorithm is constructed. The liquid cooling heat dissipation control model includes a policy function, which takes the environmental state matrix as input and the heat dissipation power adjustment of the liquid cooling system as the action output. A reward function is constructed with the heat load distribution uniformity index, heat load change trend index and system energy efficiency ratio index as optimization objectives. In a specific implementation, the heat dissipation power adjustment is a normalized continuous value, ranging from [-1, 1], where positive values ​​indicate an increase, negative values ​​indicate a decrease, and the absolute value indicates the adjustment range.

[0022] In one specific implementation, the reward function is calculated using the following formula: R = -(a1U + a2T + a3E + P) safe ), In the formula, R is the output value of the reward function; U is the heat load distribution uniformity index, calculated based on the variance or standard deviation of the temperature values ​​at each monitoring point in the computer room. The smaller the U value, the more uniform the temperature distribution; a1 is the weighting coefficient of the corresponding heat load distribution uniformity index; T is the heat load change trend index, calculated based on the rate of change of the average temperature of the computer room over time; a2 is the weighting coefficient of the corresponding heat load change trend index; E is the system energy efficiency ratio index, calculated based on the ratio of the real-time power consumption of the liquid cooling system to the real-time heat load of the IT equipment. The smaller the E value, the more heat is removed per unit of power consumption; a3 is the weighting coefficient of the corresponding system energy efficiency ratio index; P safe As a safety penalty item, when the heat load distribution uniformity index, heat load change trend index, and system energy efficiency ratio index all do not exceed their respective preset safety thresholds, P safe The value is zero. When any of the heat load distribution uniformity index, heat load change trend index, and system energy efficiency ratio index exceeds their respective preset safety thresholds, the corresponding safety penalty item value is obtained through a custom index-safety penalty coefficient lookup table.

[0023] The liquid cooling heat dissipation control model is trained using historical operating data. The parameters of the policy function are iteratively updated through a reinforcement learning algorithm so that the policy function outputs the heat dissipation power adjustment amount that maximizes the output value of the reward function under a given environmental state matrix, thus obtaining the trained policy function. Specifically, the environmental state matrix in historical data is used as the input to the policy function. The heat dissipation power adjustment output by the policy function is compared with the actual adjustment in historical data, and the parameters of the policy function are iteratively updated through a reinforcement learning algorithm.

[0024] An experience replay mechanism is employed during training. The state, action, reward, and next state data generated from each interaction are stored in the experience replay pool. When the amount of data in the experience replay pool reaches a preset threshold, a small batch of data is randomly sampled for network parameter updates. The training objective is to enable the policy function, given an environmental state matrix, to output the heat dissipation power adjustment amount that maximizes the reward function's output value. During training, the model iterates through different adjustment strategies, learning from the feedback of the reward function which adjustment method can achieve better uniformity, faster trend response, and higher energy efficiency while ensuring temperature safety.

[0025] After training, a policy function with converged parameters is obtained, which is the policy function after training is complete.

[0026] The real-time constructed environmental state matrix is ​​input into the trained policy function, which outputs the optimal liquid cooling power control command and sends it to the liquid cooling system actuator. Specifically, in actual operation, online control steps are executed in each control cycle. Multi-parameter data from each monitoring point are collected in real time, and the current environmental state matrix is ​​input into the trained strategy function. The strategy function calculates forward and outputs the heat dissipation power adjustment amount. Based on the heat dissipation power adjustment amount output by the strategy function, specific liquid cooling power control commands are generated. These commands include the operating frequency adjustment value of the variable frequency cooling pump, the opening adjustment value of the regulating valve, and the speed adjustment value of the cooling fan. The control commands are sent to the actuators of the liquid cooling system, including the variable frequency cooling pump, the regulating valve, and the cooling fan. Each actuator adjusts its operating state according to the commands, achieving dynamic adjustment of the liquid cooling power.

[0027] Before issuing instructions, instruction smoothing is performed. A weighted average is calculated between the current output instruction and the actual executed instruction from the previous time step. Instruction smoothing avoids frequent starts and stops and drastic adjustments to the actuator caused by fluctuations in the strategy function output.

[0028] Example 2 This embodiment, based on Embodiment 1, further dynamically adjusts the weighting coefficients.

[0029] In actual operation, the values ​​of weighting coefficients a1, a2, and a3 are dynamically adjusted based on the operating conditions of the computer room. The specific adjustment strategy is as follows: During peak business hours, the data center is under heavy load and sensitive to temperature fluctuations. At this time, increasing the weight coefficients a1 and a2 of the heat load distribution uniformity index and the heat load change trend index makes the control model pay more attention to temperature uniformity and trend response.

[0030] During off-peak business hours, when the data center load is low and the temperature is within a safe range, increasing the weighting coefficient a3 of the system energy efficiency ratio index will make the control model pay more attention to energy saving.

[0031] When real-time monitoring detects that the temperature of a local hot spot exceeds the preset hot spot threshold, the value of the weight coefficient a1 of the heat load distribution uniformity index is increased to guide the model to eliminate local hot spots first. After the hot spots are eliminated, the original weight is restored.

[0032] Through the above dynamic weight adjustment, the control model can adapt to the optimization needs under different operating conditions, and achieve better control effect while ensuring temperature safety.

[0033] Example 3 This embodiment describes the specific implementation of the strategy function based on Embodiment 1.

[0034] The policy function is fitted using a deterministic policy network, the structure of which is as follows: The input layer receives a three-dimensional state tensor with dimensions m×n×t, where m and n correspond to the grid division of the computer room, and t corresponds to the time step.

[0035] The first layer is a three-dimensional convolutional layer with a kernel size of 3×3×3, a stride of 1, a padding method of "same", an output channel count of 32, and the activation function is ReLU.

[0036] The second layer is a three-dimensional convolutional layer with a kernel size of 2×2×2, a stride of 1, a padding method of "same", an output channel count of 64, and the activation function is ReLU.

[0037] The feature maps output from the convolutional layers are flattened before being input into the fully connected layers. The first fully connected layer has 512 neurons and uses ReLU activation; the second fully connected layer has 256 neurons and also uses ReLU activation.

[0038] The output layer is a linear layer with 1 neuron. The activation function is tanh, and the output value range is [-1, 1], corresponding to the normalized heat dissipation power adjustment.

[0039] The reinforcement learning algorithm employs a deep deterministic policy gradient algorithm and constructs a value network to evaluate the expected cumulative reward of a state-action pair. The structure of the value network is similar to that of the policy network, but its input includes the environment state matrix and the action values ​​output by the policy network, and its output is a Q-value estimate of that state-action pair.

[0040] The loss function of the value network is the mean squared error of the temporal difference error, while the loss function of the policy network is the negative expected value of the value network output. Stable model training is achieved through alternating updates of the policy network and the value network.

[0041] Example 4 This embodiment provides a detailed explanation of partition control based on Embodiment 1.

[0042] For large-scale data center server rooms, a partitioning control strategy is adopted to reduce computational complexity. The environment state matrix is ​​constructed by rack partitioning, with each partition corresponding to an independent sub-state matrix. The specific steps are as follows: Based on the physical layout of the server racks and cooling zones within the data center, the data center is divided into multiple control zones. The division is based on factors including the physical location of the racks, the topology of the cooling piping, and historical heat load distribution characteristics. Each control zone contains one or more adjacent racks, and there may be partial overlap between zones to ensure a smooth transition.

[0043] For example, the 32 cabinets are divided into 8 control areas in a 4x8 layout, with each area containing 4 cabinets in 2x2 columns.

[0044] Within each control area, the multi-parameter data of the monitoring points in that area are used to construct a sub-state matrix according to their spatial distribution. The construction method of the sub-state matrix is ​​the same as that of the global environment state matrix, but it only contains the monitoring point data within that area.

[0045] A corresponding local liquid cooling control model is constructed for each control region. The structure and algorithm of each local model are the same as those of the global model, and they can be trained independently.

[0046] The sub-state matrices of each region are input in parallel to the corresponding local models, and each local model independently outputs the heat dissipation power adjustment amount for that region.

[0047] The adjustment values ​​output by each local model are collaboratively synthesized to generate a global liquid cooling power control command. During collaborative synthesis, the thermal coupling effect between regions is considered. When there are temperature differences between adjacent regions, a cooling capacity allocation term between regions is added when synthesizing the global command to ensure the overall temperature distribution is coordinated.

[0048] By using zone control, the control problem of large-scale data centers is decomposed into multiple small-scale sub-problems, which reduces the computational complexity of single model inference, improves the real-time performance of control response, and enables precise adjustment of local hotspots.

[0049] Although the invention has been specifically shown and described in conjunction with preferred embodiments, those skilled in the art will understand that various changes in form and detail may be made to the invention without departing from the spirit and scope of the invention as defined in the appended claims, all of which shall be within the scope of protection of the invention.

Claims

1. A method for intelligent control of liquid cooling power in a computer room that integrates multiple parameters, characterized in that, Includes the following steps: Real-time acquisition of multi-parameter data characterizing the thermal state of the computer room environment; construction of an environmental state matrix. A liquid cooling heat dissipation control model based on reinforcement learning algorithm is constructed. The liquid cooling heat dissipation control model includes a policy function. The policy function takes the environmental state matrix as input and the heat dissipation power adjustment of the liquid cooling system as the action output. A reward function is constructed with the heat load distribution uniformity index, heat load change trend index and system energy efficiency ratio index as optimization objectives. The liquid cooling heat dissipation control model is trained using historical operating data. The parameters of the policy function are iteratively updated through a reinforcement learning algorithm so that the policy function outputs the heat dissipation power adjustment amount that maximizes the output value of the reward function under a given environmental state matrix, thus obtaining the trained policy function. The real-time constructed environmental state matrix is ​​input into the trained policy function, which outputs the optimal liquid cooling power control command and sends it to the liquid cooling system actuator. The reward function is calculated using the following formula: R=-(a1U + a2T + a3E+P) safe ), In the formula, R is the output value of the reward function; U represents the heat load distribution uniformity index, calculated based on the variance or standard deviation of temperature values ​​at each monitoring point in the computer room; a1 is the corresponding weighting coefficient for the heat load distribution uniformity index. T represents the heat load change trend index, calculated based on the rate of change of the average temperature in the computer room over time; a2 is the corresponding weighting coefficient for the heat load change trend index. E represents the system energy efficiency ratio index, calculated based on the ratio of the real-time power consumption of the liquid cooling system to the real-time heat load of the IT equipment; a3 is the corresponding weighting coefficient for the system energy efficiency ratio index. P safe As a safety penalty item, when the heat load distribution uniformity index, heat load change trend index, and system energy efficiency ratio index all do not exceed their respective preset safety thresholds, P safe It is zero.

2. The intelligent control method for liquid cooling power of multiple parameters in a converged data center according to claim 1, characterized in that, The multi-parameter data includes: server inlet and outlet temperatures, CPU utilization, GPU utilization, memory usage, liquid cooling medium flow rate, liquid cooling medium inlet and outlet temperatures, cold plate surface temperature, and / or pump speed.

3. The intelligent control method for liquid cooling power of multiple parameters in a converged data center according to claim 1, characterized in that, The liquid cooling system actuator includes a variable frequency cooling pump, a regulating valve, and a cooling fan. The cooling power adjustment corresponds to adjusting the operating frequency or opening degree of the liquid cooling system actuator.

4. The intelligent control method for liquid cooling power of multiple parameters in a converged data center according to claim 1, characterized in that, The weighting coefficients a1, a2, and a3 are dynamically adjusted based on the operating conditions of the computer room, including one or more of the following: During the preset peak business hours, increase the weighting coefficients a1 and a2 of the heat load distribution uniformity index and the heat load change trend index. During the pre-defined off-peak business periods, increase the value of the weighting coefficient a3 of the system energy efficiency ratio indicator; When the temperature of a local hot spot is detected to exceed the preset hot spot threshold, the value of the weighting coefficient a1 of the heat load distribution uniformity index is increased.

5. The intelligent control method for liquid cooling power of multiple parameters in a converged data center according to claim 1, characterized in that, The policy function is fitted using a deterministic policy network, which takes the environment state matrix as input and outputs the heat dissipation power adjustment amount. The reinforcement learning algorithm uses a deep deterministic policy gradient algorithm and constructs a value network to evaluate the expected cumulative reward of the state-action pair.

6. The intelligent control method for liquid cooling power of multiple parameters in a converged data center according to claim 1, characterized in that, Before the optimal liquid cooling power control command is issued, it undergoes command smoothing processing, specifically: a weighted average of the output command at the current moment and the actual execution command at the previous moment.

7. The intelligent control method for liquid cooling power of multiple parameters in a converged data center according to claim 1, characterized in that, The environment state matrix is ​​constructed by rack partitioning, with each partition corresponding to an independent sub-state matrix.

8. The intelligent control method for liquid cooling power of multiple parameters in a converged data center according to claim 7, characterized in that, The specific implementation of rack-based partitioning includes: Based on the physical layout of the server racks and cooling zones within the server room, the server room is divided into multiple control zones, each containing one or more adjacent server racks. Within each control area, the multi-parameter data of the monitoring points in that area are constructed into a sub-state matrix according to the spatial location distribution of that area; The sub-state matrix of each region is input in parallel to the corresponding local liquid cooling heat dissipation control model, and each local liquid cooling heat dissipation control model independently outputs the heat dissipation power adjustment amount of that region. The adjustment values ​​output by each local liquid cooling heat dissipation control model are collaboratively synthesized to generate a global liquid cooling heat dissipation power control command.

Citation Information

Cited By

  • Reinforcement learning-based data center high-density computing resource adaptive scheduling system

    CN122240283A