Control system and method for heat dissipation of machine room
By introducing gas-liquid heat pipe modules, monitoring units and control units into the computer room cooling system, combined with load monitoring and prediction models, the cooling mode and resource allocation are dynamically adjusted, solving the problems of delayed cooling system response and energy waste in existing technologies, and achieving efficient, energy-saving and safe cooling effects.
Patent Information
- Application Number
- CN202511007197.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2025-10-17
AI Technical Summary
Existing technologies rely on temperature differences to determine cooling modes and are unable to detect sudden increases in heat caused by load fluctuations in advance, resulting in energy waste and the risk of local overheating.
It uses gas-liquid heat pipe modules, monitoring units and control units, combined with load monitoring data and prediction models, to dynamically adjust the cooling mode and resource allocation, including air cooling, liquid cooling and gas-liquid collaborative cooling modes, and perform load prediction and cooling resource optimization through edge computing nodes.
It achieves rapid response of the cooling system, reduces energy waste, avoids local overheating, improves heat dissipation efficiency and system robustness, and adapts to various application scenarios of high-density server clusters.
Smart Images

Figure CN120812906A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer room heat dissipation, and in particular to a control system and method for computer room heat dissipation. BACKGROUND
[0002] With the rapid development of 5G, artificial intelligence and high-performance computing technology, the server density and power consumption of data centers and power plant computer rooms have increased significantly, and the traditional fixed threshold cooling system has been difficult to meet the needs of dynamic load changes. The existing technology (such as CN114980669A) realizes heat dissipation through a gas-liquid cooperative cooling mode (air cooling, liquid cooling, gas-liquid mixed cooling), but still has the following problems:
[0003] Response lag: relying on temperature difference (ΔT3-1, ΔT4-2) to judge the switching mode, unable to perceive the sudden increase in heat caused by load fluctuations in advance;
[0004] Waste of resources: continuous operation of high-energy liquid cooling mode at low load, causing energy waste;
[0005] Local overheating risk: when the server load suddenly increases, the cooling system cannot quickly adjust the resource allocation, which is easy to cause local hot spots.
[0006] Therefore, there is an urgent need for a cooling system that combines load prediction and dynamic resource scheduling to improve heat dissipation efficiency, reduce energy consumption and ensure equipment safety. SUMMARY
[0007] Therefore, the technical problems to be solved by the present application are: relying on temperature difference (ΔT3-1, ΔT4-2) to judge the switching mode, unable to perceive the sudden increase in heat caused by load fluctuations in advance; continuous operation of high-energy liquid cooling mode at low load, causing energy waste; when the server load suddenly increases, the cooling system cannot quickly adjust the resource allocation, which is easy to cause local hot spots.
[0008] The above technical problems are solved by the following technical solutions: the present application proposes a control system for computer room heat dissipation, which includes a refrigeration unit, which includes a gas-liquid heat dissipation heat pipe module arranged in the server heat generation area, and at least one cold source of refrigeration mode;
[0009] A monitoring unit for collecting temperature and load data during server operation;
[0010] A control unit having at least two cooling dynamic adjustment logics and a processor for executing the two cooling dynamic adjustment logics;
[0011] One cooling adjustment logic is a coarse adjustment logic based on the temperature difference of the gas / liquid entering and leaving the gas-liquid heat dissipation heat pipe module, and the other cooling adjustment logic is a fine adjustment logic based on load monitoring data and a prediction model, and the coarse adjustment logic is executed first to determine whether to start the cooling mode, and then the fine adjustment logic is executed to optimize the cooling resource allocation.
[0012] In a preferred embodiment of the control system for heat dissipation of a machine room according to the present application, the cold source comprises an air cooling unit, a liquid cooling unit and an outdoor cooling unit; wherein:
[0013] The air cooling unit is connected to the gas-liquid heat dissipation heat pipe module and is used for heat dissipation through gas circulation.
[0014] The liquid cooling unit is connected to the gas-liquid heat dissipation heat pipe module and is used for heat dissipation through cooling liquid circulation.
[0015] The outdoor cooling unit is connected to the air cooling unit and the liquid cooling unit respectively and is used for providing a cold source for the cooling medium.
[0016] In a preferred embodiment of the control system for heat dissipation of a machine room according to the present application, the monitoring unit comprises:
[0017] The load monitoring module is installed on the server and reads the CPU / GPU utilization and power consumption of the server through an interface.
[0018] The temperature acquisition module acquires the inlet temperature T1, the outlet temperature T3, the inlet liquid temperature T2 and the outlet liquid temperature T4 of the gas-liquid heat dissipation heat pipe module in real time.
[0019] In a preferred embodiment of the control system for heat dissipation of a machine room according to the present application, the control unit deploys an edge computing node, which has a built-in prediction model, the prediction model is a lightweight load-heat generation mapping model, the input variables include CPU / GPU utilization, memory occupancy and network traffic, the output is ΔP and hotspot area temperature trend, and the Modbus / TCP protocol is used for communication with the refrigeration unit.
[0020] In a preferred embodiment of the control system for heat dissipation of a machine room according to the present application, the cooling dynamic adjustment logic strategy divides the cooling area according to the server priority, and allocates independent liquid cooling flow channels for different priority areas.
[0021] In a preferred embodiment of the control system for heat dissipation of a machine room according to the present application, the gas-liquid collaborative cooling efficiency is optimized according to the server density, and the optimization mode is to dynamically adjust the area ratio of the liquid cooling flow channel and the air cooling flow channel, and the area ratio of the liquid cooling flow channel is 20%-70%, and the area ratio of the air cooling flow channel is 30%-80%.
[0022] The technical problems are solved by the following technical solutions: The application further provides a control method for heat dissipation of a machine room, comprising the following steps:
[0023] S1: system startup, initialization parameters (T Air -1(max), T Liquid -1(max), T Liquid -2(min), T combine-air , T combine-Liquid );
[0024] S2: start the air cooling mode, collect the air inlet temperature T1 and the air outlet temperature T3 of the gas-liquid heat dissipation heat pipe module, and calculate the gas temperature difference ΔT3-1;
[0025] S3: based on the server load data collected by the load monitoring module, run the load-heat generation mapping model to predict the heat generation ΔP;
[0026] S4: dynamically switch the cooling mode according to the ΔP value:
[0027] When ΔP> 50W / s, start the liquid cooling mode and close the air cooling unit;
[0028] When ΔP fluctuates in the range of 20-50W / s, enter the mixed mode (air cooling + liquid cooling);
[0029] When ΔP<5W / s, return to the air cooling mode;
[0030] S5: in the liquid cooling or mixed mode, further optimize the cooling mode switching in combination with the liquid inlet and outlet temperature difference ΔT4-2 of the gas-liquid heat dissipation heat pipe module;
[0031] S6: in the gas-liquid collaborative cooling mode, synchronously adjust the running power of the air cooling unit and the liquid cooling unit, and dynamically allocate the cooling resources based on the formula Pliquid cooling = α.ΔT + β.ΔP, wherein α is a temperature sensitivity coefficient, and β is a load fluctuation coefficient;
[0032] S7: continuously monitor the temperature difference of the gas-liquid heat dissipation heat pipe module and the load prediction result until a cooling cycle is completed.
[0033] In a preferred embodiment of the control method for heat dissipation of a machine room, the condition for the control unit to start the liquid cooling mode is If (ΔT>Tmax) ∨ (ΔP>Pthreshold).
[0034] In a preferred embodiment of the control method for heat dissipation of a machine room, in the gas-liquid collaborative cooling mode, the power pump power of the outdoor cooling unit is increased to 2P, and the cold air and the cooling liquid respectively pass through the upper and lower layer flow channels for counterflow heat exchange.
[0035] In a preferred embodiment of the control method for heat dissipation of a machine room according to the present application: in the hybrid mode, the power dynamic adjustment range of the liquid cooling unit is 0.5P-1.5P, the air cooling unit maintains the basic power range of 0.3P-1.0P, and the sum of the power of the liquid cooling unit and the power of the air cooling unit does not exceed the maximum power capacity of the system.
[0036] The present application has the following advantages: 1. The present application integrates BMC / IPMI interface and virtualization platform API to collect multi-dimensional load data such as CPU / GPU utilization, memory occupancy, network traffic of the server in real time, and uses a lightweight AI inference engine (such as LSTM or XGBoost model) on the edge computing node to perform load prediction to generate ΔP (predicted heat generation). This method not only relies on temperature difference (ΔT3-1, ΔT4-2) to determine the cooling mode switching, but also combines the prediction of future load change trend to perceive the sudden increase of heat caused by load fluctuation in advance, thereby realizing the rapid response of the cooling system. For example, the liquid cooling mode is started before the load surge to avoid the risk of overheating caused by delay.
[0037] 2. The traditional cooling system continuously runs the high-energy liquid cooling mode at low load, causing energy waste. The present application divides the cooling area according to the server priority through a dynamic resource allocation strategy, and allocates independent liquid cooling channels for different priority areas. In the case of low load, the system automatically reverts to air cooling mode or maintains hybrid mode with basic power, reducing unnecessary energy consumption. For example, when ΔP<5W / s, the system turns off the liquid cooling unit and relies only on the air cooling unit for heat dissipation, with energy saving effect up to 15%-20%. In addition, the system can dynamically adjust the area ratio of the liquid cooling channel according to the location of the hot spot area, optimize the spatial allocation efficiency of cooling resources, and further improve the energy efficiency ratio.
[0038] 3. The existing technology cannot quickly adjust resource allocation when the server load surges, which easily causes local hot spots and threatens equipment safety. The present application can start local intensive cooling measures in time by monitoring the temperature trend (such as rising, falling or stable) of the hot spot area in real time and combining the ΔP prediction value. For example, when it is predicted that the temperature trend of the hot spot area is "rising", the control unit immediately increases the power pump power of the liquid cooling channel in this area (such as from 1.0P to 1.5P), and adjusts the opening degree of the liquid cooling control valve (such as from 50% to 80%), to concentrate on cooling the hot spot. This precise local cooling resource allocation mechanism effectively avoids the risk of local overheating and ensures that the server is always within a safe working temperature range.
[0039] 4. The application realizes dynamic switching of cooling modes through composite condition judgment (If (ΔT > T_max) ∨ (ΔP > P_threshold) → start liquid cooling mode), and combines industrial-level operating systems (such as Linux RT) and Modbus TCP protocols to realize the issuance of real-time control instructions. The system is also equipped with an abnormality detection algorithm (such as Isolation Forest) that can identify sensor faults or data abnormalities and continuously optimize model parameters through an online learning mechanism to ensure stable operation when the server load suddenly changes or hardware fails. This intelligent closed-loop control system not only improves heat dissipation efficiency, but also enhances the robustness and reliability of the system.
[0040] 5. The cooling system proposed in the application can not only cope with the sudden high heat problem in high-density server clusters, but also achieve high efficiency and energy saving in low-load scenarios. Whether facing high-performance computing tasks with a large increase in GPU utilization or low-load states in daily maintenance, the system can dynamically adjust the cooling mode and resource allocation strategy according to real-time load data and prediction results. For example, in hybrid mode, the power of air cooling units and liquid cooling units is allocated in proportion (such as liquid cooling unit power 1.2P, air cooling unit power 0.5P), realizing the collaborative optimization of air cooling and liquid cooling. This design that covers a variety of application scenarios provides a flexible and efficient heat dissipation solution for data centers. BRIEF DESCRIPTION OF DRAWINGS
[0041] In order to more clearly illustrate the technical solutions of the embodiments of the application, the drawings of the embodiments of the application will be briefly introduced below. Obviously, the drawings described below only relate to some embodiments of the application, rather than limiting the application. Among them:
[0042] Figure 1 The location structure diagram of the control system of the computer room cooling is shown;
[0043] Figure 2 The overall working flow chart of the control of the computer room cooling is shown;
[0044] Figure 3 The intelligent decision-making flow chart of the cooling mode of the control of the computer room cooling is shown;
[0045] Figure 4 The dynamic resource allocation flow chart of the control of the computer room cooling is shown.
[0046] In the drawings:
[0047] 100, refrigeration unit; 101, gas-liquid heat dissipation heat pipe module; 102, cold source; 102a, air cooling unit; 102b, liquid cooling unit; 102c, outdoor cooling unit; 200, monitoring unit; 300, control unit; m, server. DETAILED DESCRIPTION
[0048] For those skilled in the art to better understand the present application, the present application will be further described in detail below in conjunction with the specific embodiments and the accompanying drawings.
[0049] The terms used in the present application are those general terms currently widely used in the art in consideration of the functions about the present application, but the terms can be changed according to the intention of those skilled in the art, precedents, or new technology in the art. In addition, specific terms can be selected by the applicant, and in this case, the detailed meaning thereof will be described in the detailed description of the present application. Therefore, the terms used in the specification should not be understood as simple names, but based on the meaning of the terms and the overall description of the present application.
[0050] Reference Figure 1 The present embodiment provides a control system for heat dissipation of a machine room, comprising a refrigeration unit 100, which comprises a gas-liquid heat dissipation heat pipe module 101 arranged in a server heat generation area, and at least one cold source 102 of a refrigeration mode; a monitoring unit 200 for collecting temperature and load data of the server during operation; a control unit 300 having at least two cooling dynamic adjustment logics and a processor for executing the two cooling dynamic adjustment logics; wherein one cooling adjustment logic is a coarse adjustment logic based on the temperature difference of the gas / liquid entering and leaving the gas-liquid heat dissipation heat pipe module 101, and the other cooling adjustment logic is a fine adjustment logic based on load monitoring data and a prediction model, and the coarse adjustment logic is executed first to determine whether to start the cooling mode, and then the fine adjustment logic is executed to optimize the cooling resource allocation.
[0051] The cold source 102 comprises an air cooling unit 102a, a liquid cooling unit 102b, and an outdoor cooling unit 102c; wherein: the air cooling unit 102a is connected with the gas-liquid heat dissipation heat pipe module 101, and is used for heat dissipation through gas circulation; the liquid cooling unit 102b is connected with the gas-liquid heat dissipation heat pipe module 101, and is used for heat dissipation through cooling liquid circulation; and the outdoor cooling unit 102c is connected with the air cooling unit 102a and the liquid cooling unit 102c respectively, and is used for providing a cold source for the cooling medium.
[0052] The monitoring unit 200 comprises: a load monitoring module installed on the server to read the CPU / GPU utilization and power consumption of the server through an interface; and a temperature acquisition module for acquiring the inlet temperature T1, the outlet temperature T3, the inlet liquid temperature T2, and the outlet liquid temperature T4 of the gas-liquid heat dissipation heat pipe module 101 in real time.
[0053] A specific control system for heat dissipation of a computer room includes a gas-liquid heat dissipation heat pipe module 101 arranged on a server to be cooled. The gas-liquid heat dissipation heat pipe module 101 includes a heat pipe heat conduction unit and a heat exchange unit. The heat pipe heat conduction unit includes a heat sink and a heat pipe. The heat exchange unit includes a gas-liquid cold plate and a sealing plate cover. The heat sink is provided with a through groove accommodating an evaporation section of the heat pipe. The evaporation section of the heat pipe is fixed to the heat sink by soldering, so as to ensure that a flattened plane obtained by flattening the evaporation section of the heat pipe is flush with the bottom surface of the heat sink. Bolt holes on the heat sink are used to fix the entire gas-liquid heat dissipation heat pipe module on the surface of a CPU coated with heat-conductive silicone grease. The heat pipe includes an evaporation section, an adiabatic section, and a condensation section. The evaporation section is flat and has a flattened bottom. The wick structure inside the heat pipe is of a groove type or a composite type. The gas-liquid cold plate is provided with a through groove accommodating the condensation section of the heat pipe. The condensation section of the heat pipe is fixedly connected to the gas-liquid cold plate by soldering. The upper layer of the gas-liquid cold plate is a liquid cooling flow channel and is provided with a liquid inlet and a liquid outlet. The lower layer of the gas-liquid cold plate is an air cooling flow channel and is provided with an air inlet and an air outlet. The upper and lower flow channels are symmetrically distributed and are not connected to each other. The volume of the air cooling flow channel is greater than that of the liquid cooling flow channel. The heat sink collects heat generated by the CPU to the evaporation section of the heat pipe, so that the working medium in the evaporation section absorbs heat to evaporate into steam. The steam carries heat and flows through the inner core of the adiabatic section of the heat pipe to the condensation section under the action of pressure difference. The steam in the condensation section transfers the heat to the heat exchange unit and then liquefies again under the action of wick capillary force and gravity, and flows back to the evaporation section of the heat pipe to complete a heat transfer cycle.
[0054] The air cooling unit 102a is connected to the gas-liquid heat dissipation heat pipe module 101 and is used for heat dissipation through gas circulation. The air cooling unit 102a includes an air compressor, an air storage tank, an air cooling control valve, an air cooling filter, an air cooling plate heat exchanger, a gas flow meter, an air inlet temperature sensor, and an air outlet temperature sensor. The air compressor compresses gas and stores it in the air storage tank. The air cooling control valve is used to disconnect or close the air cooling circuit. The air cooling filter is used to filter impurities mixed in the gas after circulation. The air cooling plate heat exchanger cools the gas flowing out of the liquid storage tank. The gas flow meter is used to monitor the gas flow in the circulation loop of the air cooling unit 102a. The air inlet temperature sensor and the air outlet temperature sensor are used to monitor the gas temperature at the air inlet and the air outlet of the gas-liquid heat dissipation heat pipe module 101, respectively.
[0055] The liquid cooling unit 102b is connected with the air-liquid heat dissipation heat pipe module 101, and is used for heat dissipation through circulation of cooling liquid; wherein the liquid cooling unit 102b comprises a liquid storage tank, a power pump, a liquid cooling filter, a liquid cooling control valve, a liquid flow meter, an inlet liquid temperature sensor, an outlet liquid temperature sensor and a liquid cooling plate heat exchanger. The liquid storage tank is used for storing and recycling the cooling liquid; the power pump is used as the power source of the circulation loop of the liquid cooling unit 102b; the liquid cooling filter is used for filtering impurities mixed in the cooling liquid after circulation for multiple times; the liquid cooling control valve is used for disconnecting or closing the liquid cooling loop; the liquid flow meter is used for monitoring the liquid flow in the circulation loop of the liquid cooling unit 102b; the inlet liquid temperature sensor and the outlet liquid temperature sensor are respectively used for monitoring the liquid temperature at the inlet and outlet of the air-liquid heat dissipation heat pipe module 101; and the liquid cooling plate heat exchanger is used for cooling the cooling liquid after absorbing heat.
[0056] The outdoor cooling unit 102c is connected with the air cooling unit 102a and the liquid cooling unit 102b respectively, and is used for providing a cold source for the cooling medium; wherein the outdoor cooling unit 102c comprises a cooling tower and a power pump.
[0057] The load monitoring module integrates a BMC (baseboard management controller) / IPMI (intelligent platform management interface) interface and a virtualization platform API, and is used for collecting the CPU / GPU utilization and power consumption data of the server m in real time. The load monitoring module communicates with the server m hardware through a standard protocol such as IPMI2.0, and directly obtains key operating parameters including but not limited to processor core temperature, CPU occupancy, GPU load, memory usage, network throughput and whole machine power consumption from the BMC chip; at the same time, the module also accesses the virtualization platform (such as VMware vSphere, Kubernetes, OpenStack, etc.) through the SNMP protocol or the RESTful API, and realizes fine-grained load monitoring at the containerized application and virtual machine instance level. The load monitoring module has multi-channel concurrent collection capability, can support simultaneous access of multiple physical server m nodes and virtual resources, and transmits the collected data to the edge computing node in the control unit 300 at a fixed period (for example, once per second). The module, as an important part of the system perception layer, provides a high-precision and low-delay input basis for subsequent cooling strategy decision-making.
[0058] The edge computing node is deployed inside the control unit 300, adopts an embedded high-performance computing platform, carries a lightweight AI inference engine, and is used to run a load-heat generation mapping model constructed based on machine learning. The model is trained based on historical operation data, and the input variables include but are not limited to CPU / GPU utilization, memory occupancy, network traffic, storage IO rate, etc., and the output variable is the predicted heat generation ΔP of the server m in a unit of time (unit: watt W) and the temperature change trend of the local hotspot area. The edge computing node has online learning capability and can continuously optimize the model prediction accuracy according to the actual operation feedback data. The output result not only includes the current heat generation prediction value, but also includes the heat generation change trend in the future time window, thereby providing a forward-looking basis for the dynamic allocation of cooling resources.
[0059] Further, the edge computing node generates a corresponding cooling resource allocation strategy based on the predicted heat generation ΔP and in combination with the current cooling system operation state (such as air cooling mode, liquid cooling mode, or gas-liquid collaborative cooling mode), environmental temperature and humidity, fluid circulation efficiency, and other factors. The strategy includes but is not limited to: adjustment of the fan speed of the air cooling unit 102a; adjustment of the power of the liquid cooling unit 102b; redistribution of the cooling medium flow; automatic switching of the cooling mode (such as switching from air cooling to liquid cooling or entering the gas-liquid collaborative cooling mode); and cooling resource scheduling between different priority server m areas. The cooling resource allocation strategy is fed back to the control unit 300 through ModbusTCP, CAN bus or other industrial communication protocols, and the control unit 300 coordinates the actuators (such as air compressors, power pumps, control valves, etc.) to complete the cooling system response action, thereby realizing closed-loop regulation and control of the heat load of the server m.
[0060] The control unit 300 is connected with the air cooling unit 102a, the liquid cooling unit 102b and the outdoor cooling unit 102c, respectively, for realizing dynamic switching of the air cooling mode, the liquid cooling mode and the gas-liquid collaborative cooling mode according to the temperature difference of the inlet and outlet gas and / or the temperature difference of the inlet and outlet liquid of the gas-liquid heat dissipation heat pipe module 101. The control unit 300 is electrically connected with the air compressor, the air cooling control valve, the gas flow meter, the inlet air temperature sensor and the outlet air temperature sensor of the air cooling unit 102a, and the power pump, the liquid cooling control valve, the liquid flow meter, the inlet liquid temperature sensor and the outlet liquid temperature sensor of the liquid cooling unit 102b, and the power pump of the outdoor cooling unit 102c. The control unit 300 is provided with an embedded main control chip (such as ARMCortex-A55 or IntelAtomx64 architecture), an industrial-grade operating system (such as LinuxRT), a communication interface module (such as CAN, ModbusTCP, Ethernet) and a data processing module, which can receive real-time data from various sensors, combine the load prediction result output by the edge computing node, and execute the cooling mode switching logic.
[0061] The division of labor between the edge computing node and the control unit 300 is clear: the edge computing node is responsible for running the load-heat generation mapping model, predicting the heat generation ΔP, and generating a cooling resource allocation strategy; the control unit 300 is responsible for receiving the strategy and performing specific cooling mode switching and resource adjustment operations. The interaction between the two is realized through the Modbus TCP protocol, and the data transmission format uses register address-data value pairs, for example, the predicted ΔP value is stored in register address 40001, and the cooling mode instruction is stored in register address 40002. The real-time requirement is to update once every second, and the edge computing node sends the prediction results and strategy to the control unit 300 within the last 10 ms of each second, and the control unit 300 completes parsing and performs corresponding operations within 100 ms after receiving the data, ensuring the real-time response capability of the system.
[0062] Its core function is to intelligently judge and execute one of the following cooling modes based on the gas inlet and outlet temperature difference (ΔT3-1 = T3-T1) and / or liquid inlet and outlet temperature difference (ΔT4-2 = T4-T2) of the gas-liquid heat dissipation heat pipe module 101, combined with the server m predicted heat generation ΔP:
[0063] Air cooling mode: only the air cooling unit 102a is turned on, suitable for low load and low heat generation scenarios;
[0064] Liquid cooling mode: only the liquid cooling unit 102b is turned on, suitable for high load and sudden heat generation scenarios;
[0065] Gas-liquid collaborative cooling mode: both air cooling and liquid cooling units 102b are turned on, and the power distribution is dynamically adjusted based on the formula Pliquid cooling = a.ΔT + β.ΔP, suitable for scenarios with large load fluctuations or local hotspot risks.
[0066] It should be noted that the gas inlet and outlet temperature difference (ΔT3-1 = T3-T1) refers to the difference between the inlet temperature T1 and the outlet temperature T3 in the gas-liquid heat dissipation heat pipe module 101. T1 is the initial temperature of the cooling air before entering the heat pipe module, and T3 is the temperature after heat exchange when leaving the heat pipe module. This temperature difference reflects how much heat the air has absorbed during the heat exchange process through the heat pipe module. A larger ΔT3-1 indicates that the air has taken away more heat, and at the same time it also indicates that the current air cooling unit 102a may need higher cooling capacity to maintain the normal operating temperature of the server m. Therefore, ΔT3-1 is one of the important indicators for evaluating the performance of the air cooling system, which helps the control system to determine whether to enhance or weaken the cooling intensity in the air cooling mode, so as to ensure that the server m will not be overheated and cause performance degradation or hardware damage.
[0067] The liquid inlet temperature difference (ΔT4-2 = T4-T2) refers to the difference between the inlet temperature T2 and the outlet temperature T4 of the gas-liquid heat dissipation heat pipe module 101. T2 is the initial temperature of the cooling liquid before entering the heat pipe module, and T4 is the temperature of the cooling liquid after heat exchange and leaving the heat pipe module. ΔT4-2 reflects how much heat the cooling liquid absorbs during the flow through the heat pipe module. If ΔT4-2 is large, it means that the cooling liquid absorbs a large amount of heat, and the current liquid cooling system may not be sufficient, and the working power of the liquid cooling unit 102b needs to be increased or switched to a stronger cooling mode (such as gas-liquid cooperative cooling mode). On the contrary, if ΔT4-2 is small, it means that the cooling liquid absorbs less heat, and the system can consider reducing the working power of the liquid cooling unit 102b to save energy. Therefore, ΔT4-2 is a key parameter for measuring the efficiency of the liquid cooling system, used to dynamically adjust the cooling strategy in the liquid cooling mode, to ensure that the server m is always in a safe working temperature range.
[0068] The server m predicts the heat generation amount ΔP, which is the amount of heat generated by the server m per unit time predicted by the load-heat generation mapping model running on the edge computing node based on the data collected by the load monitoring module in real time. This prediction model considers multiple factors such as CPU utilization, GPU utilization, memory usage, network traffic, and uses machine learning algorithms (such as XGBoost, LSTM, etc.) to establish the mapping relationship between load and heat generation. The value of ΔP represents the additional heat generated by the server m in the future period of time, and the unit is usually watts (W). A higher ΔP value means that the server m will generate more heat, and stronger cooling measures may be needed to prevent overheating; while a lower ΔP value indicates that the heat generation of the server m is stable or decreasing, and the existing cooling strategy may be sufficient. By predicting ΔP in real time and adjusting the cooling resource allocation strategy accordingly, the system can respond to potential overheating risks in advance, improve overall heat dissipation efficiency and prolong equipment life.
[0069] Further, the control unit 300 adopts a composite judgment condition to make cooling mode switching decisions, and the specific logic is as follows: If (AT > Tmax) V (AP > Pthreshold) -> start the corresponding cooling mode, where AT represents the temperature difference of gas or liquid, T_max is the preset maximum allowed temperature difference threshold, AP is the predicted heat generation rate, and P_threshold is the set power threshold. The gas-liquid collaborative cooling mode dynamically allocates the power of the air cooling and liquid cooling units 102b according to the real-time monitored AT and AP through the formula Pliquidcool = a.AT + b.AP to achieve the best heat dissipation effect. The mixed mode means that the air cooling unit 102a and the liquid cooling unit 102b run at the same time, but do not perform power dynamic allocation, and only maintain fixed power operation. For example, in the mixed mode, the liquid cooling unit 102b power is fixed at 1.0P, and the air cooling unit 102a power is fixed at 0.5P. The unit of AP (heat generation rate of change) is watt per second (W / s), which represents the change speed of the heat generation of the server m in a unit of time. Its calculation method is: AP = (current heat generation - last time heat generation) / sampling time interval. The unit of AT (temperature difference) is Celsius (°C), and the threshold values are set for the gas temperature difference (AT3-1 = T3-T1) and the liquid temperature difference (AT4-2 = T4-T2) respectively, for example, the maximum allowed value of the gas temperature difference is 10°C, and the maximum allowed value of the liquid temperature difference is 9°C.
[0070] In the cooling mode switching logic, the priority of ΔP is higher than that of ΔT. That is, when ΔP≥50W / s, the system is switched to the liquid cooling mode. When ΔP<50W / s, if ΔT4-2>9℃, the system enters the gas-liquid collaborative cooling mode, and the power of the air cooling unit 102b is dynamically adjusted according to the value of ΔP. For example, if ΔP is between 20-50W / s, the power of the liquid cooling unit 102b is adjusted to 0.5P-1.5P, and the air cooling unit 102a maintains the basic power of 0.3P-1.0P; if ΔP<20W / s, the system tries to reduce ΔT4-2by increasing the power of the air cooling unit 102a, and if ΔT4-2still exceeds 9℃ for more than 10 seconds, it is switched to the liquid cooling mode. The control unit 300 continuously receives and processes real-time data of ΔT and ΔP, and performs the following logic: when ΔT3-1>ΔT3-1_max(such as 9℃) or ΔT4-2>ΔT4-2_max(such as 9℃), it indicates that the current heat dissipation capacity is insufficient, and the more efficient liquid cooling mode needs to be started. For example, if ΔT4-2>9℃, it indicates that the heat absorbed by the liquid cooling unit 102b exceeds its design capacity, and the cooling effect needs to be enhanced; when ΔP>P_threshold(such as 50W / s), it indicates that the server m is about to generate a sudden high heat (such as GPU utilization jumps to 90%), and only the existing cooling mode may not be able to respond in time, so the liquid cooling mode needs to be started in advance. The control unit 300 performs Boolean logic operations in real time through an industrial operating system (such as LinuxRT), and if any of the above conditions is true (ΔT>T_max∨ΔP>P_threshold), it immediately sends a start instruction to the liquid cooling unit 102b through the ModbusTCP protocol, and closes the air cooling unit 102a. The specific operation includes: when the air cooling unit 102a is closed, a close instruction is sent to the air cooling control valve, the air cooling circuit is disconnected, and the operation of the air compressor and the air cooling plate heat exchanger is stopped; when the liquid cooling unit 102b is started, an open instruction is sent to the liquid cooling control valve, the liquid cooling circuit is established, and the power of the power pump is increased to 1.5P(maximum power), the cooling liquid flow is increased to 80% of the maximum value, so as to quickly remove heat; when optimizing the liquid cooling flow distribution, the hot spot area is located according to the ΔP prediction result, the liquid cooling flow area ratio is dynamically adjusted (such as from 30% to 50%), the flow ratio adjustment is realized through the electric partition or variable cross-section valve, and the local cooling capacity is enhanced. In terms of hardware cooperation, the sensor feedback continuously receives data of the inlet / outlet temperature sensors (T2, T4) of the liquid cooling unit 102b, verifies the operation effect of the liquid cooling unit 102b, the power pump adjustment controls the frequency of the power pump through the frequency converter (such as from 50Hz to 60Hz), dynamically matches the cooling demand, and the control valve opening adjustment adjusts the opening of the control valve to 80% according to the data of the liquid cooling flow pressure sensor, to ensure uniform distribution of the cooling liquid. Through the above process and hardware cooperation, the system realizes closed-loop control from condition triggering to liquid cooling mode starting, which not only guarantees the stability of the temperature of the server m, but also realizes precise allocation and energy efficiency optimization of the cooling resources.
[0071] Through this composite condition, the control unit 300 can predict the risk of overheating before the load mutation, start the efficient cooling mechanism in advance, and improve the system response speed and stability. In addition, the control unit 300 also has remote communication capability, which can be connected with the data center management system (DCIM) through Ethernet or wireless module, realize centralized monitoring, remote configuration and abnormal alarm function, further enhance the intelligent level and operation convenience of the system.
[0072] Further need to explain is: the input variables of the load-heat mapping model include CPU / GPU utilization, memory occupancy, network traffic, and the output variables are predicted heat (unit: W) and hotspot area temperature trend.
[0073] The load-heat mapping model of the application realizes high-precision prediction through real-time collection and processing of multi-dimensional input variables. The collection of CPU / GPU utilization depends on the server m hardware management interface (BMC / IPMI) to directly read real-time data, such as using IPMI2.0 protocol to obtain CPU core temperature and GPU load percentage and other key parameters, and through normalization processing (0-100%) combined with historical data to generate sliding window (such as 10 seconds window), calculate the average value and volatility to capture the instantaneous change and trend of load; The collection of memory occupancy is through the virtualization platform API (such as Kubernetes REST API or VMware vSphere SDK) to obtain the memory usage data of container / virtual machine level, further to process the usage of physical memory and virtual memory in layers, and combined with memory swap rate (SwapRate) to calculate memory pressure index, so as to quantify the memory load state of system running; The collection of network traffic includes through the statistical interface of network interface card (NIC) (such as SNMP protocol) or virtualization platform API to obtain real-time network throughput (unit: Gbps), and through time series analysis to extract peak traffic, burst traffic proportion and traffic distribution mode as supplementary features of model input, finally the three kinds of input variables are processed cooperatively, to provide accurate data basis for subsequent heat prediction and hotspot area temperature trend analysis.
[0074] The load-heat mapping model is built based on machine learning algorithms such as XGBoost, LSTM, or CNN-LSTM hybrid model. The training data of the model comes from the historical running logs of server m, including load data such as CPU / GPU utilization, memory occupancy, network traffic, and corresponding heat measurement values (collected by thermocouples or infrared thermometers). In terms of feature engineering, the original data is normalized, sliding window average is calculated, and volatility is analyzed to generate standardized feature vectors. The model structure consists of input layer, hidden layer and output layer: the input layer receives the normalized CPU / GPU utilization, memory occupancy, network traffic and other feature vectors, which are collected in real time by sensors or virtualization platforms and generated through data preprocessing to capture the dynamic characteristics of server m load; the hidden layer uses a multi-layer neural network or tree model structure, which extracts time series dependence (such as trends and periodicity of load changes) through LSTM layers, and uses XGBoost layers to handle the combined effects of discrete features (such as the cooperative load mode of different hardware components), thereby establishing a nonlinear relationship between load features and heat; the output layer generates two target variables: one is the predicted heat ΔP (unit: W), representing the heat change per unit time generated by server m; the other is the hotspot area temperature trend, which is predicted through spatial heat map analysis or local temperature gradient calculation, predicting the temperature change direction (such as rising, falling or stable) of the hotspot area in the next 10 seconds to 1 minute. The model parameter settings are as follows: taking LSTM as an example, the network contains 3 hidden layers with 128 nodes each, and the hyperparameters are set to learning rate 0.001, batch size 32, and training period 100; the loss function uses mean square error (MSE); the evaluation indicators include the coefficient of determination (R 2 ) and the mean absolute error (MAE), the R 2 of the model on the test set reaches 0.95, and the MAE is 3.2W. The online learning mechanism of the model is realized through incremental learning, and the model parameters are updated every hour with newly collected data to dynamically adapt to the long-term changes in the load mode of server m. The anomaly detection algorithm uses IsolationForest to identify and filter the prediction bias caused by sensor failure or data anomaly: when the deviation between the predicted value and the actual value exceeds the set threshold (such as 10%), an abnormal alarm is triggered and the model is retrained to ensure the reliability of the model output.
[0075] The generation of the predicted heat generation ΔP (unit: W) is based on the weighted regression calculation of the input features (CPU / GPU utilization, memory occupancy, network traffic) by the model, and the specific formula is ΔP = α · CPU / GPU utilization + β · memory occupancy + γ · network traffic, where α, β, γ are the weight coefficients obtained by model training, reflecting the contribution degree of different load features to the heat generation. This predicted value serves as the core basis for cooling mode switching: when ΔP > 50W / s, the control unit 300 triggers the liquid cooling mode to cope with high-load heating; when ΔP fluctuates in the range of 20-50W / s, the system enters the hybrid mode (air cooling + liquid cooling) to balance the heat dissipation demand and energy consumption; when ΔP < 5W / s, it reverts to the air cooling mode to save energy. In addition, the generation of the hotspot area temperature trend relies on local thermal map analysis or spatial distribution data of temperature sensors, by comparing the temperature changes of the hotspot area in the future time period (such as defining "rising trend" if rising more than 2℃ within 10 seconds, "falling trend" if falling more than 1℃, and "stable trend" if fluctuating less than 0.5℃), the potential overheating risk can be accurately warned. In application scenarios, if the hotspot area temperature trend is predicted to be "rising", the control unit 300 will immediately start local intensive cooling measures, such as increasing the liquid cooling flow rate or adjusting the air cooling direction to concentrate cooling; at the same time, according to the location of the hotspot area, the distribution proportion of the liquid cooling flow channel is dynamically adjusted (such as increasing the liquid cooling flow channel area ratio of this area from 30% to 50%), so as to optimize the spatial allocation efficiency of cooling resources. Through the dual driving of ΔP prediction and hotspot temperature trend, the system can actively intervene before the load mutation, significantly reduce the local overheating risk and improve the energy efficiency ratio of the overall heat dissipation system.
[0076] The conversion of the model output to the cooling strategy is based on the dual driving of the ΔP threshold judgment and the hotspot area temperature trend response. First, the control unit 300 compares the predicted ΔP value with the preset threshold to determine the switching of the cooling mode: when ΔP > 50 W / s, the liquid cooling mode is triggered, the air cooling unit 102a is turned off and the liquid cooling unit 102b is started, the power of the power pump is increased to 1.5P, and the cooling liquid flow is maximized to cope with high load heat; when ΔP fluctuates in the range of 20-50 W / s, the mixed mode (air cooling + liquid cooling) is entered, the power of the liquid cooling unit 102b is dynamically adjusted (0.5P-1.5P), and the air cooling unit 102a maintains the basic power (0.3P-1.0P); when ΔP < 20 W / s, the air cooling mode is returned to, and the liquid cooling unit 102b is turned off to save energy. Second, in response to the hotspot area temperature trend, when it is predicted that the hotspot area temperature trend is "rising" (such as the temperature rising by more than 2°C within 10 seconds), the control unit 300 sends instructions to the liquid cooling unit 102b through the ModbusTCP protocol to immediately increase the power of the power pump of the liquid cooling flow channel in the region (such as from 1.0P to 1.5P), and to enhance the local cooling capacity by adjusting the opening degree of the liquid cooling control valve (such as from 50% to 80%); at the same time, the area ratio of the liquid cooling flow channel to the air cooling flow channel is dynamically adjusted according to the hotspot area position (such as increasing the area ratio of the liquid cooling flow channel from 30% to 50%), to optimize the spatial allocation efficiency of the cooling resources. In the execution mechanism of the control unit 300, the system processes the model output result with millisecond-level delay, and communicates with the air cooling unit 102a and the liquid cooling unit 102b through the ModbusTCP protocol, and the specific operations include: adjusting the power pump frequency according to the ΔP value (such as from 50Hz to 60Hz) to dynamically match the cooling demand; adjusting the opening degree of the liquid cooling control valve (such as the opening degree of the control valve in a specific region from 50% to 80%) to accurately control the cooling liquid flow; in the mixed mode, the power of the air cooling unit 102a is proportionally distributed (such as the power of the liquid cooling unit 102b being 1.2P and the power of the air cooling unit 102a being 0.5P) to realize the collaborative optimization of air cooling and liquid cooling. Through the closed-loop control of the above logical chain, the system can actively intervene before the load mutation, which not only guarantees the stability of the server temperature, but also realizes the on-demand allocation and energy efficiency maximization of the cooling resources.
[0077] In terms of hardware support, the system realizes data acquisition and real-time processing through the deployment of high-precision sensor networks and edge computing nodes. The sensor network includes PT1000 temperature sensors, electromagnetic flowmeters, and pressure sensors, which are used to monitor the inlet / outlet temperature (T1, T3) of the gas-liquid heat pipe module 101, the inlet / outlet temperature (T2, T4) of the liquid cooling medium, and the flow rate of the cooling medium, respectively, to ensure the accuracy and real-time performance of the data. The edge computing node uses an embedded high-performance computing platform (such as NVIDIA Jetson AGXXavier) and a lightweight AI inference engine (such as TensorRT), which supports model inference updated once per second and meets the low-latency requirements in high-load scenarios. In terms of software support, the system includes three core modules: the data preprocessing module performs denoising, normalization, and feature engineering processing on the collected raw data to generate standardized model input feature vectors; the model inference module generates ΔP prediction values and hotspot area temperature trends based on the pre-trained LSTM-XGBoost hybrid model; and the control decision module dynamically generates cooling mode switching instructions (such as liquid cooling, hybrid, and air cooling modes) and resource allocation strategies (such as liquid cooling flow channel power adjustment and air cooling unit 102a power distribution) based on the ΔP prediction values and temperature trends. Through the deep collaboration of hardware and software, the system realizes full-link closed-loop control from data acquisition to decision execution.
[0078] In terms of implementation effect, the system significantly improves the prediction accuracy through multi-dimensional feature input (CPU / GPU utilization, memory occupancy, network traffic) and online learning mechanism, with an absolute value error of the model for ΔP less than 5%, effectively supporting the accurate formulation of cooling strategies. The control unit 300 can complete the full-process response from load prediction to cooling mode switching within 500ms, significantly reducing the risk of performance degradation or hardware damage of the server m due to overheating. In terms of dynamic optimization capability, the system realizes precise allocation of local cooling resources based on the hotspot area temperature trend (such as "rising", "falling", and "stable"), such as increasing the power of the liquid cooling flow channel pump in a specific area (from 1.0P to 1.5P) or adjusting the area ratio of liquid cooling / air cooling flow channels (such as increasing the liquid cooling flow channel ratio from 30% to 50%), which improves cooling efficiency while reducing energy waste, with energy saving effect reaching 15%-20%. In addition, the system uses an anomaly detection algorithm (such as IsolationForest) to identify sensor faults or data anomalies in real time, and combines a model adaptive update mechanism (such as online incremental learning) to ensure stable operation when the server m load suddenly changes or hardware fails, significantly enhancing the robustness and reliability of the system.
[0079] Referring to Figure 2 and Figure 3 , the embodiment provides a control method for data center cooling, including the following steps:
[0080] S1: System startup, initialize parameters (T Air -1(max), T Liquid -1(max), T Liquid -2(min), T combine-air , T combine-Liquid );
[0081] S2: Start air cooling mode, collect air inlet temperature T1 and air outlet temperature T3 of the air-liquid heat dissipation heat pipe module 101, and calculate the gas temperature difference ΔT3-1;
[0082] S3: Based on the server m load data collected by the load monitoring module, run the load-heat generation mapping model to predict the heat generation ΔP;
[0083] S4: According to the ΔP value, dynamically switch the cooling mode:
[0084] When ΔP> 50W / s, start liquid cooling mode and close air cooling unit 102a;
[0085] When ΔP fluctuates in the range of 20-50W / s, enter the mixed mode (air cooling+liquid cooling);
[0086] When ΔP<5W / s, fall back to air cooling mode;
[0087] S5: In liquid cooling or mixed mode, combine the inlet and outlet liquid temperature difference ΔT4-2 of the air-liquid heat dissipation heat pipe module 101 to further optimize the cooling mode switching;
[0088] S6: In the air-liquid collaborative cooling mode, the running power of the air cooling unit 102a and the liquid cooling unit 102b is adjusted synchronously, and the cooling resources are dynamically allocated based on the formula Pliquid cooling=α.ΔT+β.ΔP, where α is the temperature sensitivity coefficient and β is the load fluctuation coefficient;
[0089] S7: Continuously monitor the temperature difference of the air-liquid heat dissipation heat pipe module 101 and the load prediction result until the completion of a cooling cycle.
[0090] The aforementioned gas-liquid synergistic cooling mode is not simply a simultaneous use of both air cooling and liquid cooling, but rather an optimization of their cooperation through intelligent monitoring and dynamic adjustment to achieve the best cooling effect. In contrast, the hybrid mode is a more basic parallel working mode, which simultaneously enables both air cooling and liquid cooling units 102b, but lacks in-depth optimization of their interaction. Specifically, in the gas-liquid synergistic cooling mode, the system intelligently adjusts the power distribution of air cooling and liquid cooling units 102b, or even closes one of the units, according to the real-time monitored temperature difference (such as the air temperature difference ΔT3-1 and the liquid temperature difference ΔT4-2) and the predicted heat generation change (ΔP), to adapt to different cooling needs. In the hybrid mode, such detailed power adjustment or unit closing operation is usually not performed, and more often, the continuous operation of the two systems is maintained to cover a wide range of load conditions.
[0091] When the system is in the gas-liquid synergistic cooling mode, the key to determining whether to switch to the pure liquid cooling mode or continue to maintain the current mode lies in the comprehensive consideration of multiple factors. First, the system needs to monitor the air temperature difference ΔT3-1 and the liquid temperature difference ΔT4-2 in real time. If it is found that ΔT4-2 exceeds the set maximum allowable value T Liquid -1(max), which indicates that the current liquid cooling capacity is insufficient to cope with the heat generated by server m, at which point it is considered to switch to the pure liquid cooling mode, and the power of the liquid cooling unit 102b may need to be increased. Conversely, if ΔT4-2 is lower than the set minimum allowable value T Liquid -2(min), it means that the cooling provided by the liquid cooling is excessive, at which point it can be considered to fall back to the air cooling mode or maintain the current gas-liquid synergistic cooling mode, but reduce the liquid cooling power to save energy.
[0092] In addition, the predicted heat generation ΔP is also an important reference index. Through the load-heat generation mapping model on the edge computing node, the system can predict the heat generation change of server m in the future period of time. If the predicted heat generation ΔP exceeds a preset threshold P_threshold (e.g., 50W / s), it means that server m will generate a large amount of additional heat, and the gas-liquid synergistic cooling may not be able to meet the cooling demand, so it is necessary to switch to the pure liquid cooling mode to provide stronger cooling capacity. Conversely, if ΔP fluctuates between 20-50W / s and the temperature difference is also within a reasonable range, the gas-liquid synergistic cooling mode can be maintained, and the power distribution of the air cooling and liquid cooling units 102b can be dynamically adjusted according to the actual situation to ensure the best cooling efficiency.
[0093] In the gas-liquid collaborative cooling mode, how the adjustment part realizes is the key to ensure that the system can dynamically respond to server m load changes and temperature differences after the analysis of the judgment condition. Once the current working state is determined (such as whether to switch to pure liquid cooling mode or maintain gas-liquid collaborative cooling mode), the next step is the specific adjustment step.
[0094] When the system is in the gas-liquid collaborative cooling mode, if it is monitored that the liquid temperature difference ΔT4-2 between the inlet and outlet exceeds the set maximum allowed value T Liquid -1(max), which indicates that the current liquid cooling unit 102b's cooling capacity is insufficient to meet the server m heat dissipation demand. At this time, the control unit 300 will start a series of adjustment measures to enhance the cooling effect. First, it will increase the power of the power pump to increase the flow rate and flow of the cooling liquid, so as to take away heat faster. At the same time, it may also need to adjust the working parameters of the liquid cooling plate heat exchanger, such as increasing the heat exchange area or improving the heat exchange efficiency, to ensure that the cooling liquid can absorb more heat in the shortest time. In addition, in order to further optimize the cooling effect, the control unit 300 may dynamically adjust the proportion of liquid cooling flow and air cooling flow according to the actual temperature feedback, such as allocating more cooling resources to the liquid cooling path and reducing the working load of the air cooling unit 102a, to ensure that the heat dissipation efficiency of the whole system is maximized.
[0095] On the other hand, if it is monitored that ΔT4-2 is lower than the set minimum allowed value T Liquid -2(min), it means that the liquid cooling provides excess cooling, at which time it can be considered to fall back to the air cooling mode or maintain the current gas-liquid collaborative cooling mode but reduce the liquid cooling power to save energy. In this case, the control unit 300 will gradually reduce the power of the power pump, reduce the flow rate and flow of the cooling liquid, and may close part of the liquid cooling flow to achieve the purpose of energy saving. At the same time, the air cooling unit 102a continues to run to ensure that the basic heat dissipation demand is met. Through this flexible power adjustment mechanism, the system not only can effectively cope with the heat dissipation demand under different load conditions, but also can maximize the reduction of energy consumption and prolong the service life of the equipment.
[0096] The predicted heat generation ΔP is also an important reference index. If the load-heat generation mapping model on the edge computing node predicts that the heat generation ΔP exceeds a preset threshold P_threshold (for example, 50 W / s), it means that the server m will soon generate a large amount of additional heat, and only relying on gas-liquid cooperative cooling may not be able to meet the heat dissipation demand. At this time, the control unit 300 will immediately take action, switch to the pure liquid cooling mode, and may increase the power of the power pump and open all available liquid cooling channels to ensure that there is enough cooling liquid covering the key heat generation area. At the same time, the air cooling equipment such as the air compressor will be temporarily turned off or the power will be reduced to avoid unnecessary energy consumption. Conversely, if the ΔP fluctuation range is between 20-50 W / s and the temperature difference is also within a reasonable range, the gas-liquid cooperative cooling mode can be maintained, and the power distribution of the air cooling and liquid cooling unit 102b can be dynamically adjusted according to the actual situation. For example, in this case, the control unit 300 can use an adaptive algorithm to accurately calculate the optimal power distribution scheme according to real-time temperature data and load prediction results, so that the air cooling and liquid cooling unit 102b can work cooperatively, ensuring efficient heat dissipation performance and achieving effective use of energy.
[0097] During the whole process, the control unit 300 continuously monitors the system state through various integrated sensors (such as the inlet air temperature sensor, the outlet air temperature sensor, the inlet liquid temperature sensor, and the outlet liquid temperature sensor) and makes rapid response and adjustment according to these real-time data. For example, when detecting that the temperature of a certain local hotspot area abnormally rises, the control unit 300 can quickly respond to increase the cooling liquid flow of the liquid cooling channel corresponding to the area or adjust the air flow direction to concentrate on cooling the hotspot. This design concept based on real-time monitoring and intelligent control not only improves the flexibility and response speed of the system, but also greatly enhances the adaptability and reliability of the data center when facing complex and variable load environments. Ultimately, through such a precisely designed adjustment mechanism, the gas-liquid cooperative cooling mode can ensure the stable operation of the server m while achieving the ideal balance between high performance and low energy consumption.
[0098] In the present embodiment, specifically:
[0099] 1. System startup and initialization (S1)
[0100] System startup: When the server m in the data center or power plant room starts, the emergency control system also starts.
[0101] Initialization parameters: Set and initialize a series of key parameters, including:
[0102] T Air -1(max): Maximum allowable value of gas temperature difference (air cooling mode), usually set to 8-10℃.
[0103] T Liquid -1(max): maximum allowable value of liquid temperature difference (liquid cooling mode), usually set to 8-10℃.
[0104] T Liquid -2(min): minimum allowable value of liquid temperature difference (liquid cooling mode), usually set to 3-5℃.
[0105] T combine-air : minimum allowable value of gas temperature difference in gas-liquid cooperative cooling mode, usually set to 7-9℃.
[0106] T combine-Liquid : minimum allowable value of liquid temperature difference in gas-liquid cooperative cooling mode, usually set to 5-7℃.
[0107] These parameters will be used for subsequent temperature and load judgment to determine the optimal cooling mode.
[0108] 2. Start air cooling mode (S2)
[0109] Start air cooling mode: the system first starts air cooling mode, using air compressor, air tank, air cooling plate heat exchanger and other equipment, through gas circulation for preliminary heat dissipation.
[0110] Collect inlet and outlet temperatures: the control unit 300 collects the inlet temperature T1 and outlet temperature T3 of the gas-liquid heat dissipation heat pipe module 101, and calculates the gas temperature difference ΔT3-1 = T3-T1.
[0111] 3. Load data acquisition and heat generation prediction (S3)
[0112] Load monitoring module: integrated BMC / IPMI interface and virtualization platform API, real-time acquisition of server m's CPU / GPU utilization, power consumption data.
[0113] Edge computing node runs load-heat generation mapping model: running a lightweight prediction model (such as LSTM neural network, XGBoost) on the edge computing node, updating the load prediction result every second, outputting the predicted heat generation ΔP (unit: W) and hotspot area temperature trend.
[0114] 4. Dynamically switch cooling mode according to heat generation ΔP (S4)
[0115] Judge ΔP value:
[0116] When ΔP > 50 W / s: Directly start the liquid cooling mode, and close the air cooling unit 102a. At this time, the power pump will draw the cooling liquid from the storage tank, pass through the liquid cooling filter, the liquid cooling control valve, the liquid flow meter, enter the liquid cooling flow channel of the air-liquid heat dissipation heat pipe module 101, and return to the storage tank after absorbing heat.
[0117] When ΔP fluctuates in the range of 20-50 W / s: Enter the mixed mode (air cooling + liquid cooling). At this time, the air cooling unit 102a and the liquid cooling unit 102b operate simultaneously, the power of the liquid cooling unit 102b is dynamically adjusted (0.5P-1.5P), and the air cooling unit 102a maintains the basic power (0.3P-1.0P).
[0118] When ΔP < 5 W / s: Fall back to the air cooling mode, close the liquid cooling unit 102b, and reduce energy consumption.
[0119] 5. Further optimize the cooling mode switching (S5)
[0120] In the liquid cooling or mixed mode, continue to monitor the liquid temperature difference ΔT4-2 of the air-liquid heat dissipation heat pipe module 101:
[0121] When ΔT4-2 > T Liquid -1(max): Indicates that the current liquid cooling capacity is still insufficient, triggers the air-liquid collaborative cooling mode, and further enhances the heat dissipation effect.
[0122] When ΔT4-2 < T Liquid -2(min): Indicates that the cooling capacity provided by the liquid cooling is excessive, stops the liquid cooling mode, and falls back to the air cooling mode.
[0123] 6. Synchronous adjustment in the air-liquid collaborative cooling mode (S6)
[0124] In the air-liquid collaborative cooling mode, the operating power of the air cooling unit 102a and the liquid cooling unit 102b is synchronously adjusted, and the cooling resources are dynamically allocated based on the formula Pliquid cooling = a · ΔT + β · ΔP, where a is the temperature sensitivity coefficient and β is the load fluctuation coefficient.
[0125] Specific operation: The air compressor and the power pump work simultaneously, the liquid cooling control valve and the air cooling control valve are opened, the power of the outdoor cooling unit 102c is increased to 2P, and the cold air and the cooling liquid respectively pass through the upper and lower flow channels (air cooling flow channel and liquid cooling flow channel) for counterflow heat exchange.
[0126] 7. Continuous monitoring and optimization (S7)
[0127] Continuous monitoring of temperature difference and load prediction results: The control unit 300 continuously monitors the temperature difference (ΔT3-1 and ΔT4-2) of the air-liquid heat dissipation heat pipe module 101 and the load prediction results, to ensure that the system is always in the best cooling state.
[0128] Adjust optimization according to actual situation:
[0129] If further optimization is needed: for example, in liquid cooling mode, the area ratio of liquid cooling flow channel can be adjusted (such as 3:7) to optimize the gas-liquid collaborative cooling efficiency; in mixed mode, the power distribution of air cooling and liquid cooling units 102b can be adjusted to optimize the efficiency of the mixed mode; in air cooling mode, the power of the air cooling unit 102a can be adjusted to optimize the efficiency of the air cooling mode.
[0130] After optimization, return to the main flow, and finally merge into the unified end node W.
[0131] 8. End flow
[0132] All paths eventually merge into a unified end node W: whether through liquid cooling mode, mixed mode, or air cooling mode to complete the heat dissipation task, the system will turn off the corresponding cooling unit after ensuring the temperature stability of the server m, and end the cooling period.
[0133] Reference Figure 4 Further, another improved embodiment of the present application is that the dynamic resource allocation strategy divides the cooling area according to the priority of the server m, and allocates independent liquid cooling flow channels for different priority areas.
[0134] Implementation of dynamic resource allocation strategy
[0135] 1. Server m priority division and cooling area division
[0136] The dynamic resource allocation strategy of the present application divides the servers m in the computer room into different priority areas based on the functional importance, business continuity requirements and data sensitivity of the servers m, and allocates independent liquid cooling flow channels for each priority area. The specific implementation is as follows:
[0137] (1) Priority division standard.
[0138] The dynamic resource allocation strategy divides the servers m in the computer room into different priority areas based on the functional importance, business continuity requirements and data sensitivity of the servers m. The specific quantitative standard is as follows:
[0139] Core priority (P1): including core database servers m, key business application servers m, real-time transaction processing servers m, etc., whose functions are crucial to business continuity, with a quantitative standard of server m downtime cost exceeding 1000 yuan / hour, and data sensitivity level being the highest (such as involving user privacy, financial transaction data, etc.).
[0140] Medium priority (P2): including web server m, cache server m, middleware server m, etc., the function has certain influence on business continuity, but has certain redundancy ability, the quantitative standard is that the shutdown cost is between 100-1000 yuan / hour, and the data sensitivity level is medium.
[0141] Low priority (P3): including test server m, non-critical application server m, edge computing node, etc., the function has less influence on business continuity, the quantitative standard is that the shutdown cost is less than 100 yuan / hour, and the data sensitivity level is low.
[0142] (2) Cooling area division.
[0143] Each priority area corresponds to a group of independent liquid cooling channels (such as liquid cooling channel segmentation control). For example:
[0144] P1 area: allocate independent liquid cooling channel A, the liquid cooling channel area ratio is 50%, the power pump power is 1.5P, and the cooling liquid flow is 80% of the maximum value.
[0145] P2 area: allocate independent liquid cooling channel B, the liquid cooling channel area ratio is 30%, the power pump power is 1.0P, and the cooling liquid flow is 50% of the maximum value.
[0146] P3 area: allocate independent liquid cooling channel C, the liquid cooling channel area ratio is 20%, the power pump power is 0.5P, and the cooling liquid flow is 30% of the maximum value.
[0147] The cooling area realizes independence through physical isolation (such as cabinet partition, liquid cooling channel partition), and ensures that the cooling resources of each area do not interfere with each other.
[0148] 2. Implementation of dynamic resource allocation strategy.
[0149] The dynamic resource allocation strategy realizes the real-time adjustment of cooling resources through the cooperative action of the control unit 300, the load monitoring module and the edge computing node. The specific steps are as follows:
[0150] Step 1: Load data acquisition and priority identification:
[0151] The load monitoring module acquires the CPU / GPU utilization, memory occupancy, network traffic and other data of each server m in real time through the BMC / IPMI interface and virtualization platform API.
[0152] The edge computing node runs the load-heat mapping model to predict the heat ΔP of each priority area, and generates a cooling demand matrix in combination with the business priority label (P1 / P2 / P3).
[0153] Step 2: Cooling resource allocation decision.
[0154] The control unit 300 dynamically adjusts the liquid cooling channel parameters of each priority region according to the cooling demand matrix and pre-set resource allocation rules (e.g. Table 1):
[0155] P1 region: priority guarantee of cooling resources. When the predicted heat generation ΔP > 50 W / s, the power pump power of the liquid cooling channel A is automatically increased to 2P, the cooling liquid flow rate is increased to 90% of the maximum value, and the air cooling unit 102a is turned off.
[0156] P2 region: on-demand allocation of cooling resources. When ΔP fluctuates in the range of 20-50 W / s, the power pump power of the liquid cooling channel B is dynamically adjusted to 0.8P-1.2P, the cooling liquid flow rate is adjusted to 40%-60% of the maximum value, and the air cooling unit 102a maintains the basic power (0.3P-0.5P).
[0157] P3 region: minimization of cooling resources. When ΔP < 5 W / s, the power pump power of the liquid cooling channel C is reduced to 0.3P, the cooling liquid flow rate is adjusted to 20% of the maximum value, and the air cooling unit 102a maintains the basic power (0.1P).
[0158]
[0159] Step 3: Dynamic adjustment of liquid cooling channel parameters.
[0160] Liquid cooling channel area ratio adjustment: dynamically adjust the area ratio of each liquid cooling channel through electric partition or variable cross-section valve. For example, when the P1 region load surges, the area ratio of the liquid cooling channel A expands from 50% to 60%, while the liquid cooling channel area ratios of the P2 and P3 regions are reduced to 25% and 15% respectively.
[0161] Power pump power adjustment: control the power pump power of the liquid cooling channel through the frequency converter. For example, when the P1 region needs to enhance cooling, the power pump frequency is increased from 50Hz to 60Hz, and the cooling liquid flow rate is correspondingly increased.
[0162] Control valve switch adjustment: adjust the opening of the liquid cooling channel through the proportional electromagnetic valve. For example, when the P3 region load decreases, the control valve opening is reduced from 80% to 30%, reducing the cooling liquid flow rate.
[0163] 3. Cooling mode switching and redundancy guarantee.
[0164] Cooling mode switching logic:
[0165] When the ΔT4-2 (liquid inlet and outlet temperature difference) of a priority region exceeds the set threshold value (e.g. T Liquid -1(max) = 9℃), the control unit 300 triggers the cooling mode switching of the region:
[0166] P1 region: Directly enter pure liquid cooling mode, turn off air cooling unit 102a, and increase the power pump power of liquid cooling flow channel A to 2P.
[0167] P2 region: Enter mixed mode (air cooling + liquid cooling), adjust the power pump power of liquid cooling flow channel B to 1.2P, and maintain the basic power of air cooling unit 102a.
[0168] P3 region: Maintain air cooling mode, only start liquid cooling flow channel C when ΔT4-2>T Liquid -1(max).
[0169] Redundancy guarantee mechanism:
[0170] Liquid cooling flow channel redundancy: Each priority region of the liquid cooling flow channel is equipped with a standby flow channel (such as liquid cooling flow channel A', B', C'), when the main liquid cooling flow channel fails, the control unit 300 automatically switches to the standby flow channel, and adjusts the power pump power to maintain the cooling capacity.
[0171] Cross-region resource allocation: In the case of extreme load (such as P1 region ΔP>100W / s), the control unit 300 can temporarily allocate cooling resources from P2 and P3 regions, for example, reduce the cooling liquid flow of liquid cooling flow channel B in P2 region to 30%, and introduce part of the cooling liquid into liquid cooling flow channel A in P1 region.
[0172] The adjustment of the area ratio of the liquid cooling flow channel is realized by an electric partition. The opening angle of the electric partition has a linear relationship with the area ratio of the liquid cooling flow channel, and the opening angle range is 0°-90°, corresponding to the area ratio of the liquid cooling flow channel from 0%-100%. The control logic is: according to the server m load data collected by the load monitoring module and the cooling demand matrix output by the edge computing node, the control unit 300 controls the driving motor of the electric partition through pulse width modulation (PWM) signal to accurately adjust the area ratio of the liquid cooling flow channel. For example, when the P1 region load suddenly increases, the control unit 300 sends a PWM signal to drive the electric partition to adjust from the initial angle of 45° (corresponding to the area ratio of the liquid cooling flow channel 50%) to 54° (corresponding to the area ratio of the liquid cooling flow channel 60%).
[0173] 4. Implementation effect and advantage.
[0174] Efficient heat dissipation guarantee: Core server m always obtains priority cooling resources, avoiding business interruption caused by local overheating. Energy efficiency optimization: Reduce cooling resource consumption in low-priority regions under low load, reducing overall system energy consumption by 15%-20%. Fast response: Based on real-time load data and prediction models, the control unit 300 can complete dynamic allocation and mode switching of cooling resources within 500ms, significantly improving system response speed. Redundancy reliability: Liquid cooling flow channel redundancy and cross-region resource allocation mechanism ensure stable operation of key server m in the case of single-point failure or extreme load.
[0175] 5. Hardware and software are implemented cooperatively.
[0176] Hardware support: liquid cooling channel design: adopt modular liquid cooling channel structure, each channel is equipped with independent power pump, control valve, temperature sensor and flow meter. Control unit 300 configuration: embedded main control chip (such as ARM Cortex-A55) runs real-time operating system (Linux RT), communicates with edge computing node through Modbus TCP protocol, and realizes millisecond level control instruction issuing.
[0177] Software support: load-heat quantity mapping model: based on LSTM neural network, input variables include CPU / GPU utilization, memory occupancy, network traffic, and output variables are predicted heat quantity ΔP and hotspot area temperature trend. Dynamic allocation algorithm: weighted priority scheduling algorithm is adopted, and the load prediction result, cooling resource availability and business priority label are considered comprehensively to generate the optimal cooling strategy.
[0178] Through the above technical scheme, the application realizes the dynamic resource allocation strategy based on the server m priority, significantly improves the intelligent level and energy efficiency ratio of the data center heat dissipation system, and provides reliable technical support for the high-density, dynamic load computer room environment.
[0179] Finally, it should be pointed out that the above detailed description of the method and device is only an embodiment, and those skilled in the art can modify the embodiment in different ways without departing from the scope of the application.
Claims
1. A control system for heat dissipation in a computer room, characterized by: include, A refrigeration unit (100) comprising a gas-liquid heat dissipation heat pipe module (101) provided in a heat generating area of a server, and a cold source (102) of at least one refrigeration mode; A monitoring unit (200) for collecting temperature and load data of the server during operation; A control unit (300) having at least two cooling dynamic adjustment logics and a processor for executing the two cooling dynamic adjustment logics; Among them, one cooling adjustment logic is a coarse adjustment logic based on the temperature difference between the gas / liquid entering and exiting the gas-liquid heat pipe module (101), and the other cooling adjustment logic is a fine adjustment logic based on load monitoring data and a prediction model, and the coarse adjustment logic is first executed to determine whether to start the cooling mode, and then the fine adjustment logic is executed to optimize the cooling resource allocation.
2. The control system for heat dissipation in a computer room according to claim 1, characterized in that: The cold source (102) includes an air cooling unit (102a), a liquid cooling unit (102b) and an outdoor cooling unit (102c); wherein: An air cooling unit (102a) is connected to the gas-liquid heat dissipation heat pipe module (101) and is used for dissipating heat through gas circulation; A liquid cooling unit (102b) is connected to the gas-liquid heat dissipation heat pipe module (101) and is used for dissipating heat through cooling liquid circulation; The outdoor cooling unit (102c) is connected to the air cooling unit (102a) and the liquid cooling unit (102b) respectively, and is used to provide a cold source for the cooling medium.
3. The control system for heat dissipation in a computer room according to claim 1 or 2, characterized in that: The monitoring unit (200) comprises: The load monitoring module is installed on the server and reads the server CPU / GPU utilization and power consumption through the interface; The temperature acquisition module acquires the air inlet temperature T1, the air outlet temperature T3, the liquid inlet temperature T2 and the liquid outlet temperature T4 of the air-liquid heat dissipation heat pipe module (101) in real time.
4. The control system for heat dissipation in a computer room according to claim 1, characterized in that: The control unit (300) deploys an edge computing node with a built-in prediction model, which is a lightweight load-heat mapping model. The input variables include CPU / GPU utilization, memory occupancy and network traffic, and the output is ΔP and hot spot area temperature trend. The model communicates with the refrigeration unit (100) via the Modbus / TCP protocol.
5. The control system for heat dissipation in a computer room according to any one of claims 1, 2, and 4, characterized in that: The cooling dynamic adjustment logic strategy divides cooling areas according to server priorities and allocates independent liquid cooling channels to areas with different priorities.
6. The control system for heat dissipation in a computer room according to claim 5, characterized in that: The gas-liquid collaborative cooling efficiency is optimized according to the server density. The optimization method is to dynamically adjust the area ratio of the liquid cooling channel to the air cooling channel, and the liquid cooling channel area accounts for 20%-70%, and the air cooling channel area accounts for 30%-80%.
7. A control method for heat dissipation in a computer room, characterized by: The control system for heat dissipation in a computer room as described in any one of claims 1 to 6 further includes the following steps: S1: System startup, initialization parameters (T Air -1(max), T Liquid -1(max), T Liquid -2(min), T combine-air 、T combine-Liquid ); S2: Start the air cooling mode, collect the air inlet temperature T1 and the air outlet temperature T3 of the gas-liquid heat dissipation heat pipe module (101), and calculate the gas temperature difference ΔT3-1; S3: Based on the server load data collected by the load monitoring module, the load-heat mapping model is run to predict the heat output ΔP; S4: Dynamically switch cooling mode according to ΔP value: When ΔP>50W / s, start the liquid cooling mode and turn off the air cooling unit; When the ΔP fluctuation range is 20-50W / s, it enters the air-liquid mixing mode; When ΔP<5W / s, it returns to air cooling mode; S5: In liquid cooling or air-liquid mixed mode, collecting the liquid inlet temperature T2 and the liquid outlet temperature T4, and combining the inlet and outlet liquid temperature difference ΔT4-2 of the gas-liquid heat pipe module (101), further optimizing the cooling mode switching; S6: In the gas-liquid coordinated cooling mode, the operating power of the air cooling unit (102a) and the liquid cooling unit (102b) are synchronously adjusted, and cooling resources are dynamically allocated based on the formula P liquid cooling = α.ΔT + β.ΔP, where α is the temperature sensitivity coefficient and β is the load fluctuation coefficient; S7: Continuously monitor the temperature difference and load prediction result of the gas-liquid heat pipe module (101) until a cooling cycle is completed.
8. The control method for heat dissipation in a computer room according to claim 7, characterized in that: The condition for the control unit (300) to start the liquid cooling mode is: If (ΔT>T max )∨(ΔP>P threshold ).
9. The control method for heat dissipation in a computer room according to claim 7, characterized in that: In the gas-liquid coordinated cooling mode, the power pump of the outdoor cooling unit (102c) is increased to 2P, and the cold air and the cooling liquid exchange heat through the upper and lower flow channels respectively.
10. The control method for heat dissipation in a computer room according to any one of claims 7 to 9, characterized in that: In the hybrid mode, the power dynamic adjustment range of the liquid cooling unit (102b) is 0.5P-1.5P, the air cooling unit (102a) maintains a basic power range of 0.3P-1.0P, and the sum of the power of the liquid cooling unit (102b) and the power of the air cooling unit (102a) does not exceed the maximum power capacity of the system.
Citation Information
Patent Citations
Data center gas-liquid heat dissipation system and control method
CN114980669A
Cited By
Multi-mode water-cooling heat dissipation system and method
CN121397981A
A multi-modal water-cooled heat dissipation system and method
CN121397981B