Intelligent control method for heat dissipation of liquid-cooled data center based on dynamic flow regulation

By acquiring real-time CPU power and temperature distribution data of data center servers, generating adaptive hydraulic resistance, and dynamically adjusting the cooling flow of liquid-cooled data centers, the problem of not being able to dynamically adjust cooling flow in existing technologies is solved, achieving more efficient heat dissipation management and ensuring stable server operation.

CN120916411BActive Publication Date: 2026-01-27HEFEI YINGFAN ELECTRONIC TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511446943.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-11
Publication Date
2026-01-27
Estimated Expiration
2045-10-11

AI Technical Summary

Technical Problem

Existing liquid-cooled data center heat dissipation control methods cannot dynamically adjust the cooling flow rate according to the actual operating status of the server. This results in excessive cooling flow rate when the CPU power is low, leading to energy waste, and insufficient cooling flow rate when the CPU power jumps, failing to remove heat in time and affecting server performance and stability.

Method used

By acquiring real-time data on CPU power changes and temperature distribution of data center servers, an adaptive hydraulic impedance is generated. An impedance modulator and a digital potentiometer are used to control the electromagnetic flow valve, dynamically adjusting the flow parameters of the cooling branch to achieve real-time adjustment according to the server's heat dissipation requirements.

Benefits of technology

It enables real-time adjustment of cooling flow based on server heat dissipation needs, improving the intelligence and flexibility of the heat dissipation system and ensuring that the server maintains a stable operating temperature under various operating conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120916411B_ABST
    Figure CN120916411B_ABST
Patent Text Reader

Abstract

The application discloses a liquid cooling data center heat dissipation intelligent regulation and control method based on dynamic flow regulation, relates to the field of data center thermal management, and comprises the following steps: acquiring CPU power change data and temperature distribution data in real time; acquiring initialization flow parameters of each cooling branch; adapting and adjusting the initialization flow parameters to the temperature distribution data to generate a first group of adaptive hydraulic impedance; performing power jump prediction according to the CPU power change data, and adapting the first group of adaptive hydraulic impedance to the temperature distribution data as a starting point according to the jump prediction information to generate a second group of adaptive hydraulic impedance; and generating a digital control instruction through an impedance modulator, and issuing the digital control instruction to the opening degree of a digital potentiometer control electromagnetic flow valve of each cooling branch through a digital potentiometer control interface. The application solves the technical problem that the existing heat dissipation regulation and control cannot dynamically adjust the cooling flow according to the actual operation state of a server, and achieves the technical effect of adjusting the cooling flow in real time according to the heat dissipation demand.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data center thermal management, and in particular to a method for intelligent control of heat dissipation in liquid-cooled data centers based on dynamic flow regulation. Background Technology

[0002] In the data center field, efficient thermal management is crucial for ensuring stable server operation, extending equipment lifespan, and reducing energy costs. Poor heat dissipation can easily lead to server overheating failures, causing serious consequences such as data loss and business interruption. Currently, the main method for solving data center heat dissipation problems, especially in liquid-cooled data centers, is to adopt a fixed flow rate control strategy. This involves pre-setting the flow rate parameters of each cooling branch of the liquid cooling system and keeping them relatively constant throughout the operation, only requiring manual intervention when significant anomalies occur. However, this fixed flow rate control method fails to fully consider the dynamic changes in CPU power and the real-time differences in internal temperature distribution of data center servers. This results in either excessive cooling flow when CPU power is low, leading to energy waste, or insufficient cooling flow when CPU power surges, failing to remove heat in time and causing localized overheating of the server, affecting server performance and stability.

[0003] Currently, liquid-cooled data center heat dissipation control suffers from the technical problem of being unable to dynamically adjust cooling flow based on the actual operating status of the server. Summary of the Invention

[0004] This application provides an intelligent cooling control method for liquid-cooled data centers based on dynamic flow regulation. It employs real-time acquisition of CPU power changes and temperature distribution data from data center servers, as well as initial flow parameters for each cooling branch of the liquid cooling system. Using temperature distribution data, it adjusts the initial flow parameters to generate a first set of adaptive hydraulic impedances. Based on CPU power change data, it predicts power surges and, combined with temperature distribution data, adapts the first set of adaptive hydraulic impedances to these surges, resulting in a second set of adaptive hydraulic impedances. These second set of adaptive hydraulic impedances are then used to generate digital control commands via an impedance modulator. These commands are then sent to the digital potentiometers of each cooling branch to control the opening of the electromagnetic flow valves. This method solves the technical problem of existing liquid-cooled data center cooling control methods that cannot dynamically adjust cooling flow based on the actual operating status of the server. It achieves the technical effect of adjusting cooling flow in real-time according to the server's cooling needs, improving the intelligence and flexibility of the cooling system.

[0005] This application provides a method for intelligent heat dissipation control of liquid-cooled data centers based on dynamic flow regulation, including: real-time acquisition of CPU power change data and temperature distribution data of data center servers; acquisition of initial flow parameters of each cooling branch in the liquid cooling system of the data center servers; flow resistance state adaptation adjustment of each initial flow parameter based on the temperature distribution data to generate a first set of adapted hydraulic impedance; power jump prediction based on the CPU power change data, and power jump adaptation of the first set of adapted hydraulic impedance based on the power jump prediction information using the temperature distribution data as a starting point to generate a second set of adapted hydraulic impedance; generating digital control commands from the second set of adapted hydraulic impedance through a preset impedance modulator, and sending them to the digital potentiometers of each cooling branch to control the opening of the electromagnetic flow valves through a digital potentiometer control interface.

[0006] In a possible implementation, the following processing is performed: the digital potentiometer control interface is an SPI interface, used to receive the digital control command and configure the resistance value corresponding to the digital potentiometer of each cooling branch; after the digital potentiometer receives the resistance value corresponding to the digital potentiometer of each cooling branch configured by the digital potentiometer control interface, it adjusts the resistance and outputs each voltage signal to the drive circuit of the electromagnetic flow valve of each cooling branch.

[0007] In a possible implementation, the flow resistance state is adapted and adjusted based on the temperature distribution data for each initial flow parameter to generate a first set of adapted hydraulic impedances. The following processes are performed: obtaining the target safe temperature threshold of the data center server; collecting the pipeline layout location of each cooling branch to determine each cooling area; calculating the difference between the temperature value of each cooling area and the target safe temperature threshold based on the temperature distribution data to determine the cooling requirement of each area; constructing a cooling amplitude-hydraulic impedance template; constructing a flow-impedance mapping template, identifying each initial hydraulic impedance corresponding to each initial flow parameter, inputting the cooling amplitude-hydraulic impedance template for actual cooling amplitude matching, and obtaining each matching result; by judging the consistency between each matching result and the cooling requirement of each area, flow resistance state adaptation adjustment is performed to generate the first set of adapted hydraulic impedances.

[0008] In a possible implementation, by determining the consistency between each matching result and the cooling requirements of each region, flow resistance state adaptation adjustment is performed to generate the first set of adapted hydraulic impedances. The following processing is also performed: determining whether the matching result and the cooling requirements of each region meet a preset consistency deviation, obtaining satisfied and unsatisfied regions; extracting the initial hydraulic impedance of the satisfied regions, establishing a mapping with the corresponding cooling region, and adding it to the first set of adapted hydraulic impedances; extracting the regional cooling requirements of the unsatisfied regions, inputting the cooling amplitude-hydraulic impedance template, obtaining the updated hydraulic impedance, establishing a mapping with the corresponding cooling region, and adding it to the first set of adapted hydraulic impedances.

[0009] In a possible implementation, power surge prediction is performed based on the CPU power change data. Starting from the temperature distribution data, the first set of adaptive hydraulic impedances is adapted for power surge based on the power surge prediction information to generate a second set of adaptive hydraulic impedances. The following processing is performed: collecting structural component information in each cooling area corresponding to each cooling branch, performing thermal covariance analysis based on CPU power, and establishing a thermal covariance fitter; reading the operation instructions of the data center server, combining the CPU power change data to perform power surge prediction within a preset period, and generating the power surge prediction information; initializing the thermal covariance fitter starting from the temperature distribution data, loading the power surge prediction information to perform thermal surge fitting, and generating predicted temperature information for each cooling area; updating the cooling requirements of each area based on the predicted temperature information of each cooling area, and by judging the consistency before and after the update, performing flow resistance state adaptation adjustment on the first set of adaptive hydraulic impedances to generate the second set of adaptive hydraulic impedances.

[0010] In a possible implementation, the following process is performed: the preset period is the sum of the control response delay of the liquid cooling system and the time of one liquid cooling cycle.

[0011] In a possible implementation, the operation instructions of the data center server are read, and a power surge prediction within a preset period is performed in conjunction with the CPU power change data to generate the power surge prediction information. The following processing is then performed: historical CPU processing data of the data center server is collected, and operation instructions are classified based on CPU power trends to generate multiple CPU power trends for multiple types of instructions; one type of instruction containing the operation instructions is matched among the multiple types of instructions, and the corresponding CPU power trend is extracted; the CPU power trend is used to predict the CPU power change data within a preset period to generate the power surge prediction information.

[0012] In a possible implementation, the CPU power change data is predicted within a preset period based on the CPU power trend to generate the power jump prediction information, and the following processing is performed: based on the CPU power change data, the corresponding time period trend is extracted from the CPU power trend to correct the power prediction for the known time period and establish a correction coefficient; the power prediction for the unknown time period is performed by combining the correction coefficient and the CPU power trend to generate the power jump prediction information.

[0013] In a possible implementation, the following processing is performed: including safety bypass and degradation control steps: determining the upper-layer network used to collect CPU power change data and temperature distribution data; when a fault is detected in the upper-layer network, triggering a safety bypass mechanism, and controlling the digital potentiometer control interface to switch the resistance values ​​of the digital potentiometers of each cooling branch to the pre-stored conservative resistance.

[0014] In a possible implementation, the following process is performed: after the fault in the upper-layer network is cleared, real-time CPU power and temperature data are reacquired, and the resistance value of the digital potentiometer is smoothly adjusted from the conservative resistance to the resistance value corresponding to the normal control state.

[0015] The proposed intelligent cooling control method for liquid-cooled data centers based on dynamic flow regulation first acquires real-time CPU power variation data and temperature distribution data of the data center server. Then, it acquires the initial flow parameters of each cooling branch within the liquid cooling system of the data center server. Next, it adjusts the flow resistance of each initial flow parameter according to the temperature distribution data to generate a first set of adaptive hydraulic impedance. Then, it predicts power surges based on the CPU power variation data, and adjusts the first set of adaptive hydraulic impedance according to the power surge prediction information based on the temperature distribution data to generate a second set of adaptive hydraulic impedance. Finally, it generates digital control commands from the second set of adaptive hydraulic impedance through a preset impedance modulator, and sends these commands to the digital potentiometers of each cooling branch to control the opening of the electromagnetic flow valves. This achieves the technical effect of adjusting the cooling flow in real-time according to the server's cooling requirements, improving the intelligence and flexibility of the cooling system. Attached Figure Description

[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings of the embodiments of the present invention will be briefly described below. Flowcharts are used in this application to illustrate the operations performed by the system according to the embodiments of the present application. It should be understood that the preceding or following operations are not necessarily performed precisely in sequence. Instead, various steps can be processed in reverse order or simultaneously as needed. Furthermore, other operations can be added to these processes, or one or more steps can be removed from these processes.

[0017] Figure 1 This is a flowchart illustrating the intelligent heat dissipation control method for liquid-cooled data centers based on dynamic flow regulation, as provided in an embodiment of this application.

[0018] Figure 2 This is a schematic diagram of the process for generating the first set of adaptive hydraulic impedance in the intelligent control method for heat dissipation of liquid-cooled data centers based on dynamic flow regulation provided in the embodiments of this application. Detailed Implementation

[0019] The above description is merely an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, specific embodiments of this application are given below.

[0020] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description of this application will be provided in conjunction with the accompanying drawings. The described embodiments should not be considered as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0021] In the following description, references to "some embodiments" describe a subset of all possible embodiments. However, it is understood that "some embodiments" can be the same or different subsets of all possible embodiments and can be combined with each other without conflict. The terms "first" and "second" are used merely to distinguish similar objects and do not represent a specific ordering of objects. The terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or modules not explicitly listed or inherent to these processes, methods, products, or devices. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only.

[0022] This application provides an intelligent heat dissipation control method for liquid-cooled data centers based on dynamic flow regulation, such as... Figure 1 As shown, the method includes:

[0023] Step S100: Real-time acquisition of CPU power change data and temperature distribution data of data center servers.

[0024] Specifically, by deploying power monitoring chips on the CPU of data center servers or using the server's built-in power management functions, real-time power consumption data of the CPU can be collected, i.e., the electrical power consumed by the CPU under different workloads. These chips or functions can measure the real-time power of the CPU under different workloads and transmit the data to the data processing unit via the internal bus.

[0025] Temperature sensors, such as thermocouples and infrared sensors, are installed in critical parts of data center servers, such as CPU heatsinks, memory modules, and hard drives. These sensors measure the temperature of each part in real time and convert the analog temperature signals into digital signals through analog-to-digital converters before transmitting them to the data processing unit.

[0026] Step S200: Obtain the initial flow parameters of each cooling branch in the liquid cooling system deployed in the data center server.

[0027] Specifically, flow meters, such as turbine flow meters and electromagnetic flow meters, are installed on each cooling branch of the liquid cooling system. These flow meters measure the coolant flow rate in real time and transmit the data to the data processing unit via a communication interface. During the system initialization phase, the initial flow rate value of each cooling branch is recorded as a baseline parameter.

[0028] Step S300: Adjust the flow resistance state of each initial flow parameter according to the temperature distribution data to generate the first set of adapted hydraulic resistance.

[0029] Specifically, based on real-time acquired temperature distribution data, the ideal flow rate value for each cooling branch under the current temperature distribution is calculated. Then, based on the difference between the ideal flow rate value and the initial flow rate parameters, the hydraulic resistance of the cooling branch is adjusted, such as by adjusting valve opening or changing pump speed, to make the actual flow rate closer to the ideal flow rate value. The set of hydraulic resistance values ​​for each cooling branch obtained after adjustment based on the initial temperature distribution data constitutes the first set of adaptive hydraulic resistances.

[0030] like Figure 2 As shown, in one possible implementation, the flow resistance state is adjusted based on the temperature distribution data for each initial flow parameter to generate a first set of adapted hydraulic impedance. Step S300 further includes step S310, obtaining the target safe temperature threshold for the data center server. Specifically, based on the data center design specifications, the server equipment manufacturer's recommended values, and historical operating data, a maximum allowable temperature value to ensure stable server operation is comprehensively determined, i.e., the target safe temperature threshold. This value is stored in the data center management system and can be configured and adjusted through the user interface. The target safe temperature threshold serves as the benchmark for temperature control.

[0031] Step S320: Collect the piping locations of each cooling branch to determine each cooling zone. Specifically, collect the location information of each cooling branch using data center design drawings, physical markers, or positioning devices installed on the piping, such as RFID tags. Based on the piping locations and the physical layout of the data center servers, the data center is divided into multiple cooling zones. Each zone contains one or more cooling branches responsible for cooling the servers within that specific zone.

[0032] Step S330: Based on the temperature distribution data, calculate the difference between the temperature value of each cooling zone and the target safe temperature threshold to determine the cooling requirement of each zone. Specifically, extract the average temperature value of each cooling zone from the temperature distribution data obtained in step S100, and then compare it with the target safe temperature threshold to calculate the temperature difference for each zone. Based on the magnitude and direction of the temperature difference, determine the cooling requirement for each cooling zone, i.e., the temperature value that needs to be reduced. A positive difference indicates that cooling is required, while a negative difference or zero indicates that the current temperature is safe and no additional cooling is needed.

[0033] Step S340: Construct a temperature drop amplitude-hydraulic impedance template. Specifically, based on the Darcy-Weisbach equation and heat conduction theory in fluid mechanics, a mathematical model or discretized lookup table is established between the temperature drop amplitude and hydraulic impedance through laboratory testing or CFD simulation. For example, under test conditions with a fixed coolant inlet temperature and server power consumption, the valve opening of the cooling branch is adjusted, i.e., the impedance is changed. The temperature drop values ​​of key parts of the server under different impedance values ​​are recorded, and an empirical formula is fitted. The test data is organized into a tabular form, such as "10Ω impedance corresponds to a 1°C temperature drop, 20Ω impedance corresponds to a 3°C temperature drop", and stored in the template library of the data processing unit for quick lookup.

[0034] Step S350: Construct a flow-impedance mapping template, identify the initial hydraulic impedance corresponding to each initial flow parameter, and input the cooling amplitude-hydraulic impedance template to perform actual cooling amplitude matching, obtaining each matching result. Specifically, based on the physical characteristics of the liquid cooling system and the pump performance curve, a mapping relationship between coolant flow and hydraulic impedance is established through theoretical calculation or experimental calibration. Starting from the initial flow parameters obtained in step S200, the corresponding initial hydraulic impedance is calculated using this mapping template. The initial hydraulic impedance of each cooling zone is input into the cooling amplitude-hydraulic impedance template in step S340. This template performs linear interpolation or nonlinear fitting through a pre-stored "impedance-cooling" relationship to derive the achievable cooling amplitude under the current initial hydraulic impedance.

[0035] Step S360 involves determining the consistency between each matching result and the cooling requirements of each region, and then performing flow resistance state adaptation adjustment to generate the first set of adapted hydraulic impedance. Specifically, the achievable cooling range output in step S350 is compared with the cooling requirements of each region calculated in step S330. That is, it is determined whether the initial hydraulic impedance can meet the cooling requirements. If it cannot meet the cooling requirements, the initial hydraulic impedance is adjusted according to the difference between the matching result and the cooling requirements to obtain the first set of adapted hydraulic impedance.

[0036] In one possible implementation, by judging the consistency between each matching result and the cooling requirements of each region, flow resistance state adaptation adjustment is performed to generate the first set of adapted hydraulic impedance. Step S360 further includes step S361, judging whether the matching result and the cooling requirements of each region meet a preset consistency deviation, to obtain satisfied and unsatisfied regions. Specifically, an allowable deviation range is preset to measure whether the difference between the achievable cooling range (i.e., the matching result) and the actual required cooling range (regional cooling requirement) of each region is within an acceptable range. The achievable cooling range of each cooling region output in step S350 is compared with the actual cooling requirement of the region calculated in step S330. If the difference between the achievable cooling range and the regional cooling requirement is within the preset deviation range, the region is determined to be a satisfied region; if the difference exceeds the preset deviation range, the region is determined to be an unsatisfied region. Through the judgment, it is determined which cooling regions' current hydraulic impedance settings can meet the cooling requirements and which regions need further adjustment.

[0037] Step S362: Extract the initial hydraulic impedance of the satisfied region, establish a mapping with the corresponding cooling region, and add it to the first set of adaptable hydraulic impedances. Specifically, for the cooling regions determined to be satisfied regions in step S361, obtain the initial hydraulic impedances corresponding to these satisfied regions from step S350. Then, establish a one-to-one correspondence between these initial hydraulic impedances and the corresponding cooling regions, and add this mapping relationship to the first set of adaptable hydraulic impedances, that is, retain the hydraulic impedance settings of those cooling regions that can already meet the cooling requirements as part of the final adaptable hydraulic impedances.

[0038] Step S363: Extract the cooling requirements of the unmet areas, input the cooling amplitude-hydraulic impedance template, obtain the updated hydraulic impedance, establish a mapping with the corresponding cooling areas, and add it to the first set of adaptable hydraulic impedances. Specifically, for the cooling areas determined to be unmet in step S361, extract the regional cooling requirement data for these areas. Input the extracted regional cooling requirement data into the cooling amplitude-hydraulic impedance template constructed in step S340. This template derives a new hydraulic impedance value that can meet the cooling requirements of the area, i.e., the updated hydraulic impedance. Then, establish a one-to-one correspondence between these updated hydraulic impedances and the corresponding unmet cooling areas, add this mapping relationship to the first set of adaptable hydraulic impedances, and adjust the hydraulic impedance of the cooling areas that do not meet the cooling requirements so that they can meet the actual cooling requirements.

[0039] Step S400: Based on the CPU power change data, perform power jump prediction; starting from the temperature distribution data, perform power jump adaptation on the first set of adapted hydraulic impedances based on the power jump prediction information, and generate the second set of adapted hydraulic impedances.

[0040] Specifically, historical CPU power variation data is used to predict future CPU power surges through time series analysis and machine learning models. Starting with current temperature distribution data and combining it with power surge prediction information, the ideal flow rate of each cooling branch after the power surge is recalculated. Then, based on the difference between the new ideal flow rate and the first set of adapted hydraulic impedance, the hydraulic impedance of the cooling branches is further adjusted to cope with the changes in heat dissipation demand caused by the increase in CPU power, generating a set of hydraulic impedance values ​​for each cooling branch, i.e., the second set of adapted hydraulic impedance.

[0041] In one possible implementation, power jump prediction is performed based on the CPU power change data. Starting from the temperature distribution data, the first set of adaptive hydraulic impedances is adapted for power jump based on the power jump prediction information to generate a second set of adaptive hydraulic impedances. Step S400 further includes step S410, collecting structural component information in each cooling area corresponding to each cooling branch, performing thermal covariance analysis based on CPU power, and establishing a thermal covariance fitter. Specifically, structural component information in each cooling area corresponding to each cooling branch is collected through data center design documents, equipment lists, and on-site surveys. This information includes, but is not limited to, server type, CPU model, memory capacity, number and type of hard drives, rack material and structure, and the connection method and layout of the cooling branches and these components.

[0042] Based on the collected structural component information, a thermal covariance analysis based on CPU power is performed. During data center operation, when CPU power changes, the CPU area, memory area, rack entrance / exit areas, etc., experience temperature changes first, with the CPU area showing a relatively large temperature rise. By analyzing a large amount of historical data and conducting theoretical modeling, the intrinsic relationship and patterns between CPU power changes and temperature changes in these areas are identified. For example, statistical methods are used to analyze the average and variance of temperature changes in each area under different CPU power variation amplitudes, or a heat conduction model based on physical principles is established to describe this heat transfer and temperature change process. Based on the above analysis results, mathematical methods, such as polynomial fitting and neural network fitting, are used to establish a thermal covariance fitter. This fitter can convert the input CPU power change data into predicted values ​​of temperature changes in various relevant areas.

[0043] Step S420: Read the operation instructions of the data center server, combine them with the CPU power change data to predict power surges within a preset period, and generate the power surge prediction information. The preset period is the sum of the control response delay of the liquid cooling system and the time of one liquid cooling cycle. Specifically, read the operation instructions currently received by the server from the data center management system, including various operations during server operation, such as starting, stopping, and adjusting the load of applications. Different operation instructions directly affect the CPU workload, thus changing the CPU's power consumption. For example, starting a large, computationally intensive application will significantly increase the CPU's computational load, leading to a rapid increase in power; stopping some non-critical applications will reduce the CPU load and correspondingly decrease power.

[0044] Combining the read operation instructions and the real-time CPU power change data acquired in step S100, time series analysis and machine learning prediction algorithms are used to predict CPU power surges within a preset period. The preset period is the sum of the control response delay of the liquid cooling system and the time of one liquid cooling cycle. The control response delay of the liquid cooling system refers to the time required from detecting a change in system state to the control system issuing a corresponding adjustment command and initiating action of the actuator; one liquid cooling cycle refers to the time required for the coolant to complete one complete cycle in the liquid cooling system. This preset period is set to ensure that the predicted CPU power change matches the actual adjustment cycle of the liquid cooling system, enabling the prediction results to effectively guide subsequent heat dissipation adjustment operations. Through the above analysis and calculation, power surge prediction information is generated, which includes the predicted CPU power value at each time point within the preset period, as well as key data such as the power surge magnitude and occurrence time. Power surge prediction is used to plan the liquid cooling system's adjustment strategy in advance, ensuring that data center servers maintain a stable operating temperature even during power changes.

[0045] Step S430: After initializing the thermal covariance fitter using the temperature distribution data as a starting point, the power jump prediction information is loaded to perform thermal jump fitting, generating predicted temperature information for each cooling region. Specifically, using the current temperature distribution data obtained in step S100 as a starting point, the thermal covariance fitter established in step S410 is initialized. The current temperature distribution data is input into the thermal covariance fitter, enabling it to begin subsequent calculations based on the current temperature state. Then, the power jump prediction information generated in step S420 is loaded into the initialized thermal covariance fitter. Based on the CPU power change in the power jump prediction information and the established thermal covariance relationship, the thermal covariance fitter simulates and calculates the temperature changes of each cooling region within a preset future period. Through thermal jump fitting, predicted temperature information for each cooling region within a preset period is generated. This predicted temperature information reflects the possible future temperature levels of each cooling region considering CPU power jumps.

[0046] Step S440: Update the cooling requirements of each cooling zone based on the predicted temperature information of each cooling zone. By judging the consistency before and after the update, perform flow resistance state adaptation adjustment on the first set of adapted hydraulic impedances to generate the second set of adapted hydraulic impedances. Specifically, based on the predicted temperature information of each cooling zone generated in step S430, recalculate the cooling requirements of each zone. This is similar to the method of calculating the cooling requirements based on the current temperature distribution data in step S330. Compare the average temperature value of each cooling zone in the predicted temperature information with the target safe temperature threshold to calculate the predicted temperature difference of each zone within a preset period. Based on the magnitude and direction of the temperature difference, determine the updated cooling requirements of each cooling zone under the power jump condition, i.e., the temperature value that needs to be reduced (a positive difference indicates that cooling is required, a negative difference or zero indicates that the current predicted temperature is safe and no additional cooling is required). Next, judge the consistency between the updated cooling requirements of each zone and the original cooling requirements of the zone calculated based on the current temperature distribution data in step S330. By comparing the differences in cooling requirements calculated in the two steps, assess the degree of impact of the power jump on the cooling requirements of each cooling zone. Based on the consistency judgment results, the first set of adapted hydraulic impedance generated in step S360 is adjusted for flow resistance state adaptation. If the updated cooling demand of a certain cooling area differs significantly from the original cooling demand, it indicates that the current hydraulic impedance setting cannot meet the cooling demand after the power jump. Therefore, the hydraulic impedance corresponding to that area needs to be adjusted according to the new cooling demand, referring to the methods in steps S340-S360, so that the actual flow rate can better adapt to the heat changes brought about by the power jump. After the above adaptation adjustment, a second set of adapted hydraulic impedance is generated. This set of hydraulic impedance is obtained by optimizing and adjusting the first set of adapted hydraulic impedance based on the CPU power jump prediction information. It can better cope with the heat dissipation challenges brought about by changes in data center server power, ensuring that the data center can maintain a stable operating temperature under various operating conditions.

[0047] In one possible implementation, the operation instructions of the data center server are read, and power surge prediction within a preset period is performed in conjunction with the CPU power change data to generate the power surge prediction information. Step S420 further includes step S421, which involves collecting historical CPU processing data of the data center server, performing operation instruction classification based on CPU power trends, and generating multiple CPU power trends for multiple instruction types. Specifically, historical CPU processing data of the data center server is collected through channels such as data center management system logs and server monitoring records. This historical data contains detailed information about the server at various points in time during long-term operation, such as the real-time power value of the CPU, the type of task processed by the CPU, and the task execution time.

[0048] Based on the collected historical CPU processing data, we performed operation instruction classification based on CPU power trends. We analyzed the characteristics of CPU power changes during the execution of different operation instructions, such as application startup, file transfer, and database query. For example, when starting a large graphics processing application, CPU power quickly rises to a high level and remains relatively stable; when performing a simple file copy operation, CPU power only rises briefly and slightly.

[0049] Clustering algorithms from data mining are used to classify operation instructions. By analyzing the CPU power change patterns corresponding to different operation instructions, instructions with similar power change trends are grouped together. Simultaneously, the CPU power trend corresponding to each category of operation instructions is recorded, including information such as the starting point of power increase, the rate of increase, the peak value reached, the power value after stabilization, and the pattern of power decrease. Through classification and recording, multiple CPU power trends for multiple instruction categories are generated.

[0050] Step S422: Match a category containing the operation instruction among the multiple instruction categories and extract the corresponding CPU power trend. Specifically, when the operation instruction of the current data center server is read, a matching search is performed among the multiple instruction categories generated in step S421. By comparing the characteristics of the current operation instruction, such as instruction type and the type of task involved, with the characteristics of various instruction categories, the category to which the current operation instruction belongs is determined. Once a category of instructions containing the current operation instruction is matched, the CPU power trend most relevant to the current operation instruction is extracted from the multiple CPU power trends corresponding to that category of instructions. The extraction process can be selected based on the specific details of the operation instruction. For example, if the current operation instruction is to launch a specific version of graphics processing software, then the CPU power trend corresponding to the launch process of that software version is further filtered from the instruction categories related to graphics processing software.

[0051] Step S423: Based on the CPU power trend, predict the CPU power change data within a preset period to generate the power jump prediction information. Specifically, based on the CPU power trend extracted in step S422, and combined with the CPU power change data acquired in real time in step S100, predict the CPU power within a preset period. First, analyze the similarity between the current CPU power change data and the initial part of the extracted CPU power trend. If the current power change trend matches the initial stage of a historical trend, it can be inferred that subsequent power changes will follow that historical trend. Then, based on the characteristics such as the speed and magnitude of power increase in the historical trend, and combined with the actual situation of the current power change, quantitatively predict the CPU power within the preset period. Through the above prediction process, power jump prediction information is generated.

[0052] In one possible implementation, the CPU power change data is predicted within a preset period based on the CPU power trend to generate the power jump prediction information. Step S423 further includes step S4231, which involves extracting the corresponding time period trend from the CPU power trend based on the CPU power change data, correcting the power prediction for the known time period, and establishing a correction coefficient. Specifically, the corresponding time period in the CPU power trend is determined at the current moment. For example, if the current moment is in the stage where CPU power begins to rise, the corresponding historical rising stage is found in the CPU power trend. The differences between the current CPU power change data and the corresponding time period data in the CPU power trend are analyzed. For example, the current moment when the CPU power begins to rise is earlier or later than the starting moment in the historical trend, or the current rate of power rise is inconsistent with the rate in the historical trend. Based on these differences, a correction coefficient is established. The correction coefficient can be calculated using various methods, such as linear regression analysis and proportional calculation. Taking linear regression analysis as an example, the current CPU power change data is used as the dependent variable, and the data of the corresponding time period in the CPU power trend is used as the independent variable. The relationship between the two is obtained through regression analysis, and then the correction coefficient is determined. The correction factor reflects the degree of deviation between the current actual power change and the historical trend. It is used to correct power predictions based on historical trends, so that the prediction results are more in line with the actual situation.

[0053] Step S4232: Combine the correction coefficient and the CPU power trend to predict the power increase for the unknown period, generating the power surge prediction information. Specifically, determine the corresponding position of the unknown period in the CPU power trend. Since the CPU power trend is based on historical data and includes the complete power change process, the unknown period is the future period to be predicted. Based on the duration of a preset period, find the corresponding period range in the CPU power trend. Using the CPU power trend extracted in step S422, make a preliminary prediction of the power for the corresponding period, that is, based on the patterns in historical data, assume that the future will follow a similar historical trend. For example, if the historical trend shows that the CPU power will increase at a certain rate in a certain stage, then predict the power for the unknown period according to this rate. Finally, apply the correction coefficient established in step S4231 to the preliminary prediction result, correcting the preliminary prediction result by multiplying the correction coefficient by the preliminary predicted power value or using other suitable calculation methods. The prediction result obtained in this way considers both the patterns of historical power trends and the differences between the current actual power change and the historical trend, and can more accurately reflect the actual changes in CPU power within the unknown period. Through the above process, power jump prediction information is generated.

[0054] In step S500, the second set of adapted hydraulic impedance is used to generate digital control commands through a preset impedance modulator, and then sent to the digital potentiometers of each cooling branch to control the opening of the electromagnetic flow valve through the digital potentiometer control interface.

[0055] Specifically, an impedance modulator is a hardware or software module that converts hydraulic impedance values ​​into digital control commands. The impedance modulator converts a second set of adapted hydraulic impedance values ​​into digital control commands, including specific control parameters such as voltage and current values, according to a preset mapping relationship. The digital potentiometer control interface is the communication interface connecting the impedance modulator and the digital potentiometer, used to transmit digital control commands. The digital potentiometer adjusts its resistance value according to the received digital control commands, thereby controlling the opening of the electromagnetic flow valve and achieving precise regulation of the coolant flow rate.

[0056] In one possible implementation, step S500 further includes: the digital potentiometer control interface is an SPI interface, used to receive the digital control command and configure the resistance value corresponding to the digital potentiometer of each cooling branch; the digital potentiometer receives the resistance value corresponding to the digital potentiometer of each cooling branch configured by the digital potentiometer control interface, adjusts the resistance, and outputs each voltage signal to the drive circuit of the electromagnetic flow valve of each cooling branch.

[0057] Specifically, the digital potentiometer control interface is an SPI interface, a high-speed, full-duplex, synchronous communication bus used here to receive digital control commands generated by the impedance modulator. This interface can accurately and efficiently transmit digital control commands from the impedance modulator to the digital potentiometer, ensuring the timeliness and accuracy of command transmission. Simultaneously, the SPI interface also has the function of configuring the resistance values ​​of the digital potentiometers corresponding to each cooling branch, allowing the setting of the resistance value of the digital potentiometer corresponding to each cooling branch based on relevant information in the digital control commands. Specifically, after receiving the resistance values ​​of the digital potentiometers corresponding to each cooling branch configured by the digital potentiometer control interface, the digital potentiometer begins resistance adjustment. The digital potentiometer has a corresponding circuit structure and control logic internally, capable of adjusting its own resistance based on the received resistance value parameters. Through this resistance adjustment, the digital potentiometer controls the output voltage signal. The digital potentiometer outputs various voltage signals to the drive circuits of the electromagnetic flow valves in each cooling branch. These voltage signals are parameters that drive the electromagnetic flow valves; different voltage values ​​correspond to different opening states of the electromagnetic flow valves. After receiving these voltage signals, the drive circuit converts electrical energy into mechanical energy according to the magnitude and characteristics of the voltage signals, driving the valve core of the electromagnetic flow valve to move, thereby controlling the opening of the electromagnetic flow valve and regulating the coolant flow to meet the heat dissipation needs of data center servers.

[0058] In one possible implementation, the method further includes a safety bypass and degradation control step: determining an upper-layer network for collecting CPU power change data and temperature distribution data; when a fault is detected in the upper-layer network, triggering a safety bypass mechanism, and controlling the digital potentiometer control interface to switch the resistance values ​​of the digital potentiometers of each cooling branch to a pre-stored conservative resistance.

[0059] Specifically, a health check is performed on the upper-layer network used to collect CPU power change data and temperature distribution data, including checking various indicators such as network connection stability, data transmission accuracy, and communication protocol compatibility. Through real-time monitoring and analysis of these indicators, potential problems in the upper-layer network can be identified promptly. When a fault is detected in the upper-layer network, a safety bypass mechanism is immediately activated. This mechanism is triggered based on pre-set fault judgment thresholds and conditions, such as the number of data transmission interruptions exceeding a set value or the data error rate reaching a certain proportion. After the safety bypass mechanism is triggered, the control interface of the digital potentiometer reads pre-stored conservative resistance values. These conservative resistance values ​​are calculated and experimentally verified to ensure that, in the event of a fault, the digital potentiometers of each cooling branch can switch to the safest and most stable resistance state. The control interface sends a switching command to the digital potentiometers of each cooling branch, switching the resistance value of the digital potentiometer to the pre-stored conservative resistance. Through safety bypass and degradation control steps, when the upper-layer network fails, the digital potentiometer of the cooling branch can be quickly and reliably switched to a conservative resistance state, thereby ensuring the stability and safety of the entire system and preventing the cooling system from going out of control due to network failure, which could then damage critical components such as the CPU.

[0060] In one possible implementation, the method further includes: after the fault in the upper-layer network is cleared, reacquiring real-time CPU power and temperature data, and controlling the resistance value of the digital potentiometer to be smoothly adjusted from the conservative resistance to the resistance value corresponding to the normal control state.

[0061] Specifically, when a fault occurs in the upper-layer network, continuous real-time monitoring of the upper-layer network is performed. This is achieved by periodically sending probe signals and checking data transmission channels to determine if the fault has been resolved. Once the fault in the upper-layer network is detected as resolved, the recovery and control process is immediately initiated. The first step is to reacquire real-time CPU power and temperature data. Then, using a method similar to steps S100-S500, the resistance values ​​of the digital potentiometers for each cooling branch under normal control are calculated. The resistance values ​​of the digital potentiometers are then smoothly adjusted from a conservative resistance to the resistance values ​​corresponding to the normal control state. The smooth adjustment process employs a gradual adjustment strategy to avoid impacting the cooling system due to sudden changes in resistance values. The system gradually adjusts the resistance values ​​of the digital potentiometers according to preset adjustment step sizes and time intervals. Through the recovery and control steps after the fault is resolved, the normal control state of the cooling system can be quickly and accurately restored after the upper-layer network fault is resolved, ensuring stable CPU operation under suitable temperature conditions while avoiding adverse effects on the system due to improper resistance value adjustment.

[0062] This application employs real-time acquisition of CPU power changes and temperature distribution data of data center servers, as well as initial flow parameters of each cooling branch of the liquid cooling system. It uses temperature distribution data to adjust the initial flow parameters to generate a first set of adaptive hydraulic impedance. Based on CPU power change data, it performs power surge prediction and, combined with temperature distribution data, adapts the first set of adaptive hydraulic impedance to the power surge, resulting in a second set of adaptive hydraulic impedance. The second set of adaptive hydraulic impedance is then used to generate digital control commands via an impedance modulator. These commands are then sent to the digital potentiometers of each cooling branch through a digital potentiometer control interface to control the opening of the electromagnetic flow valves. This approach solves the technical problem of existing liquid-cooled data center heat dissipation control systems being unable to dynamically adjust cooling flow according to the actual operating status of the server. It achieves the technical effect of adjusting cooling flow in real-time according to the server's heat dissipation needs, improving the intelligence and flexibility of the heat dissipation system.

[0063] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application. In some cases, the actions or steps described in this application can be performed in a different order than that shown in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

Claims

1. A method for intelligent heat dissipation control of liquid-cooled data centers based on dynamic flow regulation, characterized in that, include: Real-time acquisition of CPU power change data and temperature distribution data of data center servers; Obtain the initial flow parameters of each cooling branch in the liquid cooling system deployed in the data center server; The temperature distribution data is used to adjust the flow resistance state of each initial flow parameter to generate a first set of adapted hydraulic resistance. Based on the CPU power change data, a power jump prediction is performed. Starting from the temperature distribution data, the first set of adaptive hydraulic impedances is adapted for power jump based on the power jump prediction information to generate a second set of adaptive hydraulic impedances. The second set of adapted hydraulic impedance is used to generate digital control commands through a preset impedance modulator, and then sent to the digital potentiometers of each cooling branch through the digital potentiometer control interface to control the opening of the electromagnetic flow valve. The flow resistance state is adjusted based on the temperature distribution data to generate a first set of adapted hydraulic resistances, including: Obtain the target safe temperature threshold of the data center server; Collect the piping locations of each cooling branch to determine each cooling zone; Based on the temperature distribution data, the difference between the temperature value of each cooling zone and the target safe temperature threshold is calculated to determine the cooling requirements of each zone. Construct a template for temperature drop amplitude-hydraulic resistance; Construct a flow-impedance mapping template, identify the initial hydraulic impedance corresponding to each initial flow parameter, input the cooling range-hydraulic impedance template to perform actual cooling range matching, and obtain each matching result; By judging the consistency between each matching result and the cooling requirements of each region, the flow resistance state adaptation adjustment is performed to generate the first set of adapted hydraulic resistance. By determining the consistency between each matching result and the cooling requirements of each region, flow resistance state adaptation adjustment is performed to generate the first set of adapted hydraulic resistances, which also includes: Determine whether the matching result deviates from the preset consistency deviation of the cooling requirements of each region, and obtain the regions that meet the requirements and the regions that do not meet the requirements. Extract the initial hydraulic impedance of the satisfied region, establish a mapping with the corresponding cooling region, and add it to the first set of adaptive hydraulic impedance; Extract the cooling requirements of the unmet areas, input the cooling range-hydraulic impedance template, obtain the updated hydraulic impedance, establish a mapping with the corresponding cooling areas, and add it to the first set of adaptive hydraulic impedance. Based on the CPU power change data, a power jump prediction is performed. Starting from the temperature distribution data, the first set of adapted hydraulic impedances is adapted for power jump based on the power jump prediction information to generate a second set of adapted hydraulic impedances, including: Collect structural component information for each cooling area corresponding to each cooling branch, perform thermal covariance analysis based on CPU power, and establish a thermal covariance fitter; Read the operation instructions of the data center server, combine them with the CPU power change data to predict the power jump within a preset period, and generate the power jump prediction information; After initializing the thermal covariance fitter using the temperature distribution data as a starting point, the power jump prediction information is loaded to perform thermal jump fitting, generating predicted temperature information for each cooling region. The cooling requirements of each cooling zone are updated based on the predicted temperature information of each zone. By judging the consistency before and after the update, the flow resistance state is adjusted for the first set of adapted hydraulic impedances to generate the second set of adapted hydraulic impedances.

2. The intelligent heat dissipation control method for liquid-cooled data centers based on dynamic flow regulation as described in claim 1, characterized in that, The digital potentiometer control interface is an SPI interface, used to receive the digital control commands and configure the resistance values ​​of the digital potentiometers for each cooling branch. The digital potentiometer receives the resistance values ​​corresponding to the digital potentiometers of each cooling branch configured by the digital potentiometer control interface, adjusts the resistance, and then outputs each voltage signal to the drive circuit of the electromagnetic flow valve of each cooling branch.

3. The intelligent heat dissipation control method for liquid-cooled data centers based on dynamic flow regulation as described in claim 1, characterized in that, The preset period is the sum of the control response delay of the liquid cooling system and the time of one liquid cooling cycle.

4. The intelligent heat dissipation control method for liquid-cooled data centers based on dynamic flow regulation as described in claim 1, characterized in that, Read the operation instructions of the data center server, combine them with the CPU power change data to predict power surges within a preset period, and generate the power surge prediction information, including: Collect historical CPU processing data of the data center server, execute operation instruction classification based on CPU power trend, and generate multiple CPU power trends for multiple instruction types; Among the multiple types of instructions, match the type that contains the operation instruction, and extract the corresponding CPU power trend; Based on the CPU power trend, the CPU power change data is predicted within a preset period to generate the power jump prediction information.

5. The intelligent heat dissipation control method for liquid-cooled data centers based on dynamic flow regulation as described in claim 4, characterized in that, Based on the CPU power trend, the CPU power change data is predicted within a preset period to generate the power jump prediction information, including: Based on the CPU power change data, the corresponding time period trend is extracted from the CPU power trend, and the power prediction for the known time period is corrected to establish a correction coefficient; By combining the correction coefficient and the CPU power trend, power prediction is performed for unknown periods to generate the power surge prediction information.

6. The intelligent heat dissipation control method for liquid-cooled data centers based on dynamic flow regulation as described in claim 1, characterized in that, It also includes safety bypass and degradation control steps: Determine the upper-layer network used to collect CPU power variation data and temperature distribution data; When a fault is detected in the upper-layer network, a safety bypass mechanism is triggered, which controls the digital potentiometer control interface to switch the resistance value of the digital potentiometers in each cooling branch to a pre-stored conservative resistance.

7. The intelligent heat dissipation control method for liquid-cooled data centers based on dynamic flow regulation as described in claim 6, characterized in that, Once the fault in the upper-layer network is resolved, real-time CPU power and temperature data are reacquired, and the resistance value of the digital potentiometer is smoothly adjusted from the conservative resistance to the resistance value corresponding to the normal control state.

Citation Information

Patent Citations

  • Control method and system of data center liquid cooling heat dissipation system

    CN118338630A