Control method for heat dissipation of high-performance computing cluster platform

By building a three-dimensional thermal field model and hybrid prediction model of the high-performance computing cluster platform, the thermal risk score index is analyzed and corrected, the cluster of hot spot areas is divided and multi-level cooling response is triggered, and the problem of slow heat dissipation response of the computing cluster platform is solved, and high-precision, low energy consumption and high reliability heat dissipation control is achieved.

CN120179044AActive Publication Date: 2025-06-20TIANJIN ZHONGDA ZHITENG TECH CO LTD

Patent Information

Application Number
CN202510629070.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-06-20
Estimated Expiration
2045-05-16

AI Technical Summary

Technical Problem

The passive heat dissipation of the high-performance computing cluster platform responds slowly to burst loads and is difficult to take into account timeliness, energy efficiency and reliability, resulting in a sharp surge in chip temperature, triggering hardware protection frequency reduction and causing computing power loss.

Method used

Build a three-dimensional thermal field model of the high-performance computing cluster platform, obtain the sensing data of each node, import the mixed prediction model of LSTM and GBRT, obtain the average deviation value of the prediction data of each node, analyze and correct the thermal risk score index, divide the clusters of hot spot areas, assign calculation tasks, trigger multi-level cooling response, and perform early warning and hardware compensation through the cooling effect response value.

Benefits of technology

It realizes high-precision, low energy consumption and high reliability heat dissipation control, quickly matches burst computing tasks, avoids overheating and runs out of control, and ensures system thermal stability and computing power stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179044A_ABST
    Figure CN120179044A_ABST
Patent Text Reader

Abstract

The invention discloses a control method for heat dissipation of a high-performance computing cluster platform, which comprises the following steps of: firstly, establishing a three-dimensional temperature distribution model of a high-performance computing cluster, acquiring data such as temperature and power consumption of each node in real time, obtaining deviation data of each node by utilizing a prediction model, and dynamically evaluating a heat dissipation risk level of each node; dividing the whole cluster into different risk areas; computing tasks are intelligently distributed based on risk levels, and high-load tasks are preferentially scheduled to a low-temperature area; and then, starting graded cooling measures for the high-risk area, continuously monitoring the cooling effect, and feeding back the result to an early warning system and a hardware maintenance module to form'prediction-regulation-feedback 'closed-loop control. The heat dissipation efficiency of the high-performance computing cluster platform is remarkably improved, the overheat fault risk is reduced while the computing performance is guaranteed, the service life of key hardware is prolonged through dynamic optimization, and intelligent cooperation of heat dissipation resources and computing tasks is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The field of data processing technology The present invention relates to the field of data processing technology, and particularly to a control method for heat dissipation of a high-performance computing cluster platform. Background Art

[0002] During the operation of a high-performance computing cluster platform, a large amount of heat is generated. If the heat dissipation is not timely or uneven, it may lead to a decline in system performance, hardware damage, or even service interruption. Moreover, traditional cooling methods such as air cooling are difficult to meet the requirements of high-density computing tasks, and new cooling technologies such as liquid cooling and thermoelectric cooling are gradually applied to high-performance computing clusters to improve the heat dissipation efficiency and system energy efficiency ratio. Therefore, constructing an efficient heat dissipation control system is crucial for ensuring the stable operation of the computing cluster.

[0003] For example, a control method and system for heat dissipation of a high-performance computing cluster platform disclosed in the patent with publication number CN113721741B includes obtaining the process of job operations of a high-performance computing job scheduling system; adjusting the power of the heat dissipation devices of the server computing cards, server chassis, and / or server cabinets where the job operations are executed according to the process of job operations; the passive adjustment method includes collecting temperature data and calculating temperature warning data; adjusting the power of the heat dissipation devices of the server computing cards, server chassis, and / or server cabinets where temperature warnings occur according to the temperature warning data.

[0004] For example, a heat dissipation device and its control system for a high-performance computing cluster platform disclosed in the patent with publication number CN117032425A includes a host, a heat dissipation plate, a heat dissipation and fire extinguishing component, and a safety component. A power cord is plugged into the back of the host, the heat dissipation plate is fixedly covered on the top of the host, the heat dissipation and fire extinguishing component is fixedly arranged between the bottom of the heat dissipation plate and the back corner of the host, and a connecting plate is fixedly connected to the bottom of the heat dissipation and fire extinguishing component. A safety component for assisting in pulling out the power cord is fixedly arranged on the inner side of the connecting plate.

[0005] However, in the process of implementing the technical solutions of the present invention in the embodiments of the present application, it is found that the above technologies have at least the following technical problems: There are problems such as the slow response of passive heat dissipation of the computing cluster platform to sudden loads and the difficulty of the heat dissipation solution to balance timeliness, energy efficiency, and reliability, which may lead to the inability to quickly match the sudden jump in transient heat flux density generated by sudden computing tasks, causing the chip temperature to soar rapidly and triggering hardware-protective downclocking, resulting in computing power loss. Summary of the Invention

[0006] Aiming at the deficiencies of the prior art, the present invention provides a control method for heat dissipation of a high-performance computing cluster platform to solve the problems described in the above background art.

[0007] To achieve the above object, the present invention is realized by the following technical solutions: A control method for heat dissipation of a high-performance computing cluster platform, including constructing a three-dimensional thermal field model of the high-performance computing cluster platform, obtaining sensing data of each node of the computing cluster platform, importing it into a hybrid prediction model of LSTM and GBRT, and obtaining the average deviation value of the prediction data of each node.

[0008] Analyze and correct the prediction data of each node according to the average deviation value of the prediction data of each node, and combine the sensing data of each node to obtain the heat risk scoring index of each node.

[0009] Divide the hot spot areas of the high-performance computing cluster platform according to the heat risk scoring index of each node to obtain each risk area cluster, and allocate computing tasks according to each risk area cluster.

[0010] Monitor each node in each risk area cluster, obtain and analyze the hot spot data of each node in each risk area cluster, trigger a multi-level cooling response for each risk area cluster in each risk area cluster and monitor it, obtain the cooling effect response value of each risk area cluster, and give an early warning to the high-performance computing cluster platform according to the cooling effect response value of each risk area cluster and trigger hardware compensation.

[0011] Further, to obtain the average deviation value of the prediction data of each node, the specific process is as follows: Preset each monitoring period, and extract the prediction data of each node in each monitoring period, including temperature prediction value, heat flux density prediction value, and temperature rise prediction rate.

[0012] Obtain the actual data of each current node in each monitoring period, including average temperature, heat flux density, and temperature rise rate, compare them with the prediction data of each previous node respectively, and perform coupling after introducing a weight coefficient to obtain the average deviation value of the prediction data of each node. The average deviation value of the prediction data of each node is used to evaluate the accuracy of the prediction model.

[0013] Further, to analyze and correct the prediction data of each node, the specific process is as follows: Count the number of monitoring periods in which the prediction data deviation value of each node is higher than or equal to the preset average deviation threshold. If the number of monitoring periods in which the prediction data deviation value of each node is higher than or equal to the preset average deviation threshold is higher than the set deviation period data quantity threshold, then analyze and correct the prediction data of each node.

[0014] Further, the thermal risk scoring index of each node is obtained. The specific process is as follows: Extract the sensing data of each node in the computing cluster platform, including the total power consumption, the average fan speed, and the coolant flow rate, and compare them with their corresponding reference values respectively. After introducing the weight coefficient and the average deviation value of the prediction data of each node, perform coupling to obtain the thermal risk scoring index of each node. The thermal risk scoring index of each node is used to evaluate the thermal risk level of each node in the high-performance computing cluster platform.

[0015] Further, divide the hot spots in the high-performance computing cluster platform according to the thermal risk scoring index of each node to obtain each risk area cluster. The specific process is as follows: Extract the thermal risk scoring index of each node, and compare it with each risk area cluster corresponding to each interval of the set thermal risk scoring index of each node to obtain each risk area cluster. Each risk area cluster includes each safety area cluster, each warning area cluster, and each hot spot area cluster.

[0016] Further, allocate computing tasks according to each risk area cluster. The specific analysis process is as follows: After the scheduler sorts each risk area cluster according to the cluster-level risk, preferentially allocate compute-intensive or long-term batch processing tasks to the nodes in the safety area cluster, and issue an instruction to suspend the new task scheduling instructions for each warning area cluster and hot spot area cluster.

[0017] Further, trigger the multi-level cooling response of each risk area cluster and monitor it. The specific analysis process is as follows: Perform a first-level cooling response to the computing cluster platform in each safety area cluster, perform a second-level cooling response to the computing cluster platform in each warning area cluster and each hot spot area cluster, and continuously monitor the hot spot data of each hot spot area cluster.

[0018] Further, obtain the cooling effect response value of each risk area cluster. The specific process is as follows: Preset a monitoring time period, and extract the hot spot data of each node in each hot spot area cluster during the monitoring time period, including the maximum value of the temperature rise rate and the defined maximum value of the temperature rise rate, the cumulative number of thermal drop frequency events and the defined cumulative number of thermal drop frequency events, and the average task completion rate and the reference average task completion rate are respectively analyzed for proportion, and after introducing the weight coefficient and the thermal risk scoring index of each node, perform coupling to obtain the cooling effect response value of each risk area cluster. The cooling effect response value of each risk area cluster is used to evaluate the overall cooling effect degree of the cooling measures on the hot spot cluster.

[0019] Further, give an early warning to the high-performance computing cluster platform according to the cooling effect response value of each risk area cluster. The specific process is as follows: According to the cooling effect response value of each risk area cluster, and compare it with the set cooling effect response threshold of each risk area cluster. If the cooling effect response value of each risk area cluster is lower than or equal to the cooling effect response threshold of each risk area cluster, continuously monitor each risk area cluster. If the cooling effect response value of each risk area cluster is higher than the cooling effect response threshold of each risk area cluster, give an early warning.

[0020] Further, trigger hardware compensation. The specific process is as follows: Count the number of risk area clusters whose cooling effect response values are higher than the cooling effect response thresholds of their respective risk area clusters. If the number of risk area clusters whose cooling effect response values are higher than the cooling effect response thresholds of their respective risk area clusters is greater than or equal to the risk area cluster number threshold, trigger hardware compensation and enable immersion phase change cooling.

[0021] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: The present invention has the following beneficial effects: (1) A control method for heat dissipation of a high-performance computing cluster platform provided by the present invention first constructs a three-dimensional thermal field model of the cluster. Then, according to the distribution quantiles of all nodes' TRS, risk clusters are divided, and compute-intensive or long-duration batch tasks are preferentially deployed. On this basis, a multi-level cooling response strategy is adopted, and the nodes within the cluster are continuously monitored. Finally, the monitoring of the cooling effect response value triggers hardware compensation measures to prevent overheating and loss of control, and continuous closed-loop monitoring and feedback are carried out, and through the monitoring and response data feedback, to achieve high-precision, low-energy consumption, and high-reliability closed-loop control of the heat dissipation of the high-performance computing cluster.

[0022] (2) The present invention effectively solves the problem of prediction distortion caused by load fluctuations in the heat dissipation control of a high-performance computing cluster through the average prediction deviation value at the node level. Subsequently, the online analysis and correction process is triggered by counting the errors of each node, thereby ensuring that subsequent heat risk scoring and task scheduling based on the prediction results are always based on the output of a highly reliable model, which helps to immediately adjust the cooling strategy and maintain the thermal stability of the system.

[0023] (3) The present invention helps subsequent cluster-level clustering by obtaining the heat risk scoring index of each node, and evaluates the overall cooling effect by monitoring the energy efficiency index, accurately evaluates its heat dissipation pressure, and preferentially assigns high-load tasks to the nodes in the low-temperature safe area through the scheduling system, while restricting the computing power allocation in the high-temperature area. This strategy forms a closed-loop control of "monitoring - evaluation - scheduling", effectively achieving precise placement of heat dissipation resources by dynamically dividing risk areas and avoiding global overheating; intelligently scheduling tasks in combination with the heat risk state not only ensures computing efficiency but also alleviates local heat dissipation pressure, and anticipates potential failures in advance through heat state prediction, improving the reliability of the hardware. The entire process realizes the collaborative management of heat dissipation control and computing load.

[0024] (4) By obtaining the heat dissipation effect score values of each region, the present invention evaluates the actual effect of the heat dissipation measures through multi-dimensional index fusion, avoids misjudgment of a single temperature parameter, establishes a dynamic early warning mechanism, can intervene in advance when the heat dissipation efficiency is insufficient, prevent chain failures caused by local overheating, and correlates the heat dissipation efficiency with the task execution efficiency. While ensuring the computing performance, precise control of the heat dissipation resources is achieved. The whole set of methods improves the response speed and reliability of the heat dissipation system, and forms a dynamic balance between the heat dissipation efficiency and the computing load. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 is a schematic diagram of the method of the present invention; Figure 2 is a schematic diagram of the logical flow of the control method for heat dissipation of a high-performance computing cluster platform of the present invention; Figure 3 is a schematic diagram of the logical flow of analyzing and correcting the prediction data of each node of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0027] In the description of the present invention, it should be understood that the terms "opening", "upper", "lower", "thickness", "top", "middle", "length", "inner", "perimeter", etc. indicating the orientation or position relationship are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the components or elements referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation of the present invention.

[0028] Please refer to Figure 1 , the embodiments of the present invention provide a technical solution: a control method for heat dissipation of a high-performance computing cluster platform, including constructing a three-dimensional thermal field model of the high-performance computing cluster platform, obtaining the sensing data of each node of the computing cluster platform, and importing it into the LSTM and GBRT hybrid prediction model to obtain the average deviation value of the prediction data of each node.

[0029] Analyze and correct the prediction data of each node according to the average deviation value of the prediction data of each node, and combine the sensing data of each node to obtain the heat risk scoring index of each node.

[0030] Divide the hot spots of the high-performance computing cluster platform according to the thermal risk scoring index of each node to obtain each risk area cluster, and allocate computing tasks according to each risk area cluster.

[0031] Monitor each node in each risk area cluster, obtain and analyze the hot spot data of each node in each risk area cluster, trigger multi-level cooling responses for each risk area cluster in each risk area cluster and monitor them to obtain the cooling effect response values of each risk area cluster, and issue an early warning for the high-performance computing cluster platform and trigger hardware compensation according to the cooling effect response values of each risk area cluster.

[0032] Specifically, obtain the average deviation value of the prediction data of each node. The specific process is as follows: preset each monitoring period, and extract the prediction data of each node in each monitoring period, including the temperature prediction value, the heat flux density prediction value, and the temperature rise prediction rate.

[0033] Obtain the actual data of each current node in each monitoring period, including the average temperature, the heat flux density, and the temperature rise rate, compare them with the prediction data of each previous node respectively, and perform coupling after introducing the weight coefficient to obtain the average deviation value of the prediction data of each node. The average deviation value of the prediction data of each node is used to evaluate the accuracy of the prediction model.

[0034] ; represents the average deviation value of the prediction data of the i-th node, represents the average temperature of the i-th node, represents the predicted average temperature of the i-th node in the previous time, represents the set defined reference average temperature deviation amount, represents the heat flux density of the i-th node, represents the predicted heat flux density of the i-th node in the previous time, represents the set defined heat flux density deviation amount, represents the temperature rise rate of the i-th node, represents the predicted temperature rise rate of the i-th node in the previous time, represents the set defined temperature rise rate deviation amount, represents the weight coefficient corresponding to the set average temperature, represents the weight coefficient corresponding to the set heat flux density, represents the weight coefficient corresponding to the set temperature rise rate, i represents the number of each node, , n is the total number of nodes.

[0035] It should be noted that in the heat dissipation control of a high-performance computing cluster, the average temperature, heat flux density, and temperature rise rate are closely coupled: when the heat flux density exceeds the instantaneous processing capacity of the heat dissipation system, the temperature rise rate will rapidly climb. At the same time, the continuous high temperature rise rate will act on the average temperature, causing it to further exceed the safety threshold and exacerbating the heat flux density demand. If not intervened in time, it may lead to hotspot instability or device thermal throttling, triggering performance fluctuations and reliability risks. And when the heat flux density increases, if the heat dissipation measures are insufficient, the average temperature will rise, resulting in an accelerated temperature rise rate. Conversely, if the average temperature continues to rise, it may indicate that the heat flux density is too high or the heat dissipation efficiency has decreased, thereby triggering an increase in the temperature rise rate.

[0036] In a specific embodiment, the value ranges of the weight coefficient corresponding to the average temperature, the weight coefficient corresponding to the heat flux density, and the weight coefficient corresponding to the temperature rise rate are usually set between 0 and 1. First, by constructing a mapping table between the average temperature and the weight coefficient, the real-time detected average temperature is input into the mapping table in the database, so as to quickly obtain the weight coefficient corresponding to the average temperature. At the same time, for the heat flux density, by constructing a mapping table between the heat flux density and the weight coefficient, the real-time detected heat flux density is input into the mapping table in the database, so as to quickly obtain the weight coefficient corresponding to the heat flux density; for the weight coefficient corresponding to the temperature rise rate, a mapping table between the temperature rise rate and the weight coefficient is also established in advance, and the real-time measured temperature rise rate is input into the mapping table in the database, so as to quickly obtain the weight coefficient corresponding to the temperature rise rate.

[0037] Specifically, the prediction data of each node is analyzed and corrected. The specific process is as follows: count the number of monitoring periods in which the prediction data deviation value of each node is higher than or equal to the preset average deviation threshold. If the number of monitoring periods in which the prediction data deviation value of each node is higher than or equal to the preset average deviation threshold is higher than the set deviation period data quantity threshold, then the prediction data of each node is analyzed and corrected.

[0038] It should be noted that the specific steps for analyzing and correcting the prediction data of each node are as follows: first, regard the prediction residual sequences of the temperature, heat flux density, and temperature rise rate of all monitoring periods of this node as new time series signals, calculate the dynamic bias for the residual sequence through weighted moving average, and then subtract this bias from the subsequent temperature or heat flux density prediction values to eliminate long-term drift and systematic errors.

[0039] Specifically, the process of obtaining the thermal risk scoring index for each node is as follows: Extract the sensing data of each node in the computing cluster platform, including the total power consumption, the average fan speed, and the coolant flow rate, and compare them with their corresponding reference values respectively. After introducing the weight coefficients and the average deviation values of the prediction data of each node, perform coupling to obtain the thermal risk scoring index for each node. The thermal risk scoring index for each node is used to evaluate the thermal risk level of each node in the high-performance computing cluster platform.

[0040] It should be noted that the fan speed is usually monitored through sensors on the server motherboard. For example, the IPMI (Intelligent Platform Management Interface) tool can be used to obtain the rotation speed information of each fan. The coolant flow rate is specifically obtained by installing flow sensors at key nodes (such as cold plates and heat exchangers) where the coolant flows through to obtain the coolant flow rate at the key nodes in the cooling circuit.

[0041] It should be noted that the specific analysis conditions for the thermal risk scoring index of each node are as follows: ; In the formula, represents the thermal risk scoring index of the i-th node, represents the total power consumption of the i-th node, represents the set reference total power consumption of the node, represents the average fan speed of the i-th node, represents the set reference average fan speed of the node, represents the coolant flow rate of the i-th node, represents the set reference coolant flow rate of the node, represents the average deviation value of the prediction data of the i-th node, represents the weight coefficient corresponding to the total power consumption of the set node, represents the weight coefficient corresponding to the average fan speed of the set node, represents the weight coefficient corresponding to the set coolant flow rate, represents the weight coefficient corresponding to the average deviation value of the set prediction data. i represents the number of each node, , and n is the total number of nodes.

[0042] It should be noted that in the heat dissipation control of a high-performance computing cluster, the total power consumption, the average fan speed, the coolant flow rate, and the average deviation value of the prediction data are closely coupled. The total power consumption reflects the real-time energy consumption of all computing and cooling devices. An increase in power consumption usually means a greater heat load, which requires an increase in the fan speed to improve the air heat transfer efficiency and an increase in the circulation volume of the cooling medium to dissipate heat. The increase in the average fan speed can immediately accelerate air convection, but it will also increase the motor load and may introduce additional noise and energy consumption. At the same time, by comparing the node prediction errors before and after the heat dissipation action, the compensation effect of the current control strategy on temperature fluctuations can be quantified. If the average deviation value increases significantly, it indicates that the current adjustment of the fan and pump speeds is insufficient, and the fan speed also needs to be increased.

[0043] In a specific embodiment, the value ranges of the weight coefficients corresponding to the total power consumption of the node, the weight coefficient corresponding to the average fan speed, the weight coefficient corresponding to the coolant flow rate, and the weight coefficient corresponding to the average deviation value of the prediction data are usually set between 0 and 1. First, by constructing a mapping table between the average fan speed of the node and the weight coefficient, the real-time detected average fan speed of the node is input into the mapping table in the database, so as to quickly obtain the weight coefficient corresponding to the average fan speed. At the same time, for the total power consumption of the node, by constructing a mapping table between the total power consumption of the node and the weight coefficient, the real-time detected total power consumption of the node is input into the mapping table in the database, so as to quickly obtain the weight coefficient corresponding to the total power consumption of the node; for the weight coefficient corresponding to the coolant flow rate, a mapping table between the coolant flow rate and the weight coefficient is also established in advance, and the real-time measured coolant flow rate is input into the mapping table in the database, so as to quickly obtain the weight coefficient corresponding to the coolant flow rate; for the average deviation value of the prediction data, a mapping table between the average deviation value of the prediction data and the weight coefficient is also established in advance, and the real-time measured average deviation value of the prediction data is input into the mapping table in the database, so as to quickly obtain the weight coefficient corresponding to the average deviation value of the prediction data.

[0044] Specifically, the hot spots of the high-performance computing cluster platform are divided according to the heat risk score index of each node to obtain each risk area cluster. The specific process is as follows: extract the heat risk score index of each node and compare it with each risk area cluster corresponding to each interval of the set heat risk score index of each node to obtain each risk area cluster. Each risk area cluster includes each safe area cluster, each warning area cluster, and each hot spot area cluster.

[0045] It should be noted that all nodes are spatially clustered according to their physical locations first, and then safety zone clusters, warning zone clusters, and hot spot zone clusters are divided based on the average heat risk scoring index of the nodes within the clusters. By calculating the quantile thresholds of the heat risk scoring indices of each node, that is, after sorting the heat risk scoring indices of each node, the first threshold and the second threshold stored in the database are extracted, and the nodes with the average heat risk scoring index of the internal nodes lower than the first threshold are regarded as the safety zone clusters, the nodes with the average heat risk scoring index of the internal nodes between the two thresholds are regarded as the warning zone clusters, and the nodes with the average heat risk scoring index of the internal nodes higher than the second threshold are regarded as the hot spot zone clusters.

[0046] Specifically, calculation tasks are allocated according to each risk zone cluster. The specific analysis process is as follows: After the scheduler sorts each risk zone cluster according to the cluster-level risk, it preferentially allocates compute-intensive or long-term batch processing tasks to the nodes in the safety zone cluster and issues an instruction to suspend the new task scheduling instructions for each warning zone cluster and hot spot zone cluster.

[0047] It should be noted that the allocation of calculation tasks according to each risk zone cluster is specifically as follows: In the task scheduling stage, the system preferentially allocates newly submitted CPU / GPU-intensive computing tasks to the nodes within the safety zone cluster to avoid adding load to the existing risk zone clusters; for the nodes in the warning zone cluster, the system will carefully allocate tasks according to their current load and cooling capacity, continuously monitor the inlet and outlet air temperature difference and heat flux density of the nodes. If the temperature difference continues to rise or the predicted temperature exceeds the set safety threshold, a dynamic fine-tuning mechanism will be triggered, and some tasks will be seamlessly migrated to the safety zone cluster; while the nodes in the hot spot zone cluster will be temporarily excluded from task allocation until their heat risk score drops to the safe range.

[0048] Specifically, multi-level cooling responses of each risk zone cluster are triggered and monitored. The specific analysis process is as follows: A primary cooling response to the computing cluster platform is carried out in each safety zone cluster, a secondary cooling response to the computing cluster platform is carried out in each warning zone cluster and each hot spot zone cluster, and at the same time, the hot spot data of each hot spot area cluster is continuously monitored.

[0049] It should be noted that triggering a primary cooling response in each safety zone cluster includes issuing dynamic voltage and frequency adjustment instructions in each safety zone cluster. When a primary response is triggered in the safety zone cluster, the scheduler issues a DVFS command, usually decreasing the processor frequency in steps of 10–20% (such as from 2.5 GHz to 2.0–2.25 GHz) and synchronously reducing the core voltage to quickly suppress the heat source output.

[0050] It should be noted that after the first-level cooling response is triggered in each warning area cluster and each hot spot area cluster, the second-level cooling response is triggered continuously. This includes that the system first increases the CDU pump speed in steps of 5-10% (such as increasing from 60% PWM to 70-80%) and finely adjusts the opening of the bypass valve (increasing the opening by 5% per step). It also enables immersion phase change cooling and the backdoor heat exchanger, and seamlessly invokes the redundant heat dissipation resources of adjacent safe area clusters, migrating some loads over to enhance the coolant flow rate and heat exchange effect. If the system subsequently monitors that each warning area cluster and each hot spot area cluster return to the safe range, the pump speed, valve opening, and processor frequency are gradually stepped back in the opposite direction at the same ratio to achieve a millisecond-level closed-loop cooling and optimal energy consumption balance.

[0051] Specifically, the cooling effect response values of each risk area cluster are obtained. The specific process is as follows: A preset monitoring time period is set. During the monitoring time period, the hot spot data of each node in each hot spot area cluster is extracted, including the maximum value of the temperature rise rate and the defined maximum value of the temperature rise rate, the cumulative number of heat drop frequency events and the defined cumulative number of heat drop frequency events, and the average task completion rate and the reference average task completion rate are respectively analyzed for proportion. After introducing the weight coefficient and the heat risk scoring index of each node, they are coupled to obtain the cooling effect response values of each risk area cluster. The cooling effect response values of each risk area cluster are used to evaluate the overall cooling effect degree of the cooling measures on the hot spot cluster.

[0052] It should be noted that the specific analysis conditions for the cooling effect response values of each risk area cluster are as follows: ; In the formula, represents the cooling effect response value of the jth hot spot area cluster, represents the maximum value of the temperature rise rate of the ith node in the jth hot spot area cluster, represents the cumulative number of heat drop frequency events of the ith node in the jth hot spot area cluster, represents the average task completion rate of the ith node in the jth hot spot area cluster, represents the heat risk scoring index of the ith node, represents the defined maximum value of the temperature rise rate, represents the defined cumulative number of heat drop frequency events, represents the defined reference average task completion rate, represents the weight coefficient corresponding to the maximum value of the temperature rise rate, represents the weight coefficient corresponding to the number of heat drop frequency events, represents the weight coefficient corresponding to the average task completion rate, represents the weight coefficient corresponding to the heat risk scoring index. i represents the number of each node, , where n is the total number of nodes, and j represents the number of each hot spot area cluster, , and m is the total number of hot spot area clusters.

[0053] It should be noted that if the average task completion rate is too high, it may be accompanied by hasty scheduling or resource contention, which may instead lead to more frequent frequency reduction interventions and inhibitory cooling, thus potentially reducing the overall stability; the maximum value of the temperature rise rate represents the most severe instantaneous thermal shock suffered by the node during the monitoring period, which can accurately capture the extreme performance of the cooling system under peak load, and thus more effectively evaluate the inhibitory effect of the cooling strategy on the most dangerous thermal load.

[0054] It should be noted that in the heat dissipation control of a high-performance computing cluster, there is a tight coupling among the maximum value of the temperature rise rate, the cumulative number of thermal frequency reduction events, the average task completion rate, and the thermal risk scoring index of each node. The maximum value of the temperature rise rate reflects the speed at which the node temperature rises. If the maximum value of the temperature rise rate remains high, it may cause the processor to frequently trigger the thermal frequency reduction mechanism, thereby increasing the cumulative number of thermal frequency reduction events and affecting the system performance. And the increase in thermal frequency reduction events usually reduces the average task completion rate and prolongs the execution time of the computing task. When the thermal risk scoring index increases, it means that the node has high power consumption and insufficient adjustment of the fan and pump speeds, which will further push up the extreme value of the temperature rise rate and the number of frequency reduction times and depress the throughput; conversely, nodes with a lower thermal risk scoring index show gentle temperature rise, few frequency reductions, and high task rates.

[0055] In a specific embodiment, the value ranges of the weight coefficients corresponding to the maximum value of the temperature rise rate, the number of thermal frequency reduction events, the average task completion rate, and the thermal risk scoring index are usually set between 0 and 1. First, by constructing a mapping table between the maximum value of the temperature rise rate of the node and the weight coefficient, the detected maximum value of the temperature rise rate in real time is input into the mapping table in the database, so as to quickly obtain the weight coefficient corresponding to the maximum value of the temperature rise rate. At the same time, for the number of thermal frequency reduction events, by constructing a mapping table between the number of thermal frequency reduction events and the weight coefficient, the detected number of thermal frequency reduction events in real time is input into the mapping table in the database, so as to quickly obtain the weight coefficient corresponding to the number of thermal frequency reduction events; for the weight coefficient corresponding to the average task completion rate, a mapping table between the average task completion rate and the weight coefficient is also established in advance, and the measured average task completion rate in real time is input into the mapping table in the database, so as to quickly obtain the weight coefficient corresponding to the coolant flow rate; for the thermal risk scoring index, a mapping table between the thermal risk scoring index and the weight coefficient is also established in advance, and the measured thermal risk scoring index in real time is input into the mapping table in the database, so as to quickly obtain the weight coefficient corresponding to the thermal risk scoring index.

[0056] It should be noted that the cumulative number of thermal throttling events is the number of DVFS or clock throttling triggered by overheating during the monitoring period of the CPU and GPU.

[0057] Specifically, the high-performance computing cluster platform is warned according to the cooling effect response values of each risk area cluster. The specific process is as follows: According to the cooling effect response values of each risk area cluster and compare them with the set cooling effect response thresholds of each risk area cluster. If the cooling effect response value of each risk area cluster is lower than or equal to the cooling effect response threshold of each risk area cluster, then continuously monitor each risk area cluster. If the cooling effect response value of each risk area cluster is higher than the cooling effect response threshold of each risk area cluster, then a warning is issued.

[0058] It should be noted that issuing a warning specifically means notifying the operation and maintenance team through the monitoring platform or email / SMS with the corresponding pre-warning information, and at the same time sending the secondary alarm information to the cooling controller to automatically execute the preset and configured emergency cooling script and task migration strategy.

[0059] Specifically, triggering hardware compensation, the specific process is as follows: Count the number of risk area clusters where the cooling effect response value of each risk area cluster is higher than the cooling effect response threshold of each risk area cluster. If the number of risk area clusters where the cooling effect response value of each risk area cluster is higher than the cooling effect response threshold of each risk area cluster is higher than or equal to the risk area cluster number threshold, then trigger hardware compensation and enable immersion phase change cooling.

[0060] It should be noted that Figure 2 is a schematic diagram of the logic flow of the control method for heat dissipation of the high-performance computing cluster platform of the present invention. First, through three-dimensional thermal field modeling and inputting into the LSTM + GBRT hybrid prediction model, the prediction of each node temperature and heat flux density is obtained, and the node-level thermal risk score is analyzed. Then the system divides the nodes into three risk area clusters of safe, warning, and hot spots according to the risk score, and intelligently schedules computing tasks in combination with the partition results to avoid high-risk nodes. Continuously and frequently monitor hot data such as the extreme value of the temperature rise rate, throttling events, and task throughput recovery in each risk area cluster, and trigger multi-level cooling measures including DVFS, liquid cooling pump speed, and phase change cooling accordingly. Finally, judge the overall cooling performance through the cooling effect response value and automatically issue a warning and call for redundant hardware compensation, with the cooling effect closed-loop feedback, significantly improving the heat dissipation response speed and energy efficiency ratio, and ensuring that the high-performance computing cluster still maintains thermal stability and computing power stability under extremely high loads.

[0061] Figure 3This is a schematic diagram of the logic flow for analyzing and correcting the predicted data of each node in the present invention. The system collects the predicted values and measured values of the nodes in each monitoring cycle, calculates the deviation between the two, and statistically counts the number of cycles with excessive deviation. When the deviation of a certain node exceeds the threshold for multiple consecutive cycles, the analysis and correction of the predicted data of this node are triggered. After the correction is completed, the result is immediately applied to the parameter adjustment of the heat dissipation control strategy to enhance the accuracy of the cooling response. If the threshold is not exceeded, continuous monitoring is continued, and both finally converge to the closed-loop link of the optimized control strategy.

[0062] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0063] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0064] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0065] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0066] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made by those skilled in the art once they learn of the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present invention.

[0067] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. A control method for heat dissipation of a high-performance computing cluster platform, characterized in that: include: Construct a three-dimensional thermal field model of a high-performance computing cluster platform, obtain the sensor data of each node of the computing cluster platform, import it into the LSTM and GBRT hybrid prediction model, and obtain the average deviation value of the prediction data of each node; The predicted data of each node is analyzed and corrected according to the average deviation value of the predicted data of each node, and the thermal risk score index of each node is obtained by combining the sensor data of each node; The hot spots of the high-performance computing cluster platform are divided according to the thermal risk score index of each node to obtain risk clusters, and computing tasks are allocated according to each risk cluster; Monitor each node in each risk area cluster, obtain and analyze the hotspot data of each node in each risk area cluster, trigger a multi-level cooling response in each risk area cluster and monitor it, obtain the cooling effect response value of each risk area cluster, and issue an early warning to the high-performance computing cluster platform and trigger hardware compensation based on the cooling effect response value of each risk area cluster.

2. A control method for heat dissipation of a high performance computing cluster platform as claimed in claim 1, characterized in that: The specific process of obtaining the average deviation value of the predicted data of each node is as follows: Preset each monitoring cycle, and extract the prediction data of each node in each monitoring cycle, including temperature prediction value, heat flux prediction value and temperature rise prediction rate; In each monitoring cycle, the actual data of each current node, including the average temperature, heat flux density and temperature rise rate, are obtained, and compared with the predicted data of each node in the previous time, and coupled after introducing the weight coefficient to obtain the average deviation value of the predicted data of each node. The average deviation value of the predicted data of each node is used to evaluate the accuracy of the prediction model.

3. A control method for heat dissipation of a high performance computing cluster platform as claimed in claim 1, characterized in that: The specific process of analyzing and correcting the prediction data of each node is as follows: The number of monitoring cycles in which the predicted data deviation value of each node is higher than or equal to the preset average deviation threshold is counted; if the number of monitoring cycles in which the predicted data deviation value of each node is higher than or equal to the preset average deviation threshold is higher than the set deviation cycle data quantity threshold, the predicted data of each node is analyzed and corrected.

4. A control method for heat dissipation of a high performance computing cluster platform as claimed in claim 2, characterized in that: The specific process of obtaining the thermal risk score index of each node is as follows: The sensor data of each node in the computing cluster platform are extracted, including total power consumption, average fan speed, and coolant flow rate, which are compared with their corresponding reference values ​​respectively. The weight coefficient and the average deviation value of the predicted data of each node are introduced and coupled to obtain the thermal risk score index of each node. The thermal risk score index of each node is used to evaluate the thermal risk level of each node in the high-performance computing cluster platform.

5. A control method for heat dissipation of a high performance computing cluster platform as claimed in claim 1, characterized in that: The hotspot areas of the high-performance computing cluster platform are divided according to the thermal risk score index of each node to obtain risk area clusters. The specific process is as follows: The thermal risk score index of each node is extracted and compared with the risk area clusters corresponding to the set intervals of the thermal risk score index of each node to obtain the risk area clusters, which include safety area clusters, warning area clusters and hot spot area clusters.

6. A control method for heat dissipation of a high performance computing cluster platform as claimed in claim 4, characterized in that: The calculation tasks are allocated according to each risk area cluster, and the specific analysis process is as follows: After sorting the risk zone clusters according to cluster-level risk, the scheduler preferentially allocates computationally intensive or long-duration batch processing tasks to the nodes in the safe zone cluster, and issues instructions to suspend new task scheduling instructions for the warning zone clusters and hotspot zone clusters.

7. A control method for heat dissipation of a high performance computing cluster platform as claimed in claim 4, characterized in that: The multi-level cooling response of each risk cluster is triggered and monitored in each risk cluster. The specific analysis process is as follows: A first-level cooling response is performed on the computing cluster platform in each safety zone cluster, and a second-level cooling response is performed on the computing cluster platform in each warning zone cluster and each hotspot zone cluster. At the same time, the hotspot data of each hotspot zone cluster is continuously monitored.

8. A control method for heat dissipation of a high performance computing cluster platform as claimed in claim 1, characterized in that: The specific process of obtaining the cooling effect response value of each risk zone cluster is as follows: A monitoring time period is preset, and hotspot data of each node in each hotspot area cluster is extracted during the monitoring time period, including the maximum temperature rise rate and the defined maximum temperature rise rate, the cumulative number of thermal frequency reduction events and the defined cumulative number of thermal frequency reduction events, the average rate of task completion and the average rate of reference task completion, and a proportion analysis is performed respectively. After introducing the weight coefficient and the thermal risk score index of each node, the cooling effect response value of each risk area cluster is obtained by coupling. The cooling effect response value of each risk area cluster is used to evaluate the overall cooling effectiveness of the cooling measures on the hotspot cluster.

9. A control method for heat dissipation of a high performance computing cluster platform as claimed in claim 1, characterized in that: The specific process of early warning the high performance computing cluster platform according to the cooling effect response value of each risk zone cluster is as follows: According to the cooling effect response value of each risk area cluster, it is compared with the set cooling effect response threshold of each risk area cluster. If the cooling effect response value of each risk area cluster is lower than or equal to the cooling effect response threshold of each risk area cluster, each risk area cluster is continuously monitored. If the cooling effect response value of each risk area cluster is higher than the cooling effect response threshold of each risk area cluster, an early warning is issued.

10. A control method for heat dissipation of a high performance computing cluster platform as claimed in claim 9, characterized in that: The specific process of triggering hardware compensation is as follows: The number of risk area clusters whose cooling effect response value of each risk area cluster is higher than the cooling effect response threshold of each risk area cluster is counted; if the number of risk area clusters whose cooling effect response value of each risk area cluster is higher than the cooling effect response threshold of each risk area cluster is higher than or equal to the risk area cluster number threshold, hardware compensation is triggered and immersion phase change cooling is enabled.

Citation Information

Patent Citations

  • A control method and system for heat dissipation of a high-performance computing cluster platform

    CN113721741B

  • High-performance computing cluster platform heat dissipation device and control system thereof

    CN117032425A

  • Real-time temperature rise measurement method of electromagnetic system of low-voltage apparatus based on finite element analysis

    CN103970947A

  • Frequency conversion power distribution cabinet operation supervision system based on sensor

    CN117353465A

  • Low-voltage power distribution control box and heat dissipation early warning method thereof

    CN118763541A

Cited By

  • Server heat dissipation device and heat dissipation control method

    CN120428835A

  • Server heat dissipation device and heat dissipation control method

    CN120428835B

  • Heat dissipation control method and system for set top box and storage medium

    CN120568114A

  • Task scheduling optimization system for smelting process

    CN120598315A

  • Task scheduling optimization system for a smelting process

    CN120598315B