A control method for heat dissipation of a high-performance computing cluster platform
By constructing a three-dimensional thermal field model and a hybrid prediction model, evaluating thermal risks and implementing multi-stage cooling response and hardware compensation, the problem of slow heat dissipation response of high-performance computing cluster platforms is solved, and efficient and reliable heat dissipation control is achieved.
Patent Information
- Application Number
- CN202510629070.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-05-16
AI Technical Summary
The passive heat dissipation of the high-performance computing cluster platform responds slowly to burst loads, making it difficult to take into account timeliness, energy efficiency and reliability, resulting in a sharp surge in chip temperature, triggering hardware protection frequency reduction and causing computing power loss.
Build a three-dimensional thermal field model, use the mixed prediction model of LSTM and GBRT to obtain node data, evaluate thermal risks, divide risk zone clusters, implement multi-level cooling response and hardware compensation, and optimize computing task allocation and heat dissipation resource scheduling.
It realizes high-precision and low-energy heat dissipation control, ensures system thermal stability and computing efficiency, avoids overheating failures, and improves hardware reliability and computing performance.
Smart Images

Figure CN120179044B_ABST
Abstract
Description
[0001] Data processing technology field
[0002] The present invention relates to the field of data processing technology, and in particular to a control method for heat dissipation of a high-performance computing cluster platform. Background Art
[0003] High-performance computing cluster platforms generate a significant amount of heat during operation. Inadequate or uneven heat dissipation can lead to system performance degradation, hardware damage, and even service interruptions. Furthermore, existing heat dissipation methods, such as air cooling, are no longer sufficient for high-density computing tasks. New cooling technologies, such as liquid cooling and thermoelectric cooling, are increasingly being adopted in high-performance computing clusters to improve heat dissipation efficiency and system energy efficiency. Therefore, building an efficient heat dissipation control system is crucial to ensuring stable operation of computing clusters.
[0004] For example, the announcement number CN113721741B announces a method and system for controlling heat dissipation of a high-performance computing cluster platform, which includes obtaining the process of job operations of a high-performance computing job scheduling system; adjusting the power of the heat dissipation equipment of the server computing card, server chassis and / or the cabinet where the server is located that executes the process of the job operation according to the process of the job operation; the passive adjustment method includes collecting temperature data and calculating temperature warning data; and adjusting the power of the heat dissipation equipment of the server computing card, server chassis and / or the cabinet where the server is located that has a temperature warning according to the temperature warning data.
[0005] For example, the publication number CN117032425A discloses a high-performance computing cluster platform heat dissipation device and its control system, which includes a host, a heat sink, a heat dissipation and fire extinguishing component, and a safety component. A power cord is plugged into the back of the host, the heat sink is fixedly covered on the top of the host, the heat dissipation and fire extinguishing component is fixedly arranged between the bottom of the heat sink and the back corner of the host, the bottom of the heat dissipation and fire extinguishing component is fixedly connected to a connecting plate, and a safety component for assisting in unplugging the power cord is fixedly arranged on the inner side of the connecting plate.
[0006] However, in the process of implementing the technical solutions of the invention in the embodiments of the present application, the present application found that the above technology has at least the following technical problems:
[0007] There is a problem that the passive cooling of the computing cluster platform responds slowly to sudden loads, and the cooling solution is difficult to balance timeliness, energy efficiency and reliability. This may lead to the inability to quickly match the transient heat flux density jump generated by sudden computing tasks, causing the chip temperature to soar sharply, triggering hardware protective frequency reduction and resulting in computing power loss. Summary of the Invention
[0008] In view of the deficiencies of the prior art, the present invention provides a control method for heat dissipation of a high-performance computing cluster platform, which solves the problems designed in the above-mentioned background technology.
[0009] To achieve the above object, the present invention is realized through the following technical solutions: A control method for heat dissipation of a high-performance computing cluster platform, including constructing a three-dimensional thermal field model of the high-performance computing cluster platform, obtaining sensing data of each node of the computing cluster platform, importing it into an LSTM and GBRT hybrid prediction model, and obtaining the average deviation value of the prediction data of each node.
[0010] Analyze and correct the prediction data of each node according to the average deviation value of the prediction data of each node, and combine the sensing data of each node to obtain the heat risk scoring index of each node.
[0011] Divide the hot spots of the high-performance computing cluster platform according to the heat risk scoring index of each node to obtain each risk area cluster, and allocate computing tasks according to each risk area cluster.
[0012] Monitor each node in each risk area cluster, obtain and analyze the hot spot data of each node in each risk area cluster, trigger a multi-level cooling response of each risk area cluster in each risk area cluster and monitor it to obtain the cooling effect response value of each risk area cluster, and give an early warning to the high-performance computing cluster platform according to the cooling effect response value of each risk area cluster and trigger hardware compensation.
[0013] Further, to obtain the average deviation value of the prediction data of each node, the specific process is as follows: preset each monitoring period, and extract the prediction data of each node in each monitoring period, including temperature prediction value, heat flux density prediction value, and temperature rise prediction rate.
[0014] Obtain the actual data of each node in each monitoring period, including average temperature, heat flux density, and temperature rise rate, compare them with the prediction data of each node in the previous period respectively, and perform coupling after introducing a weight coefficient to obtain the average deviation value of the prediction data of each node. The average deviation value of the prediction data of each node is used to evaluate the accuracy of the prediction model.
[0015] Further, to analyze and correct the prediction data of each node, the specific process is as follows: count the number of monitoring periods in which the prediction data deviation value of each node is higher than or equal to the preset average deviation threshold. If the number of monitoring periods in which the prediction data deviation value of each node is higher than or equal to the preset average deviation threshold is higher than the set deviation period data quantity threshold, then analyze and correct the prediction data of each node.
[0016] Further, the thermal risk scoring index of each node is obtained. The specific process is as follows: The sensing data of each node in the computing cluster platform is extracted, including the total power consumption, the average fan speed, and the coolant flow rate, which are compared with their corresponding reference values respectively. After introducing the weight coefficient and the average deviation value of the prediction data of each node, they are coupled to obtain the thermal risk scoring index of each node. The thermal risk scoring index of each node is used to evaluate the thermal risk level of each node in the high-performance computing cluster platform.
[0017] Further, the hot spots in the high-performance computing cluster platform are divided according to the thermal risk scoring index of each node to obtain each risk area cluster. The specific process is as follows: The thermal risk scoring index of each node is extracted and compared with each risk area cluster corresponding to each interval of the set thermal risk scoring index of each node to obtain each risk area cluster. Each risk area cluster includes each safety area cluster, each warning area cluster, and each hot spot area cluster.
[0018] Further, the computing tasks are allocated according to each risk area cluster. The specific analysis process is as follows: After the scheduler sorts each risk area cluster according to the cluster-level risk, it preferentially allocates compute-intensive or long-duration batch processing tasks to the nodes in the safety area cluster and issues instructions to suspend the new task scheduling instructions of each warning area cluster and hot spot area cluster.
[0019] Further, a multi-level cooling response of each risk area cluster is triggered and monitored. The specific analysis process is as follows: A primary cooling response to the computing cluster platform is performed in each safety area cluster, and a secondary cooling response to the computing cluster platform is performed in each warning area cluster and hot spot area cluster. At the same time, the hot spot data of each hot spot area cluster is continuously monitored.
[0020] Further, the cooling effect response value of each risk area cluster is obtained. The specific process is as follows: A preset monitoring time period is set. In the monitoring time period, the hot spot data of each node in each hot spot area cluster is extracted, including the maximum value of the temperature rise rate and the defined maximum value of the temperature rise rate, the cumulative number of thermal drop frequency events and the defined cumulative number of thermal drop frequency events, and the average task completion rate and the reference average task completion rate are respectively analyzed for proportion. After introducing the weight coefficient and the thermal risk scoring index of each node, they are coupled to obtain the cooling effect response value of each risk area cluster. The cooling effect response value of each risk area cluster is used to evaluate the overall cooling effectiveness of the cooling measures for the hot spot cluster.
[0021] Further, an early warning is issued for the high-performance computing cluster platform according to the cooling effect response value of each risk area cluster. The specific process is as follows: According to the cooling effect response value of each risk area cluster, it is compared with the set cooling effect response threshold of each risk area cluster. If the cooling effect response value of each risk area cluster is lower than or equal to the cooling effect response threshold of each risk area cluster, each risk area cluster is continuously monitored. If the cooling effect response value of each risk area cluster is higher than the cooling effect response threshold of each risk area cluster, an early warning is issued.
[0022] Further, trigger hardware compensation. The specific process is as follows: Count the number of risk area clusters whose cooling effect response values are higher than the cooling effect response thresholds of their respective risk area clusters. If the number of risk area clusters whose cooling effect response values are higher than the cooling effect response thresholds of their respective risk area clusters is greater than or equal to the risk area cluster number threshold, trigger hardware compensation and enable immersion phase change cooling.
[0023] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:
[0024] The present invention has the following beneficial effects:
[0025] (1) A control method for heat dissipation of a high-performance computing cluster platform provided by the present invention first constructs a three-dimensional thermal field model of the cluster. Then, according to the distribution quantiles of the TRS of all nodes, risk clusters are divided, and computing-intensive or long-term batch processing tasks are preferentially deployed. On this basis, a multi-level cooling response strategy is adopted, and the nodes within the cluster are continuously monitored. Finally, the monitoring of the cooling effect response value triggers hardware compensation measures to prevent overheating and loss of control, and continuous closed-loop monitoring and feedback are performed, and through the monitoring and response data feedback, to achieve high-precision, low-energy consumption, and high-reliability closed-loop control of the heat dissipation of the high-performance computing cluster.
[0026] (2) The present invention effectively solves the problem of prediction distortion caused by load fluctuations in the heat dissipation control of a high-performance computing cluster through the average prediction deviation value at the node level. Subsequently, the online analysis and correction process triggered by errors is statistically analyzed for each node, thereby ensuring that subsequent heat risk scoring and task scheduling based on the prediction results are always based on the output of a highly reliable model, which helps to immediately adjust the cooling strategy and maintain the thermal stability of the system.
[0027] (3) By obtaining the heat risk scoring index of each node, the present invention helps subsequent cluster-level clustering, and evaluates the overall cooling effect by monitoring the energy efficiency index, accurately evaluates its heat dissipation pressure, and preferentially assigns high-load tasks to the nodes in the low-temperature safe area through the scheduling system, while restricting the computing power allocation in the high-temperature area. This strategy forms a closed-loop control of "monitoring - evaluation - scheduling", effectively achieving precise placement of heat dissipation resources by dynamically dividing risk areas, avoiding global overheating; intelligently scheduling tasks in combination with the heat risk state, not only ensuring computing efficiency but also alleviating local heat dissipation pressure, and avoiding potential failures in advance through heat state prediction, improving the reliability of the hardware. The entire process realizes the collaborative management of heat dissipation control and computing load.
[0028] (4) By obtaining the heat dissipation effect score values of each region, the present invention evaluates the actual effect of the heat dissipation measures through multi-dimensional index fusion, avoids misjudgment of a single temperature parameter, establishes a dynamic early warning mechanism, can intervene in advance when the heat dissipation efficiency is insufficient, prevent chain failures caused by local overheating, and correlates the heat dissipation efficiency with the task execution efficiency, achieving precise regulation of heat dissipation resources while ensuring computing performance. The whole method improves the response speed and reliability of the heat dissipation system, and forms a dynamic balance between the heat dissipation efficiency and the computing load. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 is a schematic diagram of the method of the present invention;
[0030] Figure 2 is a schematic diagram of the logical flow of the control method for heat dissipation of a high-performance computing cluster platform according to the present invention;
[0031] Figure 3 is a schematic diagram of the logical flow of analyzing and correcting the prediction data of each node according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0032] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0033] In the description of the present invention, it should be understood that the terms "opening", "upper", "lower", "thickness", "top", "middle", "length", "inner", "perimeter", etc. indicating orientation or positional relationships are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the components or elements referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be construed as a limitation of the present invention.
[0034] Please refer to Figure 1 , the embodiments of the present invention provide a technical solution: a control method for heat dissipation of a high-performance computing cluster platform, including constructing a three-dimensional thermal field model of the high-performance computing cluster platform, obtaining the sensing data of each node of the computing cluster platform, and importing it into the LSTM and GBRT hybrid prediction model to obtain the average deviation value of the prediction data of each node.
[0035] Analyze and correct the prediction data of each node according to the average deviation value of the prediction data of each node, and combine the sensing data of each node to obtain the heat risk score index of each node.
[0036] Divide the hot spots of the high-performance computing cluster platform according to the thermal risk scoring index of each node to obtain each risk area cluster, and allocate computing tasks according to each risk area cluster.
[0037] Monitor each node in each risk area cluster, obtain the hot spot data of each node in each risk area cluster and analyze it, trigger the multi-level cooling response of each risk area cluster in each risk area cluster and monitor it, obtain the cooling effect response value of each risk area cluster, and give an early warning to the high-performance computing cluster platform according to the cooling effect response value of each risk area cluster and trigger hardware compensation.
[0038] Specifically, obtain the average deviation value of the prediction data of each node. The specific process is as follows: preset each monitoring period, and extract the prediction data of each node in each monitoring period, including the temperature prediction value, the heat flux density prediction value, and the temperature rise prediction rate.
[0039] Obtain the actual data of each current node in each monitoring period, including the average temperature, the heat flux density, and the temperature rise rate, compare them with the prediction data of each previous node respectively, and perform coupling after introducing the weight coefficient to obtain the average deviation value of the prediction data of each node. The average deviation value of the prediction data of each node is used to evaluate the accuracy of the prediction model.
[0040] ;
[0041] represents the average deviation value of the prediction data of the i-th node, represents the average temperature of the i-th node, represents the predicted average temperature of the i-th node in the previous time, represents the set defined reference average temperature deviation amount, represents the heat flux density of the i-th node, represents the predicted heat flux density of the i-th node in the previous time, represents the set defined heat flux density deviation amount, represents the temperature rise rate of the i-th node, represents the predicted temperature rise rate of the i-th node in the previous time, represents the set defined temperature rise rate deviation amount, represents the weight coefficient corresponding to the set average temperature, represents the weight coefficient corresponding to the set heat flux density, represents the weight coefficient corresponding to the set temperature rise rate, i represents the number of each node, , n is the total number of nodes.
[0042] It should be noted that in the heat dissipation control of a high-performance computing cluster, the average temperature, heat flux density, and temperature rise rate are closely coupled: when the heat flux density exceeds the instantaneous processing capacity of the heat dissipation system, the temperature rise rate will rapidly climb. At the same time, the continuous high temperature rise rate will act on the average temperature, causing it to further exceed the safety threshold and exacerbating the heat flux density demand. If not intervened in time, it may lead to hotspot instability or device thermal throttling, triggering performance fluctuations and reliability risks. And when the heat flux density increases, if the heat dissipation measures are insufficient, the average temperature will rise, resulting in an accelerated temperature rise rate. Conversely, if the average temperature continues to rise, it may indicate that the heat flux density is too high or the heat dissipation efficiency has decreased, thereby triggering an increase in the temperature rise rate.
[0043] In a specific embodiment, the value ranges of the weight coefficients corresponding to the average temperature, the weight coefficients corresponding to the heat flux density, and the weight coefficients corresponding to the temperature rise rate are usually set between 0 and 1. First, by constructing a mapping table between the average temperature and the weight coefficient, the real-time detected average temperature is input into the mapping table in the database, so as to quickly obtain the weight coefficient corresponding to the average temperature. At the same time, for the heat flux density, by constructing a mapping table between the heat flux density and the weight coefficient, the real-time detected heat flux density is input into the mapping table in the database, so as to quickly obtain the weight coefficient corresponding to the heat flux density; for the weight coefficient corresponding to the temperature rise rate, a mapping table between the temperature rise rate and the weight coefficient is also established in advance, and the real-time measured temperature rise rate is input into the mapping table in the database, so as to quickly obtain the weight coefficient corresponding to the temperature rise rate.
[0044] Specifically, the prediction data of each node is analyzed and corrected. The specific process is as follows: count the number of monitoring periods in which the prediction data deviation value of each node is higher than or equal to the preset average deviation threshold. If the number of monitoring periods in which the prediction data deviation value of each node is higher than or equal to the preset average deviation threshold is higher than the set deviation period data quantity threshold, then the prediction data of each node is analyzed and corrected.
[0045] It should be noted that the specific steps for analyzing and correcting the prediction data of each node are as follows: first, regard the prediction residual sequences of the temperature, heat flux density, and temperature rise rate of all monitoring periods of this node as new time series signals, calculate the dynamic bias for the residual sequence through weighted moving average, and then subtract this bias from the subsequent temperature or heat flux density prediction values to eliminate long-term drift and systematic errors.
[0046] Specifically, the process of obtaining the thermal risk scoring index for each node is as follows: Extract the sensing data of each node in the computing cluster platform, including the total power consumption, average fan speed, and coolant flow rate, and compare them with their corresponding reference values. After introducing the weight coefficients and the average deviation values of the prediction data for each node, perform coupling to obtain the thermal risk scoring index for each node. The thermal risk scoring index for each node is used to evaluate the thermal risk level of each node in the high-performance computing cluster platform.
[0047] It should be noted that the fan speed is usually monitored through sensors on the server motherboard. For example, the IPMI (Intelligent Platform Management Interface) tool can be used to obtain the rotation speed information of each fan. The coolant flow rate is specifically obtained by installing flow sensors at key nodes (such as cold plates and heat exchangers) where the coolant flows through to obtain the coolant flow rate at the key nodes in the cooling circuit.
[0048] It should be noted that the specific analysis conditions for the thermal risk scoring index of each node are as follows:
[0049] ;
[0050] In the formula, represents the thermal risk scoring index of the i-th node, represents the total power consumption of the i-th node, represents the set reference total power consumption of the node, represents the average fan speed of the i-th node, represents the set reference average fan speed of the node, represents the coolant flow rate of the i-th node, represents the set reference coolant flow rate of the node, represents the average deviation value of the prediction data of the i-th node, represents the weight coefficient corresponding to the total power consumption of the set node, represents the weight coefficient corresponding to the average fan speed of the set node, represents the weight coefficient corresponding to the set coolant flow rate, represents the weight coefficient corresponding to the average deviation value of the set prediction data. i represents the number of each node, , and n is the total number of nodes.
[0051] It should be noted that in the heat dissipation control of a high-performance computing cluster, the total power consumption, the average fan speed, the coolant flow rate, and the average deviation value of the prediction data are closely coupled. The total power consumption reflects the real-time energy consumption of all computing and cooling devices. An increase in power consumption usually means a greater heat load, which requires an increase in the fan speed to improve the air heat transfer efficiency and an increase in the circulation volume of the cooling medium to dissipate heat. The increase in the average fan speed can immediately accelerate air convection, but it will also increase the motor load and may introduce additional noise and energy consumption. At the same time, by comparing the node prediction errors before and after the heat dissipation action, the compensation effect of the current control strategy on temperature fluctuations can be quantified. If the average deviation value increases significantly, it indicates that the current fan and pump speed adjustments are insufficient, and the fan speed also needs to be increased.
[0052] In a specific embodiment, the value ranges of the weight coefficients corresponding to the total power consumption of the node, the weight coefficient corresponding to the average fan speed, the weight coefficient corresponding to the coolant flow rate, and the weight coefficient corresponding to the average deviation value of the prediction data are usually set between 0 and 1. First, by constructing a mapping table between the average fan speed of the node and the weight coefficient, the real-time detected average fan speed of the node is input into the mapping table in the database, so as to quickly obtain the weight coefficient corresponding to the average fan speed. At the same time, for the total power consumption of the node, by constructing a mapping table between the total power consumption of the node and the weight coefficient, the real-time detected total power consumption of the node is input into the mapping table in the database, so as to quickly obtain the weight coefficient corresponding to the total power consumption of the node; for the weight coefficient corresponding to the coolant flow rate, a mapping table between the coolant flow rate and the weight coefficient is also established in advance, and the real-time measured coolant flow rate is input into the mapping table in the database, so as to quickly obtain the weight coefficient corresponding to the coolant flow rate; for the average deviation value of the prediction data, a mapping table between the average deviation value of the prediction data and the weight coefficient is also established in advance, and the real-time measured average deviation value of the prediction data is input into the mapping table in the database, so as to quickly obtain the weight coefficient corresponding to the average deviation value of the prediction data.
[0053] Specifically, the hot spot areas of the high-performance computing cluster platform are divided according to the heat risk scoring index of each node to obtain each risk area cluster. The specific process is as follows: extract the heat risk scoring index of each node and compare it with each risk area cluster corresponding to each interval of the set heat risk scoring index of each node to obtain each risk area cluster. Each risk area cluster includes each safety area cluster, each warning area cluster and each hot spot area cluster.
[0054] It should be noted that all nodes are spatially clustered according to their physical locations first, and then safety zone clusters, warning zone clusters, and hot spot zone clusters are divided based on the average heat risk scoring index of the nodes within the clusters. By calculating the quantile thresholds of the heat risk scoring indices of each node, that is, after sorting the heat risk scoring indices of each node, the first threshold and the second threshold stored in the database are extracted, and the nodes with the average heat risk scoring index of the internal nodes lower than the first threshold are regarded as safety zone clusters, the nodes with the average heat risk scoring index of the internal nodes between the two thresholds are regarded as warning zone clusters, and the nodes with the average heat risk scoring index of the internal nodes higher than the second threshold are regarded as hot spot zone clusters.
[0055] Specifically, calculation tasks are allocated according to each risk zone cluster. The specific analysis process is as follows: After the scheduler sorts each risk zone cluster according to the cluster-level risk, it preferentially allocates compute-intensive or long-term batch processing tasks to the nodes in the safety zone cluster, and issues instructions to suspend the new task scheduling instructions for each warning zone cluster and hot spot zone cluster.
[0056] It should be noted that the allocation of calculation tasks according to each risk zone cluster is specifically as follows: In the task scheduling stage, the system preferentially allocates newly submitted CPU / GPU-intensive computing tasks to the nodes within the safety zone cluster to avoid adding load to the existing risk zone clusters; for the nodes in the warning zone cluster, the system will carefully allocate tasks according to their current load and cooling capacity, continuously monitor the inlet and outlet air temperature difference and heat flux density of the nodes. If the temperature difference continues to rise or the predicted temperature exceeds the set safety threshold, a dynamic fine-tuning mechanism will be triggered, and some tasks will be seamlessly migrated to the safety zone cluster; while the nodes in the hot spot zone cluster will be temporarily excluded from task allocation until their heat risk score drops to the safe range.
[0057] Specifically, multi-level cooling responses of each risk zone cluster are triggered and monitored. The specific analysis process is as follows: A primary cooling response is carried out on the computing cluster platform in each safety zone cluster, and a secondary cooling response is carried out on the computing cluster platform in each warning zone cluster and each hot spot zone cluster, while continuously monitoring the hot spot data of each hot spot area cluster.
[0058] It should be noted that triggering a primary cooling response in each safety zone cluster includes issuing dynamic voltage and frequency adjustment instructions in each safety zone cluster. When a primary response is triggered in the safety zone cluster, the scheduler issues a DVFS command, usually decreasing the processor frequency in steps of 10–20% (such as from 2.5 GHz to 2.0–2.25 GHz) and synchronously reducing the core voltage to quickly suppress the heat source output.
[0059] It should be noted that after the first-level cooling response is triggered for each warning area cluster and each hot spot area cluster, the second-level cooling response is continued to be triggered. This includes that the system first increases the CDU pump speed in steps of 5–10% (such as from 60% PWM to 70–80%) and finely adjusts the opening of the bypass valve (increasing the opening by 5% per step). The immersion phase change cooling and the backdoor heat exchanger are also enabled, and the redundant heat dissipation resources of adjacent safe area clusters are seamlessly called to transfer part of the load over to enhance the coolant flow rate and heat exchange effect. If the system subsequently monitors that each warning area cluster and each hot spot area cluster return to the safe range, the pump speed, valve opening, and processor frequency are gradually stepped back in the opposite direction at the same ratio to achieve a millisecond-level closed-loop cooling and optimal energy consumption balance.
[0060] Specifically, the cooling effect response values of each risk area cluster are obtained. The specific process is as follows: A preset monitoring time period is set. During the monitoring time period, the hot spot data of each node in each hot spot area cluster are extracted, including the maximum value of the temperature rise rate and the defined maximum value of the temperature rise rate, the cumulative number of heat drop frequency events and the defined cumulative number of heat drop frequency events, and the average task completion rate and the reference average task completion rate are respectively analyzed for proportion. After introducing the weight coefficient and the heat risk scoring index of each node, they are coupled to obtain the cooling effect response values of each risk area cluster. The cooling effect response values of each risk area cluster are used to evaluate the overall cooling effectiveness of the cooling measures for the hot spot cluster.
[0061] It should be noted that the specific analysis conditions for the cooling effect response values of each risk area cluster are:
[0062] ;
[0063] In the formula, represents the cooling effect response value of the j-th hot spot area cluster, represents the maximum value of the temperature rise rate of the i-th node in the j-th hot spot area cluster, represents the cumulative number of heat drop frequency events of the i-th node in the j-th hot spot area cluster, represents the average task completion rate of the i-th node in the j-th hot spot area cluster, represents the heat risk scoring index of the i-th node, represents the defined maximum value of the temperature rise rate, represents the defined cumulative number of heat drop frequency events, represents the defined reference average task completion rate, represents the weight coefficient corresponding to the maximum value of the temperature rise rate, represents the weight coefficient corresponding to the number of heat drop frequency events, represents the weight coefficient corresponding to the average task completion rate, represents the weight coefficient corresponding to the heat risk scoring index, and i represents the number of each node. , where n is the total number of nodes, and j represents the number of each hot - spot area cluster, , and m is the total number of hot - spot area clusters.
[0064] It should be noted that if the average task - completion rate is too high, it may be accompanied by hasty scheduling or resource contention, which may lead to more frequent down - clocking interventions and inhibitory cooling, thus possibly reducing the overall stability; the maximum value of the temperature - rise rate represents the most severe instantaneous thermal shock suffered by the node during the monitoring period, which can accurately capture the limit performance of the cooling system under peak load, and thus more effectively evaluate the inhibitory effect of the cooling strategy on the most dangerous thermal load.
[0065] It should be noted that in the heat dissipation control of a high - performance computing cluster, there is a tight coupling among the maximum value of the temperature - rise rate, the cumulative number of thermal down - clocking events, the average task - completion rate, and the thermal - risk scoring index of each node. The maximum value of the temperature - rise rate reflects the speed at which the node temperature rises. If the maximum value of the temperature - rise rate remains too high, it may cause the processor to frequently trigger the thermal down - clocking mechanism, thereby increasing the cumulative number of thermal down - clocking events and affecting the system performance. And the increase in thermal down - clocking events usually reduces the average task - completion rate and prolongs the execution time of the computing task. When the thermal - risk scoring index increases, it means that the node has high power consumption and insufficient adjustment of fan and pump speeds, so it will further push up the extreme value of the temperature - rise rate and the number of down - clocking times and depress the throughput; conversely, nodes with a lower thermal - risk scoring index show gentle temperature rise, fewer down - clocking times, and high task rates.
[0066] In a specific embodiment, the value ranges of the weight coefficients corresponding to the maximum value of the temperature - rise rate, the number of thermal down - clocking events, the average task - completion rate, and the thermal - risk scoring index are usually set between 0 and 1. First, by constructing a mapping table between the maximum value of the temperature - rise rate of the node and the weight coefficient, the real - time detected maximum value of the temperature - rise rate is input into the mapping table in the database, so as to quickly obtain the weight coefficient corresponding to the maximum value of the temperature - rise rate. At the same time, for the number of thermal down - clocking events, by constructing a mapping table between the number of thermal down - clocking events and the weight coefficient, the real - time detected number of thermal down - clocking events is input into the mapping table in the database, so as to quickly obtain the weight coefficient corresponding to the number of thermal down - clocking events; for the weight coefficient corresponding to the average task - completion rate, a mapping table between the average task - completion rate and the weight coefficient is also established in advance, and the real - time measured average task - completion rate is input into the mapping table in the database, so as to quickly obtain the weight coefficient corresponding to the coolant flow rate; for the thermal - risk scoring index, a mapping table between the thermal - risk scoring index and the weight coefficient is also established in advance, and the real - time measured thermal - risk scoring index is input into the mapping table in the database, so as to quickly obtain the weight coefficient corresponding to the thermal - risk scoring index.
[0067] It should be noted that the cumulative number of thermal throttling events is the number of DVFS or clock throttling triggered by overheating within the monitoring periods of the CPU and GPU.
[0068] Specifically, the high-performance computing cluster platform is warned according to the cooling effect response values of each risk zone cluster. The specific process is as follows: based on the cooling effect response values of each risk zone cluster and comparing them with the set cooling effect response thresholds of each risk zone cluster, if the cooling effect response value of each risk zone cluster is lower than or equal to the cooling effect response threshold of each risk zone cluster, then each risk zone cluster is continuously monitored; if the cooling effect response value of each risk zone cluster is higher than the cooling effect response threshold of each risk zone cluster, then a warning is issued.
[0069] It should be noted that issuing a warning specifically means notifying the operation and maintenance team through the monitoring platform or email / SMS with the corresponding pre-warning information, and at the same time sending the secondary alarm information to the cooling controller to automatically execute the preset and configured emergency cooling scripts and task migration strategies.
[0070] Specifically, triggering hardware compensation. The specific process is as follows: count the number of risk zone clusters whose cooling effect response values are higher than the cooling effect response thresholds of each risk zone cluster. If the number of risk zone clusters whose cooling effect response values are higher than the cooling effect response thresholds of each risk zone cluster is higher than or equal to the risk zone cluster quantity threshold, then trigger hardware compensation and enable immersion phase change cooling.
[0071] It should be noted that Figure 2 This is the schematic diagram of the logic flow of the control method for heat dissipation of the high-performance computing cluster platform of the present invention. First, through three-dimensional thermal field modeling and inputting into the LSTM + GBRT hybrid prediction model, the prediction of the temperature and heat flux density of each node is obtained, and the node-level thermal risk score is analyzed; then the system divides the nodes into three risk zone clusters of safe, warning, and hot spots according to the risk score, and intelligently schedules computing tasks in combination with the partitioning results to avoid high-risk nodes. Continuously and frequently monitor hot data such as the extreme value of the temperature rise rate, throttling events, and task throughput recovery within each risk zone cluster, and trigger multi-level cooling measures including DVFS, liquid cooling pump speed, and phase change cooling accordingly. Finally, judge the overall cooling performance through the cooling effect response value and automatically issue a warning and call for redundant hardware compensation, with the cooling effect closed-loop feedback, significantly improving the heat dissipation response speed and energy efficiency ratio, and ensuring that the high-performance computing cluster still maintains thermal stability and computing power stability under extremely high loads.
[0072] Figure 3This is a schematic diagram of the logic flow for analyzing and correcting the prediction data of each node in the present invention. The system collects the predicted values and measured values of the nodes in each monitoring cycle, calculates the deviation between the two, and then statistically counts the number of cycles with excessive deviation. When the deviation of a certain node exceeds the threshold for multiple consecutive cycles, the analysis and correction of the prediction data of this node are triggered. After the correction is completed, the results are immediately applied to the parameter adjustment of the heat dissipation control strategy to enhance the accuracy of the cooling response. If the threshold is not exceeded, continuous monitoring is continued, and both finally converge to the closed-loop link of the optimized control strategy.
[0073] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.
[0074] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in one Figure 1 process or multiple processes and / or blocks Figure 1 or multiple blocks.
[0075] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device realizes the functions specified in one Figure 1 process or multiple processes and / or blocks Figure 1 or multiple blocks.
[0076] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in one Figure 1 process or multiple processes and / or blocks Figure 1 or multiple blocks.
[0077] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made by those skilled in the art once they learn of the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present invention.
[0078] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A control method for heat dissipation of a high-performance computing cluster platform, characterized in that, Including: Construct a three-dimensional thermal field model of a high-performance computing cluster platform, obtain the sensing data of each node of the computing cluster platform, import it into the LSTM and GBRT hybrid prediction model, and obtain the average deviation value of the prediction data of each node. Analyze and correct the prediction data of each node according to the average deviation value of the prediction data of each node, and combine the sensing data of each node to obtain the thermal risk scoring index of each node. The process of obtaining the thermal risk scoring index of each node is as follows: Extract the sensing data of each node of the computing cluster platform, including comparing the total power consumption, average fan speed, and coolant flow rate with their corresponding reference values respectively, introducing the weight coefficient and the average deviation value of the prediction data of each node, and then coupling them to obtain the thermal risk scoring index of each node. The thermal risk scoring index of each node is used to evaluate the thermal risk level of each node in the high-performance computing cluster platform. Divide the hot spots of the high-performance computing cluster platform according to the thermal risk scoring index of each node to obtain each risk area cluster, and allocate computing tasks according to each risk area cluster. Monitor each node in each risk area cluster, obtain and analyze the hot spot data of each node in each risk area cluster, trigger the multi-level cooling response of each risk area cluster in each risk area cluster and monitor it, obtain the cooling effect response value of each risk area cluster, and give an early warning to the high-performance computing cluster platform according to the cooling effect response value of each risk area cluster and trigger hardware compensation.
2. The control method for heat dissipation of a high-performance computing cluster platform according to claim 1, characterized in that: The process of obtaining the average deviation value of the prediction data of each node is as follows: Preset each monitoring period, and extract the prediction data of each node in each monitoring period, including the temperature prediction value, heat flux density prediction value, and temperature rise prediction rate. Obtain the actual data of each node in each monitoring period, including the average temperature, heat flux density, and temperature rise rate, compare them with the prediction data of each node in the previous period respectively, and introduce the weight coefficient and then couple them to obtain the average deviation value of the prediction data of each node. The average deviation value of the prediction data of each node is used to evaluate the accuracy of the prediction model.
3. The control method for heat dissipation of a high-performance computing cluster platform according to claim 1, characterized in that: The process of analyzing and correcting the prediction data of each node is as follows: Count the number of monitoring periods in which the prediction data deviation value of each node is higher than or equal to the preset average deviation threshold. If the number of monitoring periods in which the prediction data deviation value of each node is higher than or equal to the preset average deviation threshold is higher than the set deviation period data quantity threshold, analyze and correct the prediction data of each node.
4. The control method for heat dissipation of a high-performance computing cluster platform according to claim 1, characterized in that: The process of dividing the hot spots of the high-performance computing cluster platform according to the thermal risk scoring index of each node to obtain each risk area cluster is as follows: Extract the thermal risk scoring index of each node, and compare it with each risk area cluster corresponding to each interval of the set thermal risk scoring index of each node to obtain each risk area cluster. Each risk area cluster includes each safety area cluster, each warning area cluster, and each hot spot area cluster.
5. The control method for heat dissipation of a high-performance computing cluster platform according to claim 4, wherein: The specific analysis process of allocating computing tasks according to each risk area cluster is as follows: After the scheduler sorts each risk area cluster according to the cluster-level risk, preferentially allocate compute-intensive or long-duration batch processing tasks to the nodes in the safety area cluster, and issue an instruction to suspend the new task scheduling instructions of each warning area cluster and hot spot area cluster.
6. The control method for heat dissipation of a high-performance computing cluster platform according to claim 4, characterized in that: Trigger multi-level cooling responses for each risk area cluster and monitor them. The specific analysis process is as follows: Perform a first-level cooling response on the computing cluster platform in each safe area cluster, and perform a second-level cooling response on the computing cluster platform in each warning area cluster and each hotspot area cluster. At the same time, continuously monitor the hotspot data of each hotspot area cluster.
7. The control method for heat dissipation of a high-performance computing cluster platform according to claim 1, characterized in that: The process of obtaining the cooling effect response value of each risk area cluster is as follows: Preset a monitoring time period. During the monitoring time period, extract the hotspot data of each node in each hotspot area cluster, including the maximum value of the temperature rise rate and the defined maximum value of the temperature rise rate, the cumulative number of heat drop frequency events and the defined cumulative number of heat drop frequency events, and the average task completion rate and the reference average task completion rate for ratio analysis respectively. After introducing the weight coefficient and the heat risk scoring index of each node, perform coupling to obtain the cooling effect response value of each risk area cluster. The cooling effect response value of each risk area cluster is used to evaluate the overall cooling effectiveness of the cooling measures on the hotspot cluster.
8. The control method for heat dissipation of a high-performance computing cluster platform according to claim 1, characterized in that: The process of warning the high-performance computing cluster platform according to the cooling effect response value of each risk area cluster is as follows: According to the cooling effect response value of each risk area cluster, compare it with the set cooling effect response threshold of each risk area cluster. If the cooling effect response value of each risk area cluster is lower than or equal to the cooling effect response threshold of each risk area cluster, continuously monitor each risk area cluster. If the cooling effect response value of each risk area cluster is higher than the cooling effect response threshold of each risk area cluster, issue a warning.
9. The control method for heat dissipation of a high-performance computing cluster platform according to claim 8, characterized in that: The process of triggering hardware compensation is as follows: Count the number of risk area clusters whose cooling effect response value is higher than the cooling effect response threshold of each risk area cluster. If the number of risk area clusters whose cooling effect response value is higher than the cooling effect response threshold of each risk area cluster is higher than or equal to the risk area cluster number threshold, trigger hardware compensation and enable immersion phase change cooling.
Citation Information
Patent Citations
A control method and system for heat dissipation of a high-performance computing cluster platform
CN113721741B
High-performance computing cluster platform heat dissipation device and control system thereof
CN117032425A
Real-time temperature rise measurement method of electromagnetic system of low-voltage apparatus based on finite element analysis
CN103970947A
Low-voltage power distribution control box and heat dissipation early warning method thereof
CN118763541A