A Temperature-Sensing-Based Data Center Server Power Consumption Regulation Method

By constructing a rack-level heat propagation topology map and reconstructing and controlling the heat propagation path, the problem of inaccurate heat propagation of server nodes in high-density data centers was solved, enabling proactive identification and optimization of the heat propagation process and improving the accuracy and stability of thermal management.

CN122219737BActive Publication Date: 2026-07-17SHANDONG HONGHE INFORMATION TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANDONG HONGHE INFORMATION TECH CO LTD
Filing Date
2026-05-21
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

In high-density data centers, existing technologies struggle to accurately identify and regulate heat propagation and accumulation between server nodes, resulting in insufficient foresight and accuracy in thermal management control. In particular, when the rate of heat accumulation and the rate of temperature change are inconsistent along the heat propagation path, single-node temperature control methods are unable to effectively identify thermal risks.

Method used

By establishing a rack-level heat propagation topology map, combining the airflow connectivity matrix and the heat flux changes of server nodes, pseudo-steady-state thermal states are identified, and server power consumption is optimized by reconstructing control and computing task scheduling through heat propagation path reconstruction.

Benefits of technology

It improves the ability to characterize the heat transfer process inside the server rack, enhances the foresight and accuracy of thermal management control, and improves the operational stability of data center server nodes in complex thermal coupling environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122219737B_ABST
    Figure CN122219737B_ABST
Patent Text Reader

Abstract

This invention discloses a temperature-sensing-based method for regulating the power consumption of data center servers, relating to the field of server thermal management technology. The method includes: establishing an airflow connectivity matrix based on the spatial relationships and airflow organization structure of data center server nodes within a rack; and constructing a rack-level heat propagation topology map based on the airflow direction and weight to determine the upstream heat-affected set and downstream heat-affected set of each data center server node; and collecting processor power change sequences, processor temperature change sequences, and server intake air temperature change sequences for each data center server node based on the rack-level heat propagation topology map. By establishing a rack-level heat propagation topology map, this invention incorporates the airflow organization relationships, upstream heat-affected relationships, and downstream heat-affected relationships between data center server nodes within a rack into a unified heat propagation structure description, thus enabling server power consumption regulation to move beyond being limited to single-node temperature information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of server thermal management technology, specifically to a method for adjusting the power consumption of a data center server based on temperature sensing. Background Technology

[0002] As data centers evolve towards high-density deployment and high-computing-power operation, a single rack typically integrates a large number of data center server nodes. These server nodes are significantly coupled in terms of spatial location, air intake and exhaust relationships, and heat propagation paths. In this operating environment, the heat generated by the server nodes not only affects their own temperature but also continuously impacts the air intake conditions and heat dissipation capabilities of adjacent nodes through the airflow organization within the rack. Therefore, how to achieve server power consumption regulation under rack-level heat propagation constraints has become a crucial technical issue in data center thermal management.

[0003] Existing data center server power consumption regulation methods are usually based on the server node's own temperature information, processor power information, or load status information for control. For example, frequency reduction, power limiting, or task migration are performed after the temperature reaches a preset threshold. Such methods have good engineering applicability in single-node thermal control and can meet the temperature control requirements in general scenarios. However, in high-density rack deployment scenarios, there is a significant heat propagation relationship between server nodes. Simply relying on the node's own temperature for control makes it difficult to accurately reflect the dynamic process of heat propagation and accumulation along the airflow path inside the rack.

[0004] Furthermore, in actual operation, some data center server nodes may experience a discrepancy between the rate of heat accumulation and the rate of temperature change due to factors such as thermal inertia, airflow disturbance, and path congestion. That is, heat has been continuously accumulating along the heat propagation path, while the node temperature remains relatively stable. In such cases, if the power consumption adjustment method based on the single node temperature threshold is still used, it is difficult to effectively identify and adjust the relevant nodes before the thermal risk spreads. Therefore, it is necessary to propose a data center server power consumption control method that can comprehensively adjust the power consumption by combining the rack-level heat propagation topology, node-level heat flux changes, and path-level heat accumulation status. Summary of the Invention

[0005] The purpose of this invention is to provide a temperature-sensing-based method for regulating the power consumption of data center servers, in order to solve the problems mentioned in the background art.

[0006] This invention can be achieved through the following technical solution: a method for regulating the power consumption of a data center server based on temperature sensing, comprising:

[0007] Step 1: Establish an airflow connectivity matrix based on the spatial location relationship and airflow organization structure of the data center server nodes in the rack, and construct a rack-level heat propagation topology map based on the airflow connectivity matrix to determine the upstream heat impact set and downstream heat-affected set of each data center server node;

[0008] Step 2: Based on the heat propagation topology map, collect the processor power change sequence, processor temperature change sequence, and server intake air temperature change sequence of each data center server node. Calculate the node-level heat flux change value based on the processor power change sequence and processor temperature change sequence, and map the node-level heat flux change value to the corresponding heat propagation path according to the heat propagation topology map to form a path-level heat flux change sequence.

[0009] Step 3: Calculate the heat accumulation rate of each heat propagation path based on the path-level heat flux change sequence, and compare the heat accumulation rate with the temperature change rate of the corresponding data center server node. When the heat accumulation rate is consistently higher than the temperature change rate of the corresponding data center server node, it is determined that the data center server node is in a pseudo-steady-state thermal state.

[0010] Step 4: Based on the total heat dissipation capacity of the computer cabinet according to the heat propagation topology diagram, and combined with the propagation relationship of each data center server node in the heat propagation path, determine the budget value of the heat dissipation capacity of each data center server node.

[0011] Step 5: Perform heat propagation path reconstruction control on the data center server nodes in a pseudo-steady-state thermal state based on the heat dissipation capacity budget value. This involves adjusting the distribution of computing tasks among different processing cores and scheduling new computing tasks to data center server nodes with higher heat dissipation capacity budget values.

[0012] A further technical improvement of the present invention is that the step of establishing the airflow connectivity matrix in step one includes:

[0013] Based on the spatial relationship between the exhaust surface of each data center server node in the rack and the intake surface of the adjacent data center server node, as well as the spatial projection range of the structural obstruction between them, the obstruction ratio of the exhaust propagation path is calculated. When the obstruction ratio exceeds a preset threshold, the airflow connectivity weight between the corresponding data center server nodes is reduced.

[0014] Based on the airflow return paths formed by the side channels, top channels, and structural gaps of the rack, calculate the connection distance between each airflow return path and the air intake surface of the target data center server node, and establish the indirect connection relationship between the corresponding data center server nodes in the airflow connection relationship matrix when the connection distance is less than the preset return threshold.

[0015] Based on the installation height of each data center server node and the angle between the server exhaust direction and the upward direction of the hot airflow, the airflow propagation direction between adjacent data center server nodes is determined. When the airflow propagation direction points only to a single data center server node, the corresponding connection relationship is determined as a unidirectional connection relationship. When the airflow propagation direction points to two data center server nodes at the same time, the corresponding connection relationship is determined as a bidirectional connection relationship.

[0016] Based on the results of the exhaust propagation path obstruction ratio correction, the airflow return path connectivity, and the direction determination results of unidirectional and bidirectional connectivity, a final airflow connectivity matrix is ​​generated, and a rack-level heat propagation topology map is constructed based on the final airflow connectivity matrix.

[0017] A further technical improvement of this invention lies in: when constructing a rack-level heat propagation topology map based on an airflow connectivity matrix, determining the heat propagation edge weights for the node connectivity relationships in the airflow connectivity matrix, the steps of which include:

[0018] Within a continuous time window, the server operating parameter change data of the source data center server node and the target data center server node are collected, and the heat propagation correlation is identified based on the time propagation order between the exhaust temperature change of the source data center server node and the intake temperature change of the target data center server node. When the time propagation order meets the preset propagation coupling trigger threshold, the corresponding node connectivity is determined as the heat propagation correlation edge.

[0019] A heat propagation association sequence is constructed based on the occurrence of heat propagation association edges in multiple consecutive time windows. When the number of consecutive occurrences of the heat propagation association sequence reaches a preset propagation stability judgment threshold, the connectivity relationship of the corresponding nodes is determined as a stable heat propagation path.

[0020] The path propagation driving value is determined based on the change range of server operating parameters and the propagation distance between adjacent nodes in the stable heat propagation path, and the path propagation impedance value is determined based on the degree of deviation of airflow propagation direction between adjacent nodes and the airflow pressure difference inside the cabinet.

[0021] The path propagation coefficient is determined based on the nonlinear propagation relationship between the path propagation driving value and the path propagation impedance value. The heat propagation edge weight of the corresponding node connectivity relationship is determined based on the path propagation coefficient and the duration of the stable heat propagation path. The airflow connectivity matrix is ​​then weighted to generate a rack-level heat propagation topology map.

[0022] A further technical improvement of the present invention is that, when determining the weight of the heat propagation edge, a heat propagation competition determination step is also included, comprising:

[0023] Identify multiple source data center server nodes corresponding to the air intake area of ​​the same target data center server node. When the exhaust propagation direction of multiple source data center server nodes points to the air intake area at the same time, determine that a heat propagation competition relationship is formed between the corresponding nodes.

[0024] Within a continuous time window, the propagation time of the exhaust air temperature change of each source data center server node to the air intake area of ​​the target data center server node is collected, and the heat propagation competition order is determined according to the order of propagation time.

[0025] When the source data center server node with the earliest propagation time meets the preset propagation channel occupancy threshold, the node connectivity relationship corresponding to the source data center server node is determined as the priority hot propagation path, and the path propagation drive value of the other node connectivity relationships is reduced.

[0026] The path propagation coefficient is corrected based on the competition between the priority heat propagation path and the connectivity of the remaining nodes, and the corrected path propagation coefficient is then substituted back into the heat propagation edge weight calculation process to determine the final heat propagation edge weight.

[0027] A further technical improvement of the present invention is that a thermal propagation resonance determination step is performed during the process of determining the thermal propagation edge weights, including:

[0028] Within a continuous time window, server operating parameter change data of multiple source data center server nodes in the upstream thermal impact set of the same target data center server node are collected, and the exhaust heat fluctuation frequency of each source data center server node is identified based on the periodic change characteristics of exhaust temperature change within the continuous time window.

[0029] The thermal propagation frequency coupling relationship is determined based on the frequency difference between the exhaust thermal fluctuation frequencies of server nodes in each source data center. When the frequency difference is less than the preset propagation resonance judgment threshold, the thermal propagation relationship between the corresponding nodes is determined as a thermal propagation resonance relationship.

[0030] The thermal propagation resonance path is identified based on the spatial distribution of the connectivity relationships of multiple nodes forming a thermal propagation resonance relationship in the thermal propagation topology map, and the resonance propagation intensity is calculated based on the number of nodes forming a resonance relationship and the magnitude of the exhaust temperature change.

[0031] When the resonance propagation intensity exceeds the preset resonance intensity trigger threshold, the path propagation coefficient of the corresponding node connectivity is amplified and corrected, and the heat propagation edge weight is re-determined based on the corrected path propagation coefficient.

[0032] A further technical improvement of the present invention is that: in the process of calculating the node-level heat flux change value in step two, the following steps are included:

[0033] Collect the processor power change sequence, processor temperature change sequence, server intake air temperature change sequence, and corresponding heat propagation edge weights of data center server nodes within a continuous time window.

[0034] The node thermal inertia hysteresis value is determined based on the time difference between the processor power change sequence and the processor temperature change sequence, and the node thermal inertia hysteresis value is corrected for air intake disturbance based on the server intake air temperature change sequence.

[0035] Based on the rack-level heat propagation topology, the upstream heat impact set and downstream heat-affected set of the data center server node are extracted, and the node heat propagation gradient value is determined by combining the corresponding heat propagation edge weights and the temperature difference between adjacent data center server nodes.

[0036] The equivalent thermal storage capacity of a node is determined based on the node thermal inertia hysteresis value, the air inlet disturbance correction result, and the node thermal propagation gradient value.

[0037] The node-level heat flux change value of the data center server node is determined based on the power change amplitude corresponding to the processor power change sequence, the node heat propagation gradient value, and the node equivalent heat storage.

[0038] A further technical improvement of the present invention is that: in the process of mapping node-level heat flux change values ​​to heat propagation paths to form path-level heat flux change sequences in step two, the following is included:

[0039] In the rack-level heat propagation topology map, identify multiple downstream heat-affected cluster nodes corresponding to the data center server nodes, and determine the candidate heat propagation paths connected to the data center server nodes based on the heat propagation edge weights corresponding to the connectivity of each node.

[0040] The path propagation driving value is determined based on the heat propagation edge weights of the connectivity relationships of each node in the candidate heat propagation path and the temperature difference between adjacent data center server nodes. The path propagation impedance value is determined based on the degree of airflow propagation direction offset, path propagation distance, and airflow pressure difference inside the rack in the candidate heat propagation path.

[0041] The available heat transfer capacity of each candidate heat transfer path is determined based on the historical heat load changes of each candidate heat transfer path in the rack-level heat transfer topology diagram and the corresponding heat dissipation capacity budget value.

[0042] The path propagation efficiency is determined based on the path propagation driving value, the path propagation impedance value, and the available heat propagation capacity of the path. The node-level heat flux change value is then proportionally allocated to each candidate heat propagation path according to the path propagation efficiency to form a path-level heat flux change sequence.

[0043] A further technical improvement of the present invention is that, when determining in step three that the data center server node is in a pseudo-steady-state thermal state, the following steps are included:

[0044] The duration for which the thermal accumulation rate of a data center server node is continuously higher than the corresponding temperature change rate is recorded within a continuous time window. The node thermal inertia compensation time is determined based on the hysteresis relationship between the processor temperature change sequence and the server air intake temperature change sequence, thereby obtaining the dynamic pseudo steady state determination time threshold.

[0045] Based on the rack-level heat propagation topology, the upstream thermal impact set of data center server nodes is extracted, and the number of data center server nodes in the upstream thermal impact set that meet the condition that the heat accumulation rate is higher than the temperature change rate within the same continuous time window is counted.

[0046] When the duration exceeds the dynamic pseudo-steady state determination time threshold and the number of data center server nodes that meet the conditions reaches the preset topology consistency determination threshold, the data center server node is determined to be in a pseudo-steady state hot state.

[0047] A further technical improvement of the present invention is that, in step four, when calculating the budgeted heat dissipation capacity of each data center server node, the following is included:

[0048] Based on the rack-level heat propagation topology map, the number of heat propagation paths corresponding to each data center server node and the path-level heat flux change sequence of each path are counted to determine the heat propagation contribution of each data center server node.

[0049] Based on the path-level heat flux change sequence, the historical heat load changes of each heat propagation path are statistically analyzed, and the path congestion degree of the corresponding heat propagation path is determined.

[0050] The total heat dissipation capacity of the rack is allocated based on the heat propagation contribution of each data center server node and the path congestion of the corresponding heat propagation path, thereby obtaining the heat dissipation capacity budget value of each data center server node.

[0051] A further technical improvement of the present invention is that, during the execution of heat propagation path reconstruction control in step five, the following is included:

[0052] Based on the rack-level heat propagation topology map, identify the heat propagation path of the data center server node in a pseudo-steady-state thermal state, and determine the candidate computing node that is different from the heat propagation path;

[0053] The thermal inertia buffering capacity of each candidate computing node is determined based on the processor temperature change sequence of the corresponding processing core of each candidate computing node and the corresponding node-level heat flux change value.

[0054] New computing tasks are prioritized and scheduled to candidate computing nodes that are not part of the heat propagation path and have high thermal inertia buffering capacity, so as to achieve thermal migration control between processing cores.

[0055] Compared with the prior art, the present invention has the following beneficial effects:

[0056] This invention establishes a rack-level heat propagation topology diagram, incorporating the airflow organization relationship, upstream thermal influence relationship, and downstream heating relationship between data center server nodes within the rack into a unified heat propagation structure description. This allows server power consumption adjustment to no longer be limited to single-node temperature information, but to conduct path-level thermal state analysis under rack-level heat propagation constraints, thereby improving the ability to characterize the heat propagation process inside the rack.

[0057] Furthermore, this invention calculates the heat accumulation rate of the heat propagation path based on node-level heat flux change values ​​and path-level heat flux change sequences, and identifies the pseudo-steady-state thermal state of data center server nodes by combining the temperature change rate. This enables the identification of an operating state where heat is continuously accumulating before the node temperature has increased significantly. In this way, the ability of power consumption regulation to perceive early thermal risks can be enhanced, and the foresight and accuracy of thermal management control can be improved.

[0058] On the other hand, the present invention also combines the propagation relationship in the heat propagation path and the total heat dissipation capacity of the rack to determine the heat dissipation capacity budget value of each data center server node, and performs heat propagation path reconstruction control and new computing task scheduling accordingly, so that the power consumption adjustment process can coordinate with the heat propagation path inside the rack, the allocation of heat dissipation resources and the heat migration of the processing core; thereby, it is beneficial to improve the operational stability of data center server nodes in complex thermal coupling environment and improve the rack-level thermal management control effect. Attached Figure Description

[0059] To facilitate understanding by those skilled in the art, the present invention will be further described below with reference to the accompanying drawings.

[0060] Figure 1 This is a schematic diagram of the method logic of the present invention. Detailed Implementation

[0061] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided.

[0062] Please see Figure 1 As shown, this invention provides a method for regulating the power consumption of a data center server based on temperature sensing, comprising:

[0063] Step 1: Establish an airflow connectivity matrix based on the spatial relationships and airflow organization of the data center server nodes within the rack. Then, construct a rack-level heat propagation topology map based on this matrix to determine the upstream heat-affected set and downstream heat-affected set for each data center server node. This step begins with the overall spatial structure of the rack, structurally representing the airflow propagation relationships between the originally dispersed data center server nodes. By establishing the airflow connectivity matrix, it becomes clear which data center server nodes have direct or indirect airflow influence relationships. Further constructing the rack-level heat propagation topology map transforms these airflow influence relationships into heat propagation relationships, thereby determining the upstream heat-affected set and downstream heat-affected set for each data center server node. This provides a basic topological framework for subsequent heat propagation analysis, allowing subsequent steps to move beyond the limitations of individual data center server node temperature or power changes and instead analyze the direction and range of heat transfer under the constraints of rack-level heat propagation relationships.

[0064] Specifically, step one, establishing the airflow connectivity matrix, includes: calculating the obstruction ratio of the exhaust propagation path based on the spatial relationship between the exhaust surfaces of each data center server node and the intake surfaces of adjacent data center server nodes, as well as the spatial projection range of structural obstructions between them; reducing the airflow connectivity weight between the corresponding data center server nodes when the obstruction ratio exceeds a preset threshold; calculating the connection distance between each airflow return path and the intake surface of the target data center server node based on the airflow return paths formed by the side passages, top passages, and structural gaps of the rack; and establishing the connection between the corresponding data center server nodes in the airflow connectivity matrix when the connection distance is less than a preset return threshold. The indirect connectivity between data center server nodes is determined based on the installation height of each data center server node and the angle between the server exhaust direction and the upward direction of the hot airflow. When the airflow direction points only to a single data center server node, the corresponding connectivity is determined as a unidirectional connectivity; when the airflow direction points to two data center server nodes simultaneously, the corresponding connectivity is determined as a bidirectional connectivity. A final airflow connectivity matrix is ​​generated based on the exhaust path obstruction ratio correction result, the airflow return path connectivity, and the direction determination results of unidirectional and bidirectional connectivity. A rack-level heat propagation topology map is then constructed based on this final airflow connectivity matrix. In this embodiment, a 42U rack is selected, with 20 data center server nodes arranged within it. All data center server nodes are numbered, and the installation height, forward airflow position, and rear exhaust airflow position of each data center server node are recorded. Then, a line is established connecting the center point of the exhaust surface of any source data center server node to the center point of the intake surface of the adjacent target data center server node. The direction of the line is defined as the exhaust propagation direction. Centered on this line, a rectangular projection area with a width of 120mm and a height equal to the height of the corresponding node's intake surface is generated on the plane containing the exhaust propagation direction. The projected areas of cable bundles, cable management rack edges, slide rail edges, and baffle edges falling within this rectangular projection area are counted. The occlusion ratio is obtained by dividing the projected area by the total area of ​​the rectangular projection area. In this embodiment, the preset threshold is 0.30. When the occlusion ratio is greater than 0.30, the airflow connectivity weight between the corresponding data center server nodes is reduced from 1.00 to 0.65. Subsequently, starting from the exhaust surface of the source data center server node, the paths that can reach the intake surface of the target data center server node are searched along the side channel of the rack, the top channel of the rack, and the structural gaps of the rack. The length of the airflow return path is obtained by summing the lengths of the broken line segments, and the distance from the end of the airflow return path to the nearest point of the intake surface of the target data center server node is taken as the connection distance. In this embodiment, the preset return threshold is 150mm. When the connection distance is less than 150mm, an indirect connection relationship is added to the airflow connection relationship matrix.Next, the airflow propagation direction is determined based on the installation height of each data center server node and the angle between the server exhaust direction and the upward direction of the hot airflow: when the angle is less than 25°, the corresponding connectivity is determined as a unidirectional connectivity; when the angle is greater than or equal to 25° and less than 70°, the corresponding connectivity is determined as a bidirectional connectivity. After completing the above processing, the results of the exhaust propagation path obstruction ratio correction, the airflow return path connectivity, and the direction determination results of unidirectional and bidirectional connectivity are written into the airflow connectivity matrix. Then, a rack-level heat propagation topology is constructed with the data center server nodes as vertices and each connectivity as an edge, and the upstream heat impact set and downstream heat-affected set of each data center server node are extracted accordingly.

[0065] Based on the aforementioned final airflow connectivity matrix and rack-level heat propagation topology map, when constructing the rack-level heat propagation topology map based on the airflow connectivity matrix, the steps for determining the heat propagation edge weights for the node connectivity relationships in the airflow connectivity matrix include: collecting server operating parameter change data of the source data center server node and the target data center server node within a continuous time window, and identifying heat propagation correlations based on the time propagation sequence between the exhaust air temperature change of the source data center server node and the intake air temperature change of the target data center server node; when the time propagation sequence meets a preset propagation coupling trigger threshold, the corresponding node connectivity relationship is determined as a heat propagation correlation edge; based on the heat propagation correlation edges in multiple continuous time windows... The occurrence of heat propagation correlation sequences is used to construct a stable heat propagation path. When the number of consecutive occurrences of the heat propagation correlation sequence reaches a preset stability threshold, the corresponding node connectivity is determined as a stable heat propagation path. The path propagation driving value is determined based on the change amplitude of server operating parameters and the propagation distance between adjacent nodes in the stable heat propagation path, and the path propagation impedance value is determined based on the degree of airflow propagation direction offset between adjacent nodes and the airflow pressure difference inside the rack. The path propagation coefficient is determined based on the nonlinear propagation relationship between the path propagation driving value and the path propagation impedance value, and the heat propagation edge weight of the corresponding node connectivity is determined based on the path propagation coefficient and the duration of the stable heat propagation path, so as to weight the airflow connectivity matrix and generate a rack-level heat propagation topology map. In this embodiment, the continuous time window is 10 seconds. Within each continuous time window, the exhaust temperature change data of the source data center server node, the intake temperature change data of the target data center server node, and the corresponding server operating parameter change data are collected synchronously. The server operating parameter change data includes processor power change data, processor temperature change data, and fan speed change data. The propagation start time is defined as the moment when the exhaust temperature of the source data center server node first rises by 2.0°C relative to the start time of this window, and the arrival propagation time is defined as the moment when the intake temperature of the target data center server node first rises by 0.8°C. The time difference between the two is used as the time propagation order determination value. In this embodiment, the preset propagation coupling trigger threshold is set to 5 seconds. When the time propagation order determination value is less than 5 seconds, the corresponding node connectivity is determined as a heat propagation associated edge. The occurrence of the same node connectivity is then counted within 6 consecutive time windows. When the same heat propagation associated edge appears at least 4 times in 6 consecutive time windows, it is considered to meet the preset propagation stability determination threshold, and the node connectivity is determined as a stable heat propagation path.Based on this, the variation amplitude of server operating parameter changes in a stable heat propagation path is defined as the weighted sum of the peak-to-valley difference of processor power change and the peak-to-valley difference of exhaust temperature change within the same continuous time window, divided by the propagation distance between adjacent nodes to obtain the path propagation driving value. Simultaneously, the product of the airflow propagation direction offset between adjacent nodes and the internal airflow pressure difference of the cabinet is taken as the path propagation impedance value. In this embodiment, the pressure difference across the cabinet is maintained at 18 Pa. Next, the path propagation coefficient is determined based on the nonlinear propagation relationship between the path propagation driving value and the path propagation impedance value. Specifically, the path propagation coefficient is calculated as the sum of the path propagation driving value divided by 1 and the path propagation impedance value. This coefficient is then weighted in conjunction with the duration of the stable heat propagation path to obtain the heat propagation edge weights for the corresponding node connectivity relationships. These weights are then used to assign weights to the airflow connectivity matrix, resulting in a weighted cabinet-level heat propagation topology.

[0066] Based on the aforementioned heat propagation edge weights, the determination of heat propagation edge weights also includes a heat propagation competition determination step, which includes: identifying multiple source data center server nodes corresponding to the air intake area of ​​the same target data center server node; when the exhaust propagation direction of multiple source data center server nodes simultaneously points to the air intake area, it is determined that a heat propagation competition relationship has formed between the corresponding nodes; collecting the propagation time of the exhaust temperature change of each source data center server node to the air intake area of ​​the target data center server node within a continuous time window, and determining the heat propagation competition order according to the order of propagation time; when the source data center server node with the earliest propagation time meets the preset propagation channel occupancy threshold, the node connectivity relationship corresponding to that source data center server node is determined as the priority heat propagation path, and the path propagation driving value of the other node connectivity relationships is reduced; the path propagation coefficient is corrected according to the competition result between the priority heat propagation path and the other node connectivity relationships, and the corrected path propagation coefficient is re-substituted into the heat propagation edge weight calculation process to determine the final heat propagation edge weight. Specifically, a coverage analysis is performed on the air intake area of ​​the same target data center server node; when the exhaust propagation direction of at least two source data center server nodes simultaneously falls into the air intake area, it is determined that a heat propagation competition relationship has formed between these nodes. Subsequently, the propagation time of the exhaust temperature change of each source data center server node to the intake air area of ​​the target data center server node is measured within a continuous time window, and the heat propagation competition order is obtained by sorting the propagation times from smallest to largest. In this embodiment, when the lead time of the source data center server node with the earliest propagation time is greater than 2 seconds relative to the second earliest node, and the exhaust temperature change of the source node is not less than 3.5℃, it is considered to meet the preset propagation channel occupancy threshold, and the corresponding node connectivity is determined as the priority heat propagation path. Under this condition, the path propagation drive value of the remaining node connectivity is corrected to 0.70 of the original value, and the path propagation coefficient is corrected according to the competition result between the priority heat propagation path and the remaining node connectivity. The corrected path propagation coefficient is then substituted back into the heat propagation edge weight calculation process to obtain the final heat propagation edge weight after considering heat propagation competition.

[0067] After completing the heat propagation competition determination step, the heat propagation resonance determination step is performed during the process of determining the heat propagation edge weights. This includes: collecting server operating parameter change data of multiple source data center server nodes in the upstream heat impact set of the same target data center server node within a continuous time window, and identifying the exhaust heat fluctuation frequency of each source data center server node based on the periodic change characteristics of exhaust temperature change within the continuous time window; determining the heat propagation frequency coupling relationship based on the frequency difference between the exhaust heat fluctuation frequencies of each source data center server node, and determining the heat propagation relationship between the corresponding nodes as a heat propagation resonance relationship when the frequency difference is less than the preset propagation resonance determination threshold; identifying the heat propagation resonance path based on the spatial distribution of the connectivity relationships of multiple nodes forming the heat propagation resonance relationship in the heat propagation topology map, and calculating the resonance propagation intensity based on the number of nodes forming the resonance relationship and the amplitude of their exhaust temperature change; when the resonance propagation intensity exceeds the preset resonance intensity trigger threshold, amplifying and correcting the path propagation coefficient of the corresponding node connectivity relationship, and re-determining the heat propagation edge weights based on the corrected path propagation coefficient. In this embodiment, multiple source data center server nodes are selected from the upstream thermal impact set of the same target data center server node. Server operating parameter changes are continuously collected for 60 seconds, and exhaust temperature change curves are recorded at 1-second sampling intervals. For each source data center server node, the periodic peak interval method is used to identify its exhaust thermal fluctuation frequency. When the difference in exhaust thermal fluctuation frequencies between two source data center server nodes is less than 0.02Hz, it is considered to meet the preset propagation resonance threshold, and the thermal propagation relationship between the corresponding nodes is determined as a thermal propagation resonance relationship. Based on this, the thermal propagation resonance path is identified according to the spatial distribution of the connectivity relationships of multiple nodes forming the thermal propagation resonance relationship in the thermal propagation topology map. The resonance propagation intensity is then calculated by multiplying the number of nodes forming the resonance relationship by the average value of the exhaust temperature change amplitude of these nodes. When the resonance propagation intensity exceeds the preset resonance intensity trigger threshold of 5.5℃ / min, the path propagation coefficient of the corresponding node connectivity relationship is amplified and corrected by 1.15 times, and the thermal propagation edge weights are re-determined based on the corrected path propagation coefficient. After the above four stages, a final rack-level thermal propagation topology is obtained, which simultaneously considers shading propagation, backflow propagation, propagation direction, thermal propagation correlation, stable thermal propagation path, thermal propagation competition relationship, and thermal propagation resonance relationship. This final rack-level thermal propagation topology is then directly associated with the node-level heat flux change value calculation, path-level heat flux change sequence formation, and pseudo-steady-state thermal state determination process in subsequent embodiments.

[0068] Step 2: Based on the heat propagation topology map, collect the processor power change sequence, processor temperature change sequence, and server intake air temperature change sequence for each data center server node. Calculate the node-level heat flux change value based on the processor power change sequence and processor temperature change sequence, and map the node-level heat flux change value to the corresponding heat propagation path according to the heat propagation topology map to form a path-level heat flux change sequence. This step further transforms the operational status changes of a single data center server node into a heat change representation that reflects the heat propagation process. The processor power change sequence and processor temperature change sequence reflect the changes in the node's own heat source, while the server intake air temperature change sequence reflects the changes in the node's heat dissipation environment. Calculating the node-level heat flux change value based on these provides parameters that better characterize the node's heat change state than simple temperature values. Subsequently, the node-level heat flux change value is mapped to the corresponding heat propagation path according to the heat propagation topology map to form a path-level heat flux change sequence. This further transforms the heat changes that were originally at the node level into propagation changes at the path level, providing a basis for subsequent analysis of the heat accumulation process along the heat propagation path.

[0069] Step two, in calculating the nodal-level heat flux change, includes:

[0070] Collect the processor power change sequence, processor temperature change sequence, server intake air temperature change sequence, and corresponding heat propagation edge weights of data center server nodes within a continuous time window.

[0071] The node thermal inertia hysteresis value is determined based on the time difference between the processor power change sequence and the processor temperature change sequence, and the node thermal inertia hysteresis value is corrected for air intake disturbance based on the server intake air temperature change sequence.

[0072] Based on the rack-level heat propagation topology, the upstream heat impact set and downstream heat-affected set of the data center server node are extracted, and the node heat propagation gradient value is determined by combining the corresponding heat propagation edge weights and the temperature difference between adjacent data center server nodes.

[0073] The equivalent thermal storage capacity of a node is determined based on the node thermal inertia hysteresis value, the air inlet disturbance correction result, and the node thermal propagation gradient value.

[0074] The node-level heat flux change value of the data center server node is determined based on the power change amplitude corresponding to the processor power change sequence, the node heat propagation gradient value, and the node equivalent heat storage.

[0075] Specifically, based on the rack-level heat propagation topology map constructed according to the aforementioned implementation method, continuous operating parameters are collected for each data center server node to form the basic data required for subsequent calculation of node-level heat flux change values. In this embodiment, processor power change sequence, processor temperature change sequence, server intake air temperature change sequence, and corresponding heat propagation edge weights are collected for each data center server node. The sampling period is 1 second, and the continuous time window is 20 seconds, meaning 20 sets of sampling data are recorded within each continuous time window. The processor power change sequence is obtained by reading the instantaneous power value of the node's processor power interface; the processor temperature change sequence is obtained by reading the processor package temperature sensor; the server intake air temperature change sequence is obtained from temperature sampling points located in the middle of the server's intake surface; and the heat propagation edge weights are directly called from the edge weights in the weighted rack-level heat propagation topology map determined in the previous embodiment. After completing the above data collection, the processor power change sequence and processor temperature change sequence within the same continuous time window are time-aligned. The moment when the processor power first changes by more than 5W is recorded as the power change start moment, and the moment when the processor temperature first changes by more than 0.5℃ is recorded as the temperature change start moment. The time difference between the two is recorded as the change time difference. When the change time difference is greater than 2s, it is considered that there is a significant thermal inertia response, and this time difference is determined as the node thermal inertia lag value. Subsequently, the node thermal inertia lag value is corrected according to the server intake air temperature change sequence within the same continuous time window. Specifically, if the maximum fluctuation of the server intake air temperature within the continuous time window is greater than 1.2℃, the original node thermal inertia lag value is multiplied by 1.1 as the intake air disturbance correction result; if the maximum fluctuation of the server intake air temperature is not greater than 1.2℃, the original node thermal inertia lag value remains unchanged.

[0076] After obtaining the node thermal inertia hysteresis value and its intake disturbance correction result, the upstream thermal influence set and downstream heat-affected set of the current data center server node are extracted based on the established rack-level heat propagation topology map, and the node heat propagation gradient value is determined on this basis. Specifically, for any target data center server node, firstly, the adjacent data center server nodes in all its upstream thermal influence sets are read, and the temperature difference between these adjacent data center server nodes and the target data center server node is calculated; then, the temperature difference of each adjacent node is multiplied by the corresponding heat propagation edge weight to obtain a weighted temperature difference value, and the sum of all weighted temperature difference values ​​is divided by the number of adjacent nodes to obtain the node heat propagation gradient value of the current target data center server node. In this embodiment, when the temperature difference between a certain adjacent data center server node and the current target data center server node is greater than 3.0℃ and the corresponding heat propagation edge weight is greater than 0.7, the weighted temperature difference value corresponding to the adjacent node is counted as 1.2 times to increase the influence of strong heat propagation coupling paths on the node heat propagation gradient value. Subsequently, the equivalent heat storage capacity of the node is determined based on the obtained node thermal inertia hysteresis value, the inlet perturbation correction result, and the node heat propagation gradient value. In this embodiment, the node thermal inertia hysteresis value after inlet perturbation correction is multiplied by the node heat propagation gradient value to obtain the initial heat storage capacity coefficient; then, this initial heat storage capacity coefficient is multiplied by the average temperature rise value of the processor temperature change sequence within the current continuous time window to obtain the node equivalent heat storage capacity of the data center server node. If the average temperature rise value is less than 0.8℃, the node equivalent heat storage capacity is corrected by 0.9 times to avoid excessive amplification of subsequent node-level heat flux changes caused by low fluctuation windows.

[0077] After obtaining the equivalent heat storage of the nodes, the node-level heat flux change value of the data center server nodes is further calculated. Specifically, the peak-to-valley difference of the processor power change sequence is first calculated within the current continuous time window, and this peak-to-valley difference is determined as the power change amplitude; then, the power change amplitude is multiplied by the aforementioned node heat propagation gradient value to obtain the node heat propagation driving force; subsequently, the node heat propagation driving force is divided by "1 plus the node equivalent heat storage" to obtain the node-level heat flux change value corresponding to the current continuous time window. In this embodiment, to avoid the influence of extreme fluctuation data on the calculation results, when the power change amplitude is greater than 35W, it is first truncated to 35W before participating in the calculation; when the node equivalent heat storage is less than 0.5, it is included as 0.5 to maintain calculation stability. After completing the calculation of the node-level heat flux change value of all data center server nodes, the main step of "mapping the node-level heat flux change value to the corresponding heat propagation path to form a path-level heat flux change sequence" is executed. Specifically, for each data center server node, all its corresponding downstream heat propagation paths are read, and the path mapping ratio is calculated based on the heat propagation edge weights on each path and the corresponding propagation distance. Then, the node-level heat flux change value of the data center server node is distributed to each corresponding heat propagation path according to the path mapping ratio, thereby obtaining the path heat flux components of the node for each path. Finally, the path heat flux components from different data center server nodes on the same heat propagation path are accumulated and arranged in order of continuous time windows to form the path-level heat flux change sequence of the heat propagation path.

[0078] Finally, to ensure the connection between this embodiment and the implementation methods described in the subsequent specification, the node-level heat flux change values ​​and path-level heat flux change sequences obtained in this embodiment are directly used as input data for the subsequent step of "calculating the heat accumulation rate of each heat propagation path based on the path-level heat flux change sequence and comparing the heat accumulation rate with the temperature change rate of the corresponding data center server node." Specifically, the node-level heat flux change value characterizes the change in heat output of each data center server node within the current continuous time window, and the path-level heat flux change sequence characterizes the process of heat propagation and accumulation along each heat propagation path. Through the above implementation steps, the superordinate expression in the claims, namely, "calculating node-level heat flux change values ​​based on processor power change sequences and processor temperature change sequences, and mapping node-level heat flux change values ​​to corresponding heat propagation paths to form path-level heat flux change sequences based on the heat propagation topology map," has been specifically translated into executable work steps, including data acquisition, time alignment, thermal inertia hysteresis identification, air intake disturbance correction, heat propagation gradient calculation, determination of node equivalent heat storage, calculation of node-level heat flux change values, and generation of path-level heat flux change sequences. This provides a foundation for the calculation of heat accumulation rate, determination of pseudo-steady-state thermal state, and allocation of heat dissipation capacity budget values ​​in subsequent embodiments.

[0079] In step two, mapping node-level heat flux changes to heat propagation paths to form path-level heat flux change sequences includes:

[0080] In the rack-level heat propagation topology map, identify multiple downstream heat-affected cluster nodes corresponding to the data center server nodes, and determine the candidate heat propagation paths connected to the data center server nodes based on the heat propagation edge weights corresponding to the connectivity of each node.

[0081] The path propagation driving value is determined based on the heat propagation edge weights of the connectivity relationships of each node in the candidate heat propagation path and the temperature difference between adjacent data center server nodes. The path propagation impedance value is determined based on the degree of airflow propagation direction offset, path propagation distance, and airflow pressure difference inside the rack in the candidate heat propagation path.

[0082] The available heat transfer capacity of each candidate heat transfer path is determined based on the historical heat load changes of each candidate heat transfer path in the rack-level heat transfer topology diagram and the corresponding heat dissipation capacity budget value.

[0083] The path propagation efficiency is determined based on the path propagation driving value, the path propagation impedance value, and the available heat propagation capacity of the path. The node-level heat flux change value is then proportionally allocated to each candidate heat propagation path according to the path propagation efficiency to form a path-level heat flux change sequence.

[0084] After obtaining the node-level heat flux change values ​​of each data center server node in the aforementioned embodiments, starting from the outgoing edges of the current target data center server node in the rack-level heat propagation topology graph, multiple downstream heat-affected set nodes corresponding to the data center server node are identified. Specifically, all downstream connected edges of the target data center server node in the rack-level heat propagation topology graph are read, and node connections with heat propagation edge weights greater than 0.20 are selected as valid propagation edges. Subsequently, based on the valid propagation edges, node connections are searched downwards along the propagation direction in the graph for no more than two hops, and the connection links from the current data center server node to the corresponding downstream heat-affected set nodes are determined as candidate heat propagation paths. In this embodiment, if the heat propagation edge weight of any node connection in a candidate heat propagation path is less than 0.10, the path is considered to have too weak a propagation capability and is not included in the candidate range. After the selection is completed, the number of nodes, the number of edges, and the correspondence between the first and last nodes are recorded for each candidate heat propagation path, thereby obtaining the set of all candidate heat propagation paths connected to the current data center server node, providing input for subsequent calculation of path propagation driving value and path propagation impedance value.

[0085] For each candidate heat propagation path, the path propagation drive value and path propagation impedance value are calculated. In this embodiment, the heat propagation edge weights corresponding to each node connectivity segment in the candidate heat propagation path are first read, and the temperature difference between adjacent data center server nodes is calculated. Then, the heat propagation edge weights of each node connectivity segment are multiplied by the corresponding temperature difference to obtain the propagation drive component of that segment. Finally, the propagation drive components on the entire candidate heat propagation path are summed to obtain the path propagation drive value of the candidate heat propagation path. For example, if a candidate heat propagation path contains 3 node connectivity segments with heat propagation edge weights of 0.72, 0.68, and 0.55, and corresponding temperature differences of 2.4℃, 1.8℃, and 1.2℃, then the path propagation drive value of the candidate heat propagation path is: 0.72×2.4+0.68×1.8+0.55×1.2. Subsequently, the path propagation impedance value is calculated for the same candidate heat propagation path: first, the airflow propagation direction offset angle of each node connection is measured, and the average offset angle of the entire path is obtained; then, the cumulative propagation distance of the entire path is calculated; finally, the pressure difference across the front and rear of the current cabinet is read as the airflow pressure difference inside the cabinet. In this embodiment, the average offset angle, cumulative propagation distance, and airflow pressure difference inside the cabinet are normalized and then weighted and summed to obtain the path propagation impedance value of the candidate heat propagation path. Specifically, when the average offset angle is greater than 35°, the cumulative propagation distance is greater than 450mm, or the airflow pressure difference inside the cabinet is less than 16Pa, the path propagation impedance value of the corresponding candidate heat propagation path is treated as a high-resistance path.

[0086] After obtaining the path propagation drive value and path propagation impedance value of each candidate heat propagation path, the available heat propagation capacity of each candidate heat propagation path is further determined. Specifically, firstly, the historical heat load changes of each candidate heat propagation path are extracted from the historical records from the previous continuous time window to the current continuous time window. In this embodiment, the average path-level heat flux change of each candidate heat propagation path within the past 5 continuous time windows is used as the historical heat load index of the path. Secondly, the heat dissipation capacity budget value corresponding to each data center server node traversed by the candidate heat propagation path is read, and the smallest heat dissipation capacity budget value is selected as the constraint capacity of the candidate heat propagation path. Subsequently, the historical heat load index is subtracted from the constraint capacity to obtain the available heat propagation capacity of the candidate heat propagation path. For example, if a candidate heat propagation path traverses 3 data center server nodes with corresponding heat dissipation capacity budget values ​​of 42W, 38W, and 35W respectively, and the historical heat load index within the past 5 continuous time windows is 11W, then the available heat propagation capacity of the candidate heat propagation path is 35W minus 11W, i.e., 24W. If the calculated result is less than 5W, the candidate heat propagation path is marked as a low-capacity path, and only the lowest projection ratio is retained in subsequent allocations.

[0087] The path propagation efficiency of each candidate heat propagation path is determined based on the aforementioned path propagation driving value, path propagation impedance value, and available heat propagation capacity of the path. Then, the node-level heat flux change value is proportionally allocated to each candidate heat propagation path according to the path propagation efficiency, thereby forming a path-level heat flux change sequence. In this embodiment, the path propagation driving value of each candidate heat propagation path is first divided by "1 + path propagation impedance value" to obtain the basic propagation capacity of the path. Then, the basic propagation capacity of the path is multiplied by the corresponding available heat propagation capacity of the path to obtain the unnormalized path propagation efficiency of the candidate heat propagation path. Subsequently, the unnormalized path propagation efficiencies of all candidate heat propagation paths corresponding to the current data center server node are summed, and the unnormalized path propagation efficiency of each candidate heat propagation path is divided by the sum to obtain the allocation ratio of that candidate heat propagation path. Finally, the node-level heat flux change value of the current data center server node is multiplied by the allocation ratio of each candidate heat propagation path, and projected onto the corresponding candidate heat propagation path. The projection value obtained by each candidate heat propagation path within the current window is recorded in the order of continuous time windows, thereby forming the path-level heat flux change sequence of each candidate heat propagation path. In this embodiment, to avoid a single path obtaining an excessively high allocation ratio, when the allocation ratio of any candidate heat propagation path is greater than 0.70, its upper limit is limited to 0.70, and the excess part is redistributed according to the unnormalized path propagation efficiency of the remaining candidate heat propagation paths. Through the above steps, the node-level heat flux change value obtained in the previous embodiment is further converted into a path-level heat flux change sequence constrained by the heat propagation edge weight, path propagation impedance value, and path available heat propagation capacity, thereby providing direct input for subsequent calculation of the heat accumulation rate of each heat propagation path and identification of pseudo-steady-state thermal states based on the path-level heat flux change sequence.

[0088] Step 3: Calculate the heat accumulation rate for each heat propagation path based on the path-level heat flux change sequence, and compare the heat accumulation rate with the temperature change rate of the corresponding data center server node. When the heat accumulation rate is consistently higher than the temperature change rate of the corresponding data center server node, the data center server node is determined to be in a pseudo-steady-state thermal state. This step identifies abnormal thermal states that are difficult to detect directly by temperature characterization alone, from the perspective of "whether heat is accumulating." Calculating the heat accumulation rate for each heat propagation path using the path-level heat flux change sequence can determine whether heat is continuously accumulating along the heat propagation path. Comparing the heat accumulation rate with the temperature change rate of the corresponding data center server node can identify a special state where heat has been continuously accumulating along the propagation path, but the node temperature characterization has not changed significantly. Identifying this state as a pseudo-steady-state thermal state is significant because it avoids the lag caused by judging solely based on temperature changes, thus identifying data center server nodes that need regulation before thermal risks manifest as obvious temperature increases.

[0089] When determining in step three that the data center server node is in a pseudo-steady-state thermal state, the following is included:

[0090] The duration for which the thermal accumulation rate of a data center server node is continuously higher than the corresponding temperature change rate is recorded within a continuous time window. The node thermal inertia compensation time is determined based on the hysteresis relationship between the processor temperature change sequence and the server air intake temperature change sequence, thereby obtaining the dynamic pseudo steady state determination time threshold.

[0091] Based on the rack-level heat propagation topology, the upstream thermal impact set of data center server nodes is extracted, and the number of data center server nodes in the upstream thermal impact set that meet the condition that the heat accumulation rate is higher than the temperature change rate within the same continuous time window is counted.

[0092] When the duration exceeds the dynamic pseudo-steady state determination time threshold and the number of data center server nodes that meet the conditions reaches the preset topology consistency determination threshold, the data center server node is determined to be in a pseudo-steady state hot state.

[0093] After obtaining the path-level heat flux change sequence for each heat propagation path in the aforementioned embodiments, a heat accumulation rate calculation is performed for each heat propagation path. Specifically, using 10s as a continuous time window, the path-level heat flux change sequence of the same heat propagation path within the most recent 6 consecutive time windows is continuously read. The path-level heat flux value of the current consecutive time window is subtracted from the path-level heat flux value of the previous consecutive time window, and then divided by the length of the consecutive time window to obtain the instantaneous heat accumulation rate of the heat propagation path within the current consecutive time window. Subsequently, a moving average is performed on the instantaneous heat accumulation rates within the most recent 3 consecutive time windows to obtain the heat accumulation rate used for judgment. In this embodiment, when the path-level heat flux values ​​of a certain heat propagation path in the 3 consecutive time windows are 18W, 24W, and 31W respectively, the instantaneous heat accumulation rate of the path in the 3rd consecutive time window can be calculated by the difference between adjacent windows and the time window length, and then averaged with the instantaneous heat accumulation rates of the previous 2 consecutive time windows to obtain the current heat accumulation rate used for judgment. After calculating the heat accumulation rate for each heat propagation path, each heat propagation path is mapped back to its corresponding target data center server node so that it can proceed to the next step of comparing the temperature change rate with the corresponding data center server node.

[0094] After obtaining the heat accumulation rate for each heat propagation path, the temperature change rate of the corresponding data center server node is calculated synchronously, and an initial pseudo-steady-state determination is performed. Specifically, the processor temperature change sequence of each data center server node is read using the same continuous time window as described above. The processor temperature at the end of the current continuous time window is subtracted from the processor temperature at the beginning of the current continuous time window, and then divided by the length of the continuous time window to obtain the temperature change rate of the data center server node within the current continuous time window. Subsequently, the heat accumulation rate of the corresponding heat propagation path of the same target data center server node is compared with the temperature change rate of the target data center server node; when the heat accumulation rate is greater than the temperature change rate, an event of "heat accumulation rate higher than temperature change rate" is recorded within the continuous time window. To avoid misjudgment due to single sampling fluctuations, this embodiment requires that the heat accumulation rate be at least 0.08 W / s higher than the temperature change rate to be recorded as a valid event. If this valid event occurs in multiple consecutive continuous time windows, the duration of "heat accumulation rate continuously higher than the corresponding temperature change rate" of the data center server node is accumulated.

[0095] After the cumulative duration has begun, the node thermal inertia compensation time is determined based on the lag relationship between the processor temperature change sequence and the server intake air temperature change sequence, thereby obtaining the dynamic pseudo-steady-state determination time threshold. Specifically, within the same continuous time window, the moment when the server intake air temperature change sequence first rises by 0.5°C is taken as the intake air disturbance start time, and the moment when the processor temperature change sequence first rises by 0.5°C is taken as the node temperature rise response time. The time difference between the two is recorded as the original lag time. Then, the average value of the original lag time in the most recent four consecutive time windows is continuously calculated, and this average value is determined as the node thermal inertia compensation time. In this embodiment, the basic determination time is taken as 30s. The basic determination time is added to the node thermal inertia compensation time to obtain the dynamic pseudo-steady-state determination time threshold. For example, if the original lag times of a data center server node in the most recent four consecutive time windows are 6s, 7s, 8s, and 7s, then its node thermal inertia compensation time is taken as 7s, and the corresponding dynamic pseudo-steady-state determination time threshold is 37s. After obtaining the threshold, the upstream thermal impact set of the data center server node is extracted based on the aforementioned rack-level heat propagation topology map. Within the current continuous time window, the number of data center server nodes in the upstream thermal impact set that satisfy the condition of "heat accumulation rate being higher than temperature change rate" is counted. In this embodiment, two preset topology consistency judgment thresholds are set; that is, when at least two data center server nodes in the upstream thermal impact set simultaneously satisfy this condition, the topology consistency constraint is considered satisfied.

[0096] By combining the cumulative duration results, the dynamic pseudo-steady-state determination time threshold, and the preset topology consistency determination threshold, the final determination of whether a data center server node is in a pseudo-steady-state thermal state is completed. Specifically, when the cumulative duration of a target data center server node's "heat accumulation rate continuously exceeding the corresponding temperature change rate" exceeds the corresponding dynamic pseudo-steady-state determination time threshold, and the number of data center server nodes meeting the conditions in its upstream thermal influence set reaches the preset topology consistency determination threshold, the target data center server node is determined to be in a pseudo-steady-state thermal state; otherwise, the target data center server node is only marked as an observation node and is not included in the subsequent heat dissipation capacity budget value calculation step.

[0097] Step 4: Based on the total heat dissipation capacity of the computer rack according to the heat propagation topology diagram, and combined with the propagation relationship of each data center server node in the heat propagation path, determine the budgeted heat dissipation capacity of each data center server node. This step transforms the limited heat dissipation capacity inside the rack into an allocable and constrainable control basis. The total heat dissipation capacity of the rack reflects the overall capacity of the rack to handle heat release. However, the propagation relationship of different data center server nodes in the heat propagation path is different, indicating that their influence on the rack's heat propagation structure and the degree of heat propagation constraints they are subject to are also different. Determining the budgeted heat dissipation capacity of each data center server node based on these two pieces of information can further transform the abstract heat propagation state into specific node-level heat dissipation resource constraints, thus providing a quantitative basis for subsequent steps to determine which data center server nodes need to be prioritized for adjustment and which data center server nodes can undertake new computing tasks.

[0098] When calculating the budgeted heat dissipation capacity for each data center server node in step four, the following is included:

[0099] Based on the rack-level heat propagation topology map, the number of heat propagation paths corresponding to each data center server node and the path-level heat flux change sequence of each path are counted to determine the heat propagation contribution of each data center server node.

[0100] Based on the path-level heat flux change sequence, the historical heat load changes of each heat propagation path are statistically analyzed, and the path congestion degree of the corresponding heat propagation path is determined.

[0101] The total heat dissipation capacity of the rack is allocated based on the heat propagation contribution of each data center server node and the path congestion of the corresponding heat propagation path, thereby obtaining the heat dissipation capacity budget value of each data center server node.

[0102] After obtaining the rack-level heat propagation topology map and the path-level heat flux change sequence corresponding to each heat propagation path in the aforementioned embodiments, the total heat dissipation capacity of the computer rack is determined based on the heat propagation topology map, combined with the propagation relationship of each data center server node in the heat propagation path, to determine the budgeted heat dissipation capacity value of each data center server node. Specifically, data on the rack inlet air temperature, rack exhaust air temperature, and rack air supply volume are collected, and the instantaneous heat dissipation capacity of the computer rack is calculated based on the air supply volume and the inlet / exhaust air temperature difference. Subsequently, the results of multiple instantaneous heat dissipation capacities are averaged within a continuous time window to obtain the current total heat dissipation capacity of the rack. In this embodiment, the continuous time window can be 30 seconds, and the calculation results of the most recent 5 continuous time windows are averaged to reduce the impact of short-term airflow fluctuations on the heat dissipation capacity calculation, thereby obtaining a stable total heat dissipation capacity of the rack.

[0103] After obtaining the total heat dissipation capacity of the rack, the number of heat propagation paths corresponding to each data center server node is counted based on the rack-level heat propagation topology map, and the path-level heat flux change sequences corresponding to these paths are extracted. By statistically analyzing the path-level heat flux change values ​​of each heat propagation path within multiple consecutive time windows, the average path-level heat flux value of each path is calculated. Then, the average path-level heat flux values ​​belonging to the same data center server node are summarized to obtain the heat propagation contribution of that data center server node, which is used to characterize the heat output capability of that node in the overall heat propagation process of the rack.

[0104] After obtaining the heat propagation contribution of each data center server node, the historical heat load changes of each path within multiple consecutive time windows are statistically analyzed based on the path-level heat flux change sequence corresponding to each heat propagation path. Specifically, the heat load level of each heat propagation path is obtained by calculating the average value of the path-level heat flux change sequence and the change amplitude between adjacent time windows, and the path congestion level of the corresponding path is determined based on the average heat load and the change amplitude. In this embodiment, when the average path-level heat flux exceeds 1.2 times the average value of all paths in the rack, the path is marked as a high-load path; when the change amplitude of the path-level heat flux exceeds 20% of the average value, the path is marked as a fluctuating path, and the path congestion level evaluation level of the path is increased accordingly.

[0105] Having obtained the total heat dissipation capacity of the server rack, the heat propagation contribution of each data center server node, and the path congestion level of each heat propagation path, a basic allocation ratio of the total heat dissipation capacity of the server rack is determined based on the heat propagation contribution of each node. This allocation ratio is then adjusted according to the path congestion level of the corresponding heat propagation path to obtain the final heat dissipation capacity budget value for each data center server node. Specifically, the total heat dissipation capacity of the server rack can be initially allocated according to the proportion of each node's heat propagation contribution to the total contribution of all nodes. Then, the allocation ratio is appropriately reduced for nodes with higher path congestion levels, while the allocation ratio is appropriately increased for nodes with lower path congestion levels, to obtain the heat dissipation capacity budget value for each node. The heat dissipation capacity budget value then serves as an important constraint for performing heat propagation path reconstruction control and computation task scheduling in subsequent steps.

[0106] Step 5: Based on the heat dissipation capacity budget, perform heat propagation path reconstruction control on data center server nodes in a pseudo-steady-state thermal state. This involves adjusting the distribution of computing tasks among different processing cores and scheduling new computing tasks to data center server nodes with higher heat dissipation capacity budgets. This step translates the heat propagation analysis results and heat dissipation resource allocation results obtained in the previous steps into actual control actions. For data center server nodes in a pseudo-steady-state thermal state, if the original task distribution and heat propagation path are maintained, heat may continue to accumulate along the current path; therefore, heat propagation path reconstruction control is necessary. By adjusting the distribution of computing tasks among different processing cores, the distribution of heat sources and the heat diffusion method within the node can be changed. By scheduling new computing tasks to data center server nodes with higher heat dissipation capacity budgets, subsequent heat loads can be guided to locations with stronger heat dissipation capacity. The purpose of this step is to transform the aforementioned identification and budget results into actual adjustment measures, thereby achieving proactive reconstruction of the heat propagation path within the rack and reorganization of the heat load distribution.

[0107] After obtaining the heat dissipation capacity budget value of each data center server node and determining that some nodes are in a pseudo-steady-state thermal state in the aforementioned embodiments, heat propagation path reconstruction control is performed on the data center server nodes in the pseudo-steady-state thermal state according to the heat dissipation capacity budget value.

[0108] Specifically, from the data center server nodes currently identified as being in a pseudo-steady-state thermal state, nodes with a heat dissipation capacity budget lower than the average budget of rack nodes are selected as nodes to be refactored. Then, the heat propagation path of this node is read based on the rack-level heat propagation topology, and the node connectivity and corresponding heat propagation edge weights on this path are extracted. When the average path-level heat flux change of a path is higher than the average of all paths in the rack, this path is identified as the target heat propagation path requiring refactoring control.

[0109] After identifying the target heat propagation path, the heat propagation path of the data center server node in the pseudo-steady-state thermal state is identified based on the rack-level heat propagation topology map, and candidate computing nodes different from the heat propagation path are determined.

[0110] Specifically, nodes located on the target heat propagation path and nodes with direct heat propagation connectivity to the path are removed from all data center server nodes in the rack, and the remaining nodes are retained as initial candidate nodes. Then, based on the heat dissipation capacity budget value corresponding to each node, only nodes with a heat dissipation capacity budget value higher than the average budget value of the rack nodes are retained as candidate computing nodes to avoid heat load concentration again after task migration.

[0111] After obtaining the candidate computing nodes, the thermal inertia buffering capacity of each candidate computing node is determined based on the processor temperature change sequence of the corresponding processing core and the corresponding node-level heat flux change value.

[0112] Specifically, for each candidate computing node, the processor temperature change sequence and the corresponding node-level heat flux change value within the most recent consecutive time windows are read. The processor temperature rise rate per unit time is calculated, and the ratio of the temperature rise rate to the corresponding node-level heat flux change value is calculated to obtain the initial thermal inertia buffering capacity of the candidate computing node. Then, the calculation results of the most recent three consecutive time windows are averaged to obtain a stable thermal inertia buffering capacity value, and the candidate computing nodes are sorted according to this value.

[0113] Based on the ranking of candidate computing nodes by their thermal inertia buffering capacity and their corresponding heat dissipation capacity budget, the distribution of computing tasks among different processing cores is adjusted, and new computing tasks are scheduled to data center server nodes with higher heat dissipation capacity budgets.

[0114] Specifically, the newly added computing tasks that can be migrated in the node to be reconstructed are divided into task migration sets, and the task migration sets are scheduled to candidate computing nodes that simultaneously meet the criteria of "heat dissipation capacity budget value is higher than the average budget value of rack nodes and thermal inertia buffer capability is among the top of candidate nodes". Then, within the candidate computing node, the migration tasks are assigned to the processing cores with lower current processor temperature change rate, thereby realizing thermal migration control between processing cores.

[0115] On the other hand, in this embodiment, the various preset thresholds involved in each step can be determined through experimental testing or engineering experience. Specifically, under the typical operating environment of the target data center rack, multiple rounds of server load change experiments, temperature change monitoring, and airflow propagation tests can be conducted to statistically analyze parameters such as server exhaust temperature, intake temperature, power changes, and airflow propagation time. Combined with long-term operating experience, parameter ranges that can distinguish between normal heat propagation states and abnormal heat accumulation states can be determined, thereby obtaining each preset threshold. These thresholds can be appropriately adjusted according to different rack structures, server densities, and operating load conditions.

[0116] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters and thresholds in the formulas are set by those skilled in the art according to the actual situation.

[0117] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.

Claims

1. A method for regulating the power consumption of a data center server based on temperature sensing, characterized in that, include: Step 1: Establish an airflow connectivity matrix based on the spatial location relationship and airflow organization structure of the data center server nodes in the rack, and construct a rack-level heat propagation topology map based on the airflow connectivity direction and connectivity weight to determine the upstream heat impact set and downstream heat-affected set of each data center server node; Step 2: Based on the rack-level heat propagation topology map, collect the processor power change sequence, processor temperature change sequence, and server intake air temperature change sequence of each data center server node. Calculate the node-level heat flux change value based on the processor power change sequence and processor temperature change sequence, and map the node-level heat flux change value to the corresponding heat propagation path according to the heat propagation edge weights between nodes in the rack-level heat propagation topology map to form a path-level heat flux change sequence. Step 3: Calculate the heat accumulation rate of each heat propagation path within a continuous time window based on the path-level heat flux change sequence, and compare the heat accumulation rate with the temperature change rate of the corresponding data center server node. When the heat accumulation rate is consistently higher than the temperature change rate of the corresponding data center server node, it is determined that the data center server node is in a pseudo-steady-state thermal state. Step 4: Based on the rack-level heat propagation topology diagram and path-level heat flux change sequence, determine the total heat dissipation capacity of the computer rack and the propagation relationship of each data center server node in the heat propagation path to determine the budget value of the heat dissipation capacity of each data center server node. Step 5: Perform heat propagation path reconstruction control on the data center server nodes in the pseudo-steady-state thermal state according to the heat dissipation capacity budget value. This is done by adjusting the distribution of computing tasks among different processing cores and scheduling new computing tasks to data center server nodes with higher heat dissipation capacity budget values. Step 1, establishing the airflow connectivity matrix, includes the following steps: Based on the spatial relationship between the exhaust surface of each data center server node in the rack and the intake surface of the adjacent data center server node, as well as the spatial projection range of the structural obstruction between them, the obstruction ratio of the exhaust propagation path is calculated. When the obstruction ratio exceeds a preset threshold, the airflow connectivity weight between the corresponding data center server nodes is reduced. Based on the airflow return paths formed by the side channels, top channels, and structural gaps of the rack, calculate the connection distance between each airflow return path and the air intake surface of the target data center server node, and establish the indirect connection relationship between the corresponding data center server nodes in the airflow connection relationship matrix when the connection distance is less than the preset return threshold. Based on the installation height of each data center server node and the angle between the server exhaust direction and the upward direction of the hot airflow, the airflow propagation direction between adjacent data center server nodes is determined. When the airflow propagation direction points only to a single data center server node, the corresponding connection relationship is determined as a unidirectional connection relationship. When the airflow propagation direction points to two data center server nodes at the same time, the corresponding connection relationship is determined as a bidirectional connection relationship. Based on the results of the exhaust propagation path obstruction ratio correction, the airflow return path connectivity, and the direction determination results of unidirectional and bidirectional connectivity, a final airflow connectivity matrix is ​​generated, and a rack-level heat propagation topology map is constructed based on the final airflow connectivity matrix. When constructing a rack-level heat propagation topology based on an airflow connectivity matrix, the steps for determining the heat propagation edge weights for the node connectivity relationships in the airflow connectivity matrix include: Within a continuous time window, the server operating parameter change data of the source data center server node and the target data center server node are collected, and the heat propagation correlation is identified based on the time propagation order between the exhaust temperature change of the source data center server node and the intake temperature change of the target data center server node. When the time propagation order meets the preset propagation coupling trigger threshold, the corresponding node connectivity is determined as the heat propagation correlation edge. A heat propagation association sequence is constructed based on the occurrence of heat propagation association edges in multiple consecutive time windows. When the number of consecutive occurrences of the heat propagation association sequence reaches a preset propagation stability judgment threshold, the connectivity relationship of the corresponding nodes is determined as a stable heat propagation path. The path propagation driving value is determined based on the change range of server operating parameters and the propagation distance between adjacent nodes in the stable heat propagation path, and the path propagation impedance value is determined based on the degree of deviation of airflow propagation direction between adjacent nodes and the airflow pressure difference inside the cabinet. The path propagation coefficient is determined based on the nonlinear propagation relationship between the path propagation driving value and the path propagation impedance value. The heat propagation edge weight of the corresponding node connectivity relationship is determined based on the path propagation coefficient and the duration of the stable heat propagation path. The airflow connectivity matrix is ​​then weighted to generate a rack-level heat propagation topology map. Step two, in calculating the nodal-level heat flux change, includes: Collect the processor power change sequence, processor temperature change sequence, server intake air temperature change sequence, and corresponding heat propagation edge weights of data center server nodes within a continuous time window. The node thermal inertia hysteresis value is determined based on the time difference between the processor power change sequence and the processor temperature change sequence, and the node thermal inertia hysteresis value is corrected for air intake disturbance based on the server intake air temperature change sequence. Based on the rack-level heat propagation topology, the upstream heat impact set and downstream heat-affected set of the data center server node are extracted, and the node heat propagation gradient value is determined by combining the corresponding heat propagation edge weights and the temperature difference between adjacent data center server nodes. The equivalent thermal storage capacity of a node is determined based on the node thermal inertia hysteresis value, the air inlet disturbance correction result, and the node thermal propagation gradient value. The node-level heat flux change value of the data center server node is determined based on the power change amplitude corresponding to the processor power change sequence, the node heat propagation gradient value, and the node equivalent heat storage. When determining in step three that the data center server node is in a pseudo-steady-state thermal state, the following is included: The duration for which the thermal accumulation rate of a data center server node is continuously higher than the corresponding temperature change rate is recorded within a continuous time window. The node thermal inertia compensation time is determined based on the hysteresis relationship between the processor temperature change sequence and the server air intake temperature change sequence, thereby obtaining the dynamic pseudo steady state determination time threshold. Based on the rack-level heat propagation topology, the upstream thermal impact set of data center server nodes is extracted, and the number of data center server nodes in the upstream thermal impact set that meet the condition that the heat accumulation rate is higher than the temperature change rate within the same continuous time window is counted. When the duration exceeds the dynamic pseudo-steady state determination time threshold and the number of data center server nodes that meet the conditions reaches the preset topology consistency determination threshold, the data center server node is determined to be in a pseudo-steady state hot state.

2. The method for adjusting the power consumption of a data center server based on temperature sensing according to claim 1, characterized in that, The determination of heat propagation edge weights also includes a heat propagation contention determination step, which includes: Identify multiple source data center server nodes corresponding to the air intake area of ​​the same target data center server node. When the exhaust propagation direction of multiple source data center server nodes points to the air intake area at the same time, determine that a heat propagation competition relationship is formed between the corresponding nodes. Within a continuous time window, the propagation time of the exhaust air temperature change of each source data center server node to the air intake area of ​​the target data center server node is collected, and the heat propagation competition order is determined according to the order of propagation time. When the source data center server node with the earliest propagation time meets the preset propagation channel occupancy threshold, the node connectivity relationship corresponding to the source data center server node is determined as the priority hot propagation path, and the path propagation drive value of the other node connectivity relationships is reduced. The path propagation coefficient is corrected based on the competition between the priority heat propagation path and the connectivity of the remaining nodes, and the corrected path propagation coefficient is then substituted back into the heat propagation edge weight calculation process to determine the final heat propagation edge weight.

3. The method for adjusting the power consumption of a data center server based on temperature sensing according to claim 2, characterized in that, The process of determining the weights of heat propagation edges involves performing a heat propagation resonance determination step, including: Within a continuous time window, server operating parameter change data of multiple source data center server nodes in the upstream thermal impact set of the same target data center server node are collected, and the exhaust heat fluctuation frequency of each source data center server node is identified based on the periodic change characteristics of exhaust temperature change within the continuous time window. The thermal propagation frequency coupling relationship is determined based on the frequency difference between the exhaust thermal fluctuation frequencies of server nodes in each source data center. When the frequency difference is less than the preset propagation resonance judgment threshold, the thermal propagation relationship between the corresponding nodes is determined as a thermal propagation resonance relationship. The thermal propagation resonance path is identified based on the spatial distribution of the connectivity relationships of multiple nodes forming a thermal propagation resonance relationship in the thermal propagation topology map, and the resonance propagation intensity is calculated based on the number of nodes forming a resonance relationship and the magnitude of the exhaust temperature change. When the resonance propagation intensity exceeds the preset resonance intensity trigger threshold, the path propagation coefficient of the corresponding node connectivity is amplified and corrected, and the heat propagation edge weight is re-determined based on the corrected path propagation coefficient.

4. The method for adjusting the power consumption of a data center server based on temperature sensing according to claim 1, characterized in that, In step two, mapping node-level heat flux changes to heat propagation paths to form path-level heat flux change sequences includes: In the rack-level heat propagation topology map, identify multiple downstream heat-affected cluster nodes corresponding to the data center server nodes, and determine the candidate heat propagation paths connected to the data center server nodes based on the heat propagation edge weights corresponding to the connectivity of each node. The path propagation driving value is determined based on the heat propagation edge weights of the connectivity relationships of each node in the candidate heat propagation path and the temperature difference between adjacent data center server nodes. The path propagation impedance value is determined based on the degree of airflow propagation direction offset, path propagation distance, and airflow pressure difference inside the rack in the candidate heat propagation path. The available heat transfer capacity of each candidate heat transfer path is determined based on the historical heat load changes of each candidate heat transfer path in the rack-level heat transfer topology diagram and the corresponding heat dissipation capacity budget value. The path propagation efficiency is determined based on the path propagation driving value, the path propagation impedance value, and the available heat propagation capacity of the path. The node-level heat flux change value is then proportionally allocated to each candidate heat propagation path according to the path propagation efficiency to form a path-level heat flux change sequence.

5. The method for adjusting the power consumption of a data center server based on temperature sensing according to claim 1, characterized in that, When calculating the budgeted heat dissipation capacity for each data center server node in step four, the following is included: Based on the rack-level heat propagation topology map, the number of heat propagation paths corresponding to each data center server node and the path-level heat flux change sequence of the corresponding paths are counted to determine the heat propagation contribution of each data center server node. Based on the path-level heat flux change sequence, the historical heat load changes of each heat propagation path are statistically analyzed, and the path congestion degree of the corresponding heat propagation path is determined. The total heat dissipation capacity of the rack is allocated based on the heat propagation contribution of each data center server node and the path congestion of the corresponding heat propagation path, thereby obtaining the heat dissipation capacity budget value of each data center server node.

6. The method for adjusting the power consumption of a data center server based on temperature sensing according to claim 1, characterized in that, When performing heat propagation path reconstruction control in step five, the following is included: Based on the rack-level heat propagation topology map, identify the heat propagation path of the data center server node in a pseudo-steady-state thermal state, and determine the candidate computing node that is different from the heat propagation path; The thermal inertia buffering capacity of each candidate computing node is determined based on the processor temperature change sequence of the corresponding processing core of each candidate computing node and the corresponding node-level heat flux change value. New computing tasks are prioritized and scheduled to candidate computing nodes that are not part of the heat propagation path and have high thermal inertia buffering capacity, so as to achieve thermal migration control between processing cores.