Green data center liquid cooling thermal management method and system
Patent Information
- Application Number
- CN202511040667.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2045-07-28
AI Technical Summary
长期的累积效应会导致冷板内部微通道结构产生疲劳,甚至出现微裂纹,降低热传导效率,形成一种隐蔽的、慢性的硬件老化
[0069]分配策略调整模块,用于基于驱使对应的处理器降低工作频率或降低功耗之后的反馈情况,控制上层工作负载调度器调整资源分配策略;
Smart Images

Figure CN120935989B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of data center thermal management, and specifically to a green data center liquid cooling thermal management method and system. Background Technology
[0002] In modern data centers, especially green data centers where low-energy operation is a core indicator, efficient management of heat generated by servers is crucial for ensuring stable system operation and achieving energy efficiency goals. However, in actual operation, data centers often adopt a multi-tenant hosting model, which presents unique challenges to thermal management.
[0003] The recurring, intense localized temperature differences on the cold plate create differential thermal stress on the material. The long-term cumulative effect leads to fatigue in the microchannel structure within the cold plate, even causing microcracks, reducing heat transfer efficiency, and resulting in a hidden, chronic hardware aging process. This puts the operator in an irreconcilable dilemma: to absolutely guarantee tenant performance, the entire shared liquid cooling circuit must be operated at its highest power level, but this results in enormous energy consumption. Conversely, strictly implementing energy-saving strategies cannot cope with unpredictable, localized heat bursts, constantly facing the risk of tenant complaints about substandard performance or even premature hardware failure.
[0004] To address the aforementioned issues, existing technologies urgently need improvement. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this application provides a green data center liquid cooling thermal management method and system. This method has the advantages of proactively identifying and managing local hot spots in the data center liquid cooling system, effectively balancing system performance and energy efficiency through intelligent intervention and feedback mechanisms, and avoiding hardware damage and unnecessary energy consumption.
[0006] This application provides a green data center liquid cooling thermal management method, the key technical points of which are:
[0007] A green data center liquid cooling thermal management method includes:
[0008] Continuously monitor the coolant outlet temperature corresponding to the liquid cooling circuit inside the server to obtain the fluctuation characteristics of the coolant outlet temperature;
[0009] Based on the fluctuation characteristics, the cumulative heat load of the critical area of the cold plate in the liquid cooling circuit is evaluated to obtain the hardware heat capacity margin of the critical area of the cold plate.
[0010] The hardware thermal capacity margin is compared with the preset thermal capacity threshold to determine whether to trigger the intervention mechanism and obtain the intervention judgment result.
[0011] When the intervention judgment result indicates that the intervention mechanism has been triggered, a local heat dissipation bottleneck is created in the hot spot area;
[0012] By addressing localized heat dissipation bottlenecks, the processor core temperature in hot areas is raised to a preset internal temperature protection threshold, causing the corresponding processor to reduce its operating frequency or power consumption.
[0013] Based on the feedback after driving the corresponding processor to reduce its operating frequency or power consumption, the upper-layer workload scheduler is controlled to adjust the resource allocation strategy.
[0014] Continuously monitor the recovery progress of hardware thermal capacity margin, and stop creating local heat dissipation bottlenecks when the hardware thermal capacity margin is fully restored.
[0015] Through the above scheme, this method can proactively identify and manage local hot spots in the liquid cooling system of the data center. Through intelligent intervention and feedback mechanisms, it effectively balances system performance and energy efficiency, and avoids hardware damage and unnecessary energy consumption.
[0016] To further address the issue, this application also proposes that, when the intervention judgment result indicates that an intervention mechanism has been triggered, the steps for creating a localized heat dissipation bottleneck in the hotspot area include:
[0017] Monitor the instruction execution flow inside the processor and identify the execution mode of a specific instruction set;
[0018] The execution mode of a specific instruction set is compared with a dynamic threshold to determine whether a high power density region is formed, thus obtaining the region determination result;
[0019] When the region determination result indicates the formation of a high power density region, the corresponding high power density region is recorded as a hot spot region, and targeted performance intervention is performed on specific computing units inside the processor to create local heat dissipation bottlenecks.
[0020] Monitor the temperature or power density of a specific computing unit, and withdraw the targeted performance intervention once the temperature or power density of the specific computing unit returns to the normal range.
[0021] By monitoring the instruction execution flow and identifying high power density areas, the above-mentioned solution can accurately locate hotspots and perform targeted performance interventions, making the manufacturing of local heat dissipation bottlenecks more precise and efficient.
[0022] To improve the solution, this application also proposes a step of continuously monitoring the coolant outlet temperature corresponding to the liquid cooling circuit within the server and obtaining the fluctuation characteristics of the coolant outlet temperature, including:
[0023] Continuously monitor the coolant outlet temperature corresponding to the liquid cooling circuit inside the server to obtain the fluctuation characteristics of the coolant outlet temperature;
[0024] Continuously monitor the inlet and outlet pressures of the coolant flowing through the critical area of the cold plate to obtain the coolant pressure difference in the critical area of the cold plate;
[0025] Based on the trend of coolant pressure difference changes, determine whether the internal structure of the key area of the cold plate has changed, and obtain the result of the internal change judgment.
[0026] When the internal change judgment result indicates a change, adjust the mapping relationship between the fluctuation characteristics and the cumulative heat load.
[0027] By combining the above scheme with the changing trend of coolant pressure difference, the changes in the internal structure of the cold plate can be judged more accurately, thereby dynamically adjusting the heat load assessment model and improving the accuracy of heat capacity margin assessment.
[0028] To improve the solution, this application also proposes a step for controlling the upper-layer workload scheduler to adjust resource allocation strategies based on feedback after driving the corresponding processor to reduce its operating frequency or power consumption.
[0029] Continuously monitor the resource allocation behavior corresponding to the physical signals from the upper-layer workload scheduler that reduce the processor's operating frequency or power consumption;
[0030] Based on resource allocation behavior, the effectiveness of physical signals in guiding the behavior of the upper-layer workload scheduler is evaluated to obtain the guidance effectiveness;
[0031] When bootstrapping effectiveness declines, identify the adaptive strategies that upper-layer workload schedulers have developed;
[0032] Adjust the parameters of local heat dissipation bottlenecks according to the adaptive strategy;
[0033] Based on the adjusted parameters of the local heat dissipation bottleneck, a local heat dissipation bottleneck is created, causing the processor core temperature in the hot spot area to rise to the preset internal temperature protection threshold, thereby driving the corresponding processor to reduce its operating frequency or power consumption.
[0034] The physical signals that cause the processor to reduce its operating frequency or power consumption are used as feedback to control the upper-layer workload scheduler to adjust the resource allocation strategy.
[0035] By introducing the above scheme, the effectiveness of guiding the behavior of the upper-layer workload scheduler and the identification of adaptive strategies are introduced, enabling the thermal management system to dynamically adjust intervention parameters, effectively respond to the adaptive behavior of the scheduler, and achieve more intelligent resource allocation optimization.
[0036] To improve the solution, this application also proposes a method for evaluating the effectiveness of physical signals in guiding the behavior of the upper-layer workload scheduler based on resource allocation behavior. The steps for obtaining the effectiveness of this guidance include:
[0037] Monitor the adjustment behavior of the upper-layer workload scheduler in allocating computing resources and generate a sequence of adjustment behaviors;
[0038] Based on the adjustment behavior sequence, analyze and identify the response patterns of the upper-layer workload scheduler to physical signals;
[0039] The response pattern is compared with the preset response characteristics to determine whether the response pattern has changed, and the result of the response change judgment is obtained.
[0040] When the response change judgment result indicates a change, the effectiveness evaluation rules for the physical signal to guide the behavior of the upper-layer workload scheduler are redefined based on the changed response pattern.
[0041] Based on the effectiveness evaluation rules, evaluate the effectiveness of physical signals in guiding the behavior of the upper-layer workload scheduler.
[0042] The above scheme, by analyzing the scheduler's response patterns to physical signals and dynamically adjusting the evaluation rules, further improves the accuracy and adaptability of the effectiveness assessment of behavioral guidance, ensuring the continuous effectiveness of the feedback mechanism.
[0043] To improve the solution, this application also proposes steps for identifying adaptive strategies developed by the upper-layer workload scheduler when bootstrapping effectiveness declines, including:
[0044] To obtain the rate, duration, and magnitude of the decline in guiding effectiveness;
[0045] The rate, duration, and magnitude of the decline in guidance effectiveness are compared with preset thresholds to obtain the guidance effectiveness judgment results;
[0046] Based on the effectiveness of the guidance, identify the adaptive strategies that the upper-layer workload scheduler has developed.
[0047] The above scheme enables timely identification of adaptive strategies developed by the scheduler based on the rate, duration, and magnitude of the decline in guidance effectiveness, providing a basis for subsequent parameter adjustments and enhancing the robustness of the system.
[0048] To improve the solution, this application also proposes steps for adjusting the parameters of local heat dissipation bottlenecks according to an adaptive strategy, including:
[0049] Adjust the flow rate or local pressure of the coolant flowing through the hot spot area;
[0050] The strength of the local heat dissipation bottleneck is changed according to the adjusted flow rate or local pressure of the coolant, so as to dynamically adjust the parameters for manufacturing the local heat dissipation bottleneck.
[0051] The above scheme achieves fine-grained dynamic adjustment of intervention parameters by adjusting the coolant flow rate or local pressure to change the intensity of local heat dissipation bottlenecks, making thermal management more flexible and precise.
[0052] To improve the solution, this application also proposes a step to assess the cumulative heat load of the critical region of the cold plate in the liquid cooling circuit based on fluctuation characteristics, and to obtain the hardware heat capacity margin of the critical region of the cold plate, including:
[0053] Identify thermal events within fluctuation characteristics; thermal events include the magnitude of instantaneous temperature increases.
[0054] Based on the intensity of the thermal event, the contribution of the thermal event to the thermal stress in the critical area of the cold plate is quantified.
[0055] The cumulative heating stress contribution is used to obtain the cumulative heat load.
[0056] The accumulated heat load is compared with the preset thermal damage threshold of the critical area of the cold plate to obtain the hardware thermal capacity margin.
[0057] By identifying thermal events and quantifying their contribution to the thermal stress of the cold plate, the above scheme can more accurately accumulate the heat load and compare it with the thermal damage threshold, thereby obtaining a more reliable hardware thermal capacity margin and effectively warning of potential hardware risks.
[0058] To improve the solution, this application also proposes steps for targeted performance intervention on specific computing units within the processor, including:
[0059] Get the current temperature of a specific computing unit inside the processor;
[0060] Determine the target performance state of a specific computing unit based on the current temperature;
[0061] Based on the target performance state, the clock frequency or operating voltage of a specific computing unit can be adjusted to achieve targeted performance intervention for that specific computing unit.
[0062] The above scheme can precisely adjust the clock frequency or operating voltage of a specific computing unit according to its current temperature, thereby achieving precise targeted performance intervention in hotspot areas and effectively controlling local temperature.
[0063] A green data center liquid cooling thermal management system for performing green data center liquid cooling thermal management, comprising:
[0064] The fluctuation characteristic acquisition module is used to continuously monitor the coolant outlet temperature corresponding to the liquid cooling circuit in the server and acquire the fluctuation characteristics of the coolant outlet temperature.
[0065] The capacity margin assessment module is used to assess the cumulative heat load of the critical area of the cold plate in the liquid cooling circuit based on the fluctuation characteristics, and to obtain the hardware heat capacity margin of the critical area of the cold plate.
[0066] The intervention mechanism judgment module is used to compare the hardware thermal capacity margin with the preset thermal capacity threshold to determine whether the intervention mechanism is triggered and to obtain the intervention judgment result.
[0067] The heat dissipation bottleneck manufacturing module is used to create local heat dissipation bottlenecks in hot spots when the intervention judgment result indicates that the intervention mechanism has been triggered.
[0068] The bottleneck strategy execution module is used to make the processor core temperature in the hot spot area rise to a preset internal temperature protection threshold by using local heat dissipation bottlenecks, thereby driving the corresponding processor to reduce its operating frequency or power consumption.
[0069] The allocation strategy adjustment module is used to control the upper-layer workload scheduler to adjust the resource allocation strategy based on the feedback after driving the corresponding processor to reduce its operating frequency or power consumption.
[0070] The bottleneck recovery management module is used to continuously monitor the recovery progress of hardware thermal capacity margin. When the hardware thermal capacity margin is fully recovered, it stops creating local heat dissipation bottlenecks.
[0071] The above solution provides a system for implementing the aforementioned green data center liquid cooling thermal management method. Through modular design, the implementation of the method becomes more specific and feasible, facilitating deployment and management.
[0072] In summary, the green data center liquid cooling thermal management method and system provided in this application continuously monitors the temperature fluctuation characteristics of the coolant outlet, assesses the cumulative heat load and hardware thermal capacity margin of key areas of the cold plate, and creates local heat dissipation bottlenecks in hot spots when the margin is insufficient, driving the processor to reduce frequency or power consumption. Then, based on feedback, the resource allocation strategy of the upper-layer workload scheduler is adjusted, and the margin is monitored to stop intervention when it recovers. This effectively solves the thermal management problem caused by local hot spots and non-cooperative software behavior in multi-tenant environments, which traditional thermal management methods cannot handle. This method has the advantages of being able to proactively identify and manage local hot spots in data center liquid cooling systems, effectively balancing system performance and energy efficiency through intelligent intervention and feedback mechanisms, and avoiding hardware damage and unnecessary energy consumption. Attached Figure Description
[0073] Figure 1 This is a flowchart of a green data center liquid cooling thermal management method according to one embodiment of the present invention;
[0074] Figure 2This is one of the flowcharts of a green data center liquid cooling thermal management method according to another embodiment of the present invention;
[0075] Figure 3 This is a second flowchart of a green data center liquid cooling thermal management method according to another embodiment of the present invention;
[0076] Figure 4 This is the third flowchart of a green data center liquid cooling thermal management method according to another embodiment of the present invention;
[0077] Figure 5 This is the fourth flowchart of a green data center liquid cooling thermal management method according to another embodiment of the present invention;
[0078] Figure 6 This is a system block diagram of a green data center liquid-cooled thermal management system according to another embodiment of the present invention;
[0079] Explanation of reference numerals in the attached figures:
[0080] 1. Green Data Center Liquid Cooling Thermal Management System; 11. Fluctuation Characteristic Acquisition Module; 12. Capacity Margin Assessment Module; 13. Intervention Mechanism Judgment Module; 14. Heat Dissipation Bottleneck Creation Module; 15. Bottleneck Strategy Execution Module; 16. Allocation Strategy Adjustment Module; 17. Bottleneck Recovery Management Module. Detailed Implementation
[0081] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments. The components of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0082] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0083] Traditional liquid cooling thermal management methods for green data centers suffer from limitations in real-time identification of hotspot locations, intensities, and durations caused by non-cooperative workload schedulers in multi-tenant server architectures. This prevents the regionalized control of coolant flow, hindering effective thermal stress relief and resulting in overall cooling of the entire loop, thus failing to meet data center energy efficiency targets. Specifically, the central thermal management system can only sense the averaged coolant temperature, failing to detect localized temperature unevenness on the cold plates. This prevents targeted cooling adjustments and induces differential thermal stress in the hardware, leading to fatigue of the internal structure of the cold plates and reduced heat transfer efficiency.
[0084] In response, this application proposes a green data center liquid cooling thermal management method, combining... Figure 1 As shown, it includes:
[0085] S1, continuously monitor the coolant outlet temperature corresponding to the liquid cooling circuit in the server, and obtain the fluctuation characteristics of the coolant outlet temperature.
[0086] S2. Based on the fluctuation characteristics, evaluate the cumulative heat load of the critical area of the cold plate in the liquid cooling circuit to obtain the hardware heat capacity margin of the critical area of the cold plate.
[0087] S3 compares the hardware thermal capacity margin with the preset thermal capacity threshold to determine whether to trigger the intervention mechanism and obtain the intervention judgment result.
[0088] S4, when the intervention judgment result indicates that the intervention mechanism has been triggered, a local heat dissipation bottleneck is created in the hot spot area;
[0089] S5, by addressing localized heat dissipation bottlenecks, causes the processor core temperature in hot areas to rise to a preset internal temperature protection threshold, thereby driving the corresponding processor to reduce its operating frequency or power consumption.
[0090] S6, based on the feedback after driving the corresponding processor to reduce its operating frequency or power consumption, controls the upper-layer workload scheduler to adjust the resource allocation strategy.
[0091] S7 continuously monitors the recovery progress of hardware thermal capacity margin, and stops creating local heat dissipation bottlenecks when the hardware thermal capacity margin is fully restored.
[0092] The fluctuation characteristics of the coolant outlet temperature refer to the dynamic pattern of the coolant temperature change over time as it flows out of the liquid cooling circuit, including information such as instantaneous temperature rises, falls, frequency, and amplitude. This can be achieved using high-precision temperature sensors combined with a data acquisition system, such as thermistor arrays or infrared temperature sensors. Its main purpose is to capture transient changes in heat generated inside the server in real time, providing fundamental data for subsequent thermal load assessment. The cumulative thermal load of the critical area of the cold plate refers to the total heat or accumulated thermal stress borne by the cold plate area in the liquid cooling circuit that is in direct contact with high heat density components over a period of time. This can be calculated using an algorithm model based on temperature fluctuation characteristics, such as through integration or weighted averaging. Its main purpose is to quantify the potential damage that local hotspots may cause to the hardware, serving as a basis for assessing the hardware's health status. The hardware thermal capacity margin refers to the additional thermal load capacity that the critical area of the cold plate can withstand without thermal damage. The thermal capacity threshold can be obtained by comparing the preset thermal damage threshold of the critical area of the cold plate with the cumulative thermal load. This is primarily to intuitively reflect the remaining capacity of the hardware to resist thermal stress, providing a basis for determining whether intervention measures are needed. The thermal capacity threshold refers to a pre-set critical value used to determine whether the hardware's thermal capacity margin has reached the point where an intervention mechanism needs to be triggered. It can be determined based on the hardware's design life, reliability requirements, and actual operating experience. Its main purpose is to avoid over-intervention or under-intervention, ensuring the accuracy and effectiveness of the thermal management strategy. The intervention mechanism refers to a series of operations automatically initiated by the system to alleviate hotspot problems when the hardware's thermal capacity margin falls below the thermal capacity threshold. This may include steps such as creating localized heat dissipation bottlenecks and adjusting the processor's operating state. Its main purpose is to take timely measures to protect the hardware and optimize heat dissipation efficiency when potential thermal damage risks are detected. A hotspot area refers to a specific physical location on the liquid-cooled cold plate inside the server where the local power density is excessively high due to specific computing tasks or instruction execution modes, resulting in a significantly higher temperature than the surrounding area. It can be identified using methods such as internal processor temperature sensors, power consumption monitoring units, or instruction execution flow analysis. Its main purpose is to accurately pinpoint the target area requiring heat dissipation intervention, avoiding unnecessary global cooling of the entire system. A localized heat dissipation bottleneck refers to an artificially created state of limited heat dissipation in a hotspot area, designed to cause heat accumulation and thus increase the local temperature. This can be achieved by adjusting the coolant flow through the hotspot area, changing the local pressure, or intervening in the performance of specific computing units through software instructions. Its main purpose is to control the rise in local temperature, triggering the processor's own temperature protection mechanism, thereby indirectly guiding the processor to reduce power consumption. The internal temperature protection threshold refers to the upper limit preset within the processor; when the core temperature reaches this value, the processor will automatically trigger protective measures such as frequency reduction or power reduction.The thermal management system (TMS) can be configured by the processor manufacturer or adjusted according to system operating requirements. Its primary function is to utilize the processor's inherent thermal protection mechanisms to indirectly control hotspot areas and prevent hardware overheating damage. The upper-layer workload scheduler refers to a software module running above the operating system or virtualization layer, responsible for allocating computing tasks to different processor cores or computing resources. It can employ scheduling strategies based on resource utilization, task priority, or latency requirements. Its main purpose is to receive feedback from the thermal management system and adjust its resource allocation strategy to avoid assigning high-load tasks to processors in a hotspot state, thereby achieving system-level collaborative optimization.
[0093] In some preferred embodiments, this application is implemented as follows. In a liquid-cooled server deployed within a green data center, the coolant outlet temperature corresponding to the liquid cooling circuit within the server is continuously monitored. This can be achieved by installing a high-precision thermistor sensor array at the outlet of the liquid cooling circuit. The sensor array transmits real-time temperature data to a data acquisition unit. The data acquisition unit samples and analyzes this temperature data, identifying the fluctuation characteristics of the coolant outlet temperature, for example, by calculating the rate of change, peak amplitude, and fluctuation frequency of temperature within a millisecond-level time window. Based on these fluctuation characteristics, a thermal management algorithm module embedded in the server management controller (BMC) evaluates the cumulative heat load of the critical cold plate area of the liquid cooling circuit. For example, when a sharp and frequent instantaneous increase in coolant outlet temperature is detected within a short period, the algorithm module identifies it as a thermal event and quantifies its contribution to the thermal stress of the critical cold plate area based on its intensity. These thermal stress contributions are accumulated to form the cumulative heat load. Subsequently, the module compares the cumulative heat load with a preset thermal damage threshold for the cold plate material to obtain the hardware thermal capacity margin of the critical cold plate area. Next, the thermal management algorithm module compares the calculated hardware thermal capacity margin with a preset thermal capacity threshold. For example, if the preset threshold is 80% of the cold plate's thermal capacity, the system determines that an intervention mechanism needs to be triggered when the hardware thermal capacity margin falls below this threshold. When the intervention determination indicates that the intervention mechanism has been triggered, the system creates a localized heat dissipation bottleneck in the hotspot area. Specifically, the thermal management algorithm module sends instructions to the micro-pumps or micro-valves in the liquid cooling circuit, for example, by reducing the flow rate of coolant through a specific cold plate area, thereby physically limiting heat dissipation in that area. Through the created localized heat dissipation bottleneck, the processor core temperature in the hotspot area gradually rises. When the processor core temperature reaches a preset internal temperature protection threshold, the processor's internal Dynamic Voltage Frequency Scaling (DVFS) mechanism is activated, causing the corresponding processor to automatically reduce its operating frequency or power consumption. Based on the physical signal feedback after the processor reduces its operating frequency or power consumption, the thermal management algorithm module controls the upper-level workload scheduler to adjust the resource allocation strategy. For example, the system can send a prompt to the operating system kernel scheduler, instructing it to avoid assigning new computationally intensive tasks to currently downclocked processor cores, or to migrate some tasks to other processor cores with lower loads. Simultaneously, the system continuously monitors the recovery progress of hardware thermal capacity margin. As processor power consumption decreases, the thermal load in hotspot areas gradually lessens, and hardware thermal capacity margin begins to recover. When monitoring data shows that hardware thermal capacity margin has recovered to a safe level, the system will stop creating localized heat dissipation bottlenecks; for example, it will restore the normal flow of coolant, allowing the heat dissipation in hotspot areas to return to normal.
[0094] Through the above technical solution, this application can effectively solve the hotspot problem caused by server load fluctuations in a multi-tenant environment in the liquid cooling thermal management of green data centers. By continuously monitoring the fluctuation characteristics of the coolant outlet temperature, the system can identify and assess the cumulative heat load and hardware thermal capacity margin of key areas of the cold plate in real time and accurately, thereby avoiding the lag and inaccuracy of traditional average temperature monitoring. When a potential thermal damage risk is detected, a local heat dissipation bottleneck is created in the hotspot area, and the processor's own internal temperature protection mechanism is cleverly utilized to drive the processor to reduce its operating frequency or power consumption. This local, on-demand intervention method avoids ineffective and energy-intensive overall cooling of the entire liquid cooling loop, significantly reducing the data center's operating energy consumption. In addition, the physical signal of the processor reducing its operating frequency or power consumption is used as feedback to control the upper-level workload scheduler to adjust the resource allocation strategy, realizing the collaborative optimization between the thermal management system and the upper-level software scheduler, making resource allocation more reasonable, and further improving the overall energy efficiency ratio and stability of the system. Ultimately, by continuously monitoring the recovery progress of the hardware thermal capacity margin and stopping intervention in a timely manner, the accuracy and timeliness of thermal management measures were ensured, effectively alleviating the differential thermal stress on the cold plate, extending the hardware lifespan, and ensuring the timely recovery of system performance.
[0095] Optional, combined Figure 2 As shown, when the intervention judgment result indicates that the intervention mechanism has been triggered, the steps in S4 to create a local heat dissipation bottleneck in the hotspot area include:
[0096] S41 monitors the instruction execution flow inside the processor and identifies the execution mode of a specific instruction set;
[0097] S42 compares the execution mode of a specific instruction set with a dynamic threshold to determine whether a high power density region is formed, and obtains the region determination result;
[0098] S43, when the region judgment result indicates the formation of a high power density region, the corresponding high power density region is recorded as a hot spot region, and targeted performance intervention is performed on specific computing units inside the processor to create local heat dissipation bottlenecks.
[0099] S44 monitors the temperature or power density of a specific computing unit and cancels the targeted performance intervention once the temperature or power density of the specific computing unit returns to the normal range.
[0100] Instruction execution flow refers to the sequence and state of instructions flowing within the processor during program execution. This can be obtained using techniques such as hardware performance counters, instruction trace units, or software instrumentation, with the aim of capturing the processor's microscopic operating state. Specific instruction sets refer to sets of instructions with specific computational characteristics or resource consumption patterns, such as floating-point instructions, vector processing instructions, or memory access instructions. These can be predefined or dynamically identified based on processor architecture and application load characteristics, with the aim of identifying computational task types that may lead to localized heat concentration. Dynamic thresholds are judgment criteria that are adjusted in real-time based on system operating status, historical data, or environmental conditions. They can be generated using machine learning algorithms, adaptive control logic, or statistical analysis methods, with the aim of improving the accuracy and adaptability of hotspot region identification. High power density... A region refers to an area within the processor where power consumption is concentrated per unit area. This can manifest as a significant increase in local temperature or abnormal readings from power sensors. The purpose is to accurately identify the physical location of the heat concentration. A specific computing unit refers to a logical or physical unit within the processor that performs a specific computing task, such as one or more CPU cores, GPU computing units, memory controllers, or cache modules. It can be determined based on the processor architecture and the identification results of hotspot regions. The purpose is to provide fine-grained intervention in hotspot regions. Targeted performance intervention refers to performance adjustment measures taken for specific computing units within the processor. These measures may include, but are not limited to, reducing the clock frequency of the computing unit, reducing its operating voltage, adjusting its cache strategy, or limiting its instruction throughput. The purpose is to reduce the power consumption and temperature of hotspot regions without affecting the performance of other computing units.
[0101] In some preferred embodiments, this application is implemented as follows: In a multi-core processor system, a hardware performance monitoring unit can be deployed, which can capture the instruction execution flow of each processor core or integrated graphics processing unit (GPU) in real time. For example, it can monitor the instruction throughput of the floating-point unit (FPU), the frequency of memory access instructions, and the execution count of specific vector instructions. This data is transmitted to an embedded controller or management chip. The controller internally runs a hotspot identification algorithm, which compares the collected instruction execution pattern data with a pre-set dynamic threshold. This dynamic threshold can be adjusted in real time according to the overall load of the current server, the ambient temperature, and historical hotspot data. For example, the threshold can be set more leniently when the system load is low, and tightened when the load is high or the ambient temperature rises. When the algorithm determines that the instruction execution pattern of a certain core or GPU region exceeds the dynamic threshold, indicating that the region has formed a high power density region, the region is recorded as a hotspot region. Subsequently, the embedded controller immediately sends a performance intervention command to the specific computing unit corresponding to the hotspot region. For example, if the hotspot is located in a CPU core, the controller can send a command to the power management unit of that core to reduce its clock frequency or core voltage. If the hotspot is located on the integrated GPU, its rendering frequency or the voltage of the compute unit can be reduced. While implementing performance intervention, the controller continuously monitors the temperature sensor data or power density readings of the specific compute unit. Once the temperature or power density of the compute unit falls back to a preset safe range, the controller immediately retracts the previous performance intervention command, restoring the compute unit to its original or higher performance state, ensuring that system performance recovers promptly after the hotspot subsides.
[0102] Optional, combined Figure 3 As shown, the steps of S1 to continuously monitor the coolant outlet temperature corresponding to the liquid cooling circuit in the server and obtain the fluctuation characteristics of the coolant outlet temperature include:
[0103] S11, continuously monitor the coolant outlet temperature corresponding to the liquid cooling circuit in the server, and obtain the fluctuation characteristics of the coolant outlet temperature.
[0104] S12, continuously monitor the inlet and outlet pressures of the coolant flowing through the critical area of the cold plate to obtain the coolant pressure difference in the critical area of the cold plate;
[0105] S13, Based on the changing trend of the coolant pressure difference, determine whether the internal structure of the key area of the cold plate has changed, and obtain the internal change judgment result;
[0106] S14, When the internal change judgment result indicates a change, adjust the mapping relationship between the fluctuation characteristics and the cumulative heat load.
[0107] The coolant pressure difference refers to the pressure difference between the inlet and outlet of the coolant flowing through the critical area of the cold plate. This can be measured and calculated in real time by installing pressure sensors at the inlet and outlet of the critical area. Its purpose is to directly reflect the resistance encountered by the coolant as it flows inside the cold plate, thus indirectly indicating the health of the internal microchannels and other structures. Determining whether the internal structure of the critical area of the cold plate has changed involves analyzing the trend of the coolant pressure difference over time to identify any abnormalities in the internal physical structure. Specifically, this can be achieved by setting a baseline value and a threshold for the pressure difference. When the rate of increase, duration, or absolute value of the pressure difference exceeds the preset threshold, it is considered that the internal structure may have undergone deformation, blockage, or microcracks. This aims to provide a corrective basis for the hardware health status in subsequent heat load assessments, ensuring the accuracy of the assessment. Adjusting the mapping relationship between fluctuation characteristics and cumulative heat load refers to dynamically correcting the algorithm or model used to convert coolant outlet temperature fluctuation characteristics into cumulative heat load based on the judgment results of changes in the internal structure of the cold plate. This can be achieved by modifying the coefficients and weights in the mapping function, or by switching to different assessment models. For example, when it is determined that the internal structure of the cold plate has deteriorated, even if the temperature fluctuation characteristics remain unchanged, the weighting factor in the mapping relationship can be increased to correspondingly increase the assessed cumulative heat load value. The purpose is to compensate for the decrease in heat dissipation efficiency caused by changes in the internal structure of the cold plate, so that the assessment of the cumulative heat load is closer to the actual thermal stress state of the cold plate, avoiding misjudgment or omission of potential hardware damage risks.
[0108] In some preferred embodiments, to continuously monitor the inlet and outlet pressures of the coolant flowing through the critical areas of the cold plate, miniature pressure sensors, such as MEMS (Micro-Electro-Mechanical Systems) pressure sensors, can be installed at the coolant inlet and outlet pipes of the cold plate in the liquid cooling circuit of each server. These sensors can acquire pressure data at a frequency of 10 times per second or higher and transmit the data to the baseboard management controller (BMC) or a dedicated data acquisition unit inside the server. The BMC can calculate the difference between the inlet and outlet pressures, i.e., the coolant pressure difference, in real time and store and analyze it as time-series data. Specifically, to determine whether the internal structure of the critical areas of the cold plate has changed based on the trend of the coolant pressure difference, a reference pressure difference range can be set, which is obtained through calibration under normal operating conditions of the cold plate. The system can continuously monitor the average value and standard deviation of the coolant pressure difference. When the average value of the pressure difference continuously exceeds the upper limit of the reference range for a period of time (e.g., continuously for 24 hours), or its rate of increase exceeds a preset threshold (e.g., an increase of 0.5 kPa per hour), it can be determined that the internal structure of the cold plate may have changed. In addition, machine learning models can be used to more intelligently identify internal structural changes by training on normal and abnormal pressure difference patterns in historical data. When the confidence level of the model output reaches a preset threshold, the internal change is considered to have occurred. As a specific implementation, when the internal change judgment indicates a change in the internal structure of the cold plate, the mapping relationship between the fluctuation characteristics and the cumulative heat load can be adjusted. For example, if the original mapping relationship is linear, i.e., cumulative heat load = K * temperature fluctuation characteristic intensity, where K is a constant, when the internal structure of the cold plate is judged to be deteriorated, the value of K can be increased by a preset increment (e.g., 10%), or a new mapping function with a higher slope can be switched. In this way, even if the temperature fluctuation characteristic intensity remains unchanged, the assessed cumulative heat load will increase accordingly, thus more accurately reflecting the actual thermal stress of the cold plate in the deteriorated state. This adjustment can ensure that when the heat dissipation capacity of the cold plate decreases, the system can more sensitively identify potential overheating risks and trigger intervention mechanisms in a timely manner.
[0109] Optional, combined Figure 4 As shown, the steps by which S6 controls the upper-layer workload scheduler to adjust the resource allocation strategy based on feedback after driving the corresponding processor to reduce its operating frequency or power consumption include:
[0110] S61 continuously monitors the resource allocation behavior corresponding to the physical signals from the upper-layer workload scheduler that reduce the processor's operating frequency or power consumption.
[0111] S62, based on resource allocation behavior, evaluate the effectiveness of physical signals in guiding the behavior of the upper-layer workload scheduler to obtain the guidance effectiveness;
[0112] S63, when bootstrapping effectiveness declines, identify adaptive strategies that the upper-layer workload scheduler has developed;
[0113] S64 adjusts the parameters of local heat dissipation bottlenecks according to an adaptive strategy;
[0114] S65 creates local heat dissipation bottlenecks based on the adjusted parameters of the local heat dissipation bottlenecks, causing the processor core temperature in the hot area to rise to the preset internal temperature protection threshold, thereby driving the corresponding processor to reduce its operating frequency or power consumption.
[0115] S66 uses physical signals that reduce the processor's operating frequency or power consumption as feedback to control the upper-layer workload scheduler to adjust resource allocation strategies.
[0116] Physical signals refer to the performance or state changes that occur within the system when the processor reduces its operating frequency or power consumption due to increased temperature. These changes can be perceived by the upper-layer workload scheduler. Specifically, they can manifest as an actual decrease in processor core frequency, a reduction in power consumption, a decrease in instruction throughput, or an increase in task completion latency. The purpose is to convey implicit information about the processor's current thermal state to the upper-layer workload scheduler, prompting it to adjust resource allocation. Resource allocation behavior refers to the specific operations performed by the upper-layer workload scheduler after receiving the physical signals, including allocating and scheduling the computing resources it manages (such as CPU cores, memory, network bandwidth, etc.). This can include migrating tasks from overheated processor cores to other cores, reducing the priority of specific tasks, pausing some non-critical tasks, or adjusting the parallelism of tasks. The aim is to respond to the processor's thermal state by changing the distribution of workloads, thereby preventing further overheating or restoring performance. Guiding effectiveness refers to the degree to which physical signals have the expected impact on the resource allocation behavior of the upper-level workload scheduler. Specifically, it can be assessed by comparing the changes in the scheduler's resource allocation pattern before and after receiving the signal with whether they align with the expected thermal management goals. Its purpose is to quantify the control effect of thermal management strategies on the behavior of the upper-level scheduler, providing a basis for subsequent strategy adjustments. Adaptive strategies refer to new resource allocation or task scheduling patterns developed by the upper-level workload scheduler after receiving physical signals for an extended period, in order to avoid or offset the impact of these signals on its performance goals. Specifically, this may include learning and predicting the timing of thermal management interventions, adjusting task allocation logic to avoid triggering bottlenecks, or prioritizing resources for critical tasks when performance degrades. Its purpose is to describe the avoidance behaviors adopted by the upper-level scheduler to maintain its own performance goals, representing a challenge that thermal management systems need to address. The parameters of a local heat dissipation bottleneck refer to the adjustable properties of the technical means used to create a local heat dissipation bottleneck. Specifically, these may include the intensity of directional performance intervention, the local flow rate of the coolant, the local pressure, or the flow resistance of the microchannels inside the cold plate. The purpose is to change the intensity and mode of action of the local heat dissipation bottleneck by adjusting these parameters in order to cope with the adaptive strategies developed by the upper-level workload scheduler and maintain the effectiveness of the thermal management strategy.
[0117] Optionally, S62 evaluates the effectiveness of physical signals in guiding the behavior of the upper-layer workload scheduler based on resource allocation behavior. The steps to determine the effectiveness of this guidance include:
[0118] Monitor the adjustment behavior of the upper-layer workload scheduler in allocating computing resources and generate a sequence of adjustment behaviors;
[0119] Based on the adjustment behavior sequence, analyze and identify the response patterns of the upper-layer workload scheduler to physical signals;
[0120] The response pattern is compared with the preset response characteristics to determine whether the response pattern has changed, and the result of the response change judgment is obtained.
[0121] When the response change judgment result indicates a change, the effectiveness evaluation rules for the physical signal to guide the behavior of the upper-layer workload scheduler are redefined based on the changed response pattern.
[0122] Based on the effectiveness evaluation rules, evaluate the effectiveness of physical signals in guiding the behavior of the upper-layer workload scheduler.
[0123] The adjustment behavior sequence refers to a record of a series of dynamic adjustment operations performed by the upper-layer workload scheduler on the allocation of computing resources (such as CPU cores, memory bandwidth, I / O throughput, etc.) after receiving physical signals. Specifically, it can be a time-ordered set of data points consisting of timestamps, resource types, changes in allocation amounts, and operation types (increase, decrease, maintain), aiming to provide a complete historical view of the scheduler's behavior. The response pattern refers to the regularity or tendency of the upper-layer workload scheduler's resource allocation adjustment behavior when facing specific physical signals (such as processor frequency reduction or power reduction). Specifically, it can be identified by statistically analyzing parameters such as frequency, amplitude, delay, and duration in the adjustment behavior sequence, for example, a fast response pattern, a delayed response pattern, an ignored pattern, or an over-response pattern, aiming to reveal the scheduler's internal processing logic for physical signals. The preset response characteristics refer to the expected or baseline response pattern guided by physical signals to the upper-layer workload scheduler's behavior under normal system operation or ideal conditions. Specifically, it can be predefined through historical data analysis, expert experience setting, or simulation, aiming to serve as a reference standard for judging whether the scheduler's behavior has changed. The response change judgment result refers to the conclusion drawn by comparing the currently identified response pattern with the preset response characteristics regarding whether the scheduler's behavior pattern deviates from expectations. Specifically, it can be a Boolean value (such as "changed" or "not changed"), and its purpose is to indicate whether the effectiveness evaluation mechanism needs adjustment. The effectiveness evaluation rule refers to the standard or algorithm used to quantify the effect of physical signals on the behavior guidance of the upper-layer workload scheduler. Specifically, it can be a set of weighted parameters, a machine learning model, or a set of logical judgment conditions, and its purpose is to ensure that the evaluation results accurately reflect the actual guidance capability of the physical signals.
[0124] In some preferred embodiments, this application is implemented as follows: First, to monitor the adjustment behavior of the upper-layer workload scheduler in allocating computing resources and generate an adjustment behavior sequence, a lightweight agent or hook can be deployed in the operating system kernel or virtualization layer. This agent captures system calls from the scheduler in real time, such as CPU core allocation, memory bandwidth limiting, or I / O priority adjustment. Each time an adjustment behavior is captured, the agent records the timestamp, the type of resource affected, the amount of resources before and after the adjustment, and the operation type (e.g., the number of CPU cores is reduced from 8 to 4, or the memory bandwidth is reduced from 10GB / s to 5GB / s), and stores these data points in a circular buffer in chronological order to form an adjustment behavior sequence. Next, based on the adjustment behavior sequence, to analyze and identify the response pattern of the upper-layer workload scheduler to physical signals, a sliding time window can be used to process the sequence data. Within a fixed time window, such as the past 30 seconds, statistical analysis is performed on the resource allocation adjustment behavior of the scheduler after responding to physical signals (such as processor downclocking events). For example, the average response latency, the average magnitude of resource allocation reduction, and the average duration of low resource allocation after the physical signal is emitted can be calculated. These statistics can form a vector representing the current response pattern. Subsequently, to compare the response pattern with preset response features and determine whether the response pattern has changed, a response change judgment result can be obtained by calculating the Euclidean distance or cosine similarity between the currently calculated response pattern vector and one or more preset baseline response pattern vectors. The preset response features can be obtained through training on a large amount of experimental data in the early stages of stable system operation; for example, a "normal response" pattern vector and a "sluggish response" pattern vector. If the distance between the current pattern vector and the "normal response" pattern vector exceeds a certain dynamic threshold, or the similarity with the "sluggish response" pattern vector increases significantly, it can be judged that the response pattern has changed, and the response change judgment result is set to "changed". When the response change judgment result indicates a change, an adaptive learning module can be launched to redetermine the effectiveness evaluation rules for guiding the behavior of the upper-layer workload scheduler based on the changed response pattern. This module can be a reinforcement learning-based agent or a dynamic Bayesian network. It receives the currently identified changed response pattern as input and adjusts the weight parameters or logical decision branches in the evaluation rule according to a predefined policy update algorithm (e.g., Q-learning or policy gradient). For example, if the scheduler is identified as exhibiting a "sluggish response" pattern, the evaluation rule can increase the penalty weight for response delay or introduce a longer observation window to confirm its behavior.Finally, based on the redefined effectiveness evaluation rules, the system will score the scheduler's latest behavior using the updated rules to assess the effectiveness of physical signals in guiding the behavior of the upper-layer workload scheduler. For example, if the rules now emphasize response speed, a fast-responding scheduler will receive a higher effectiveness score. This score can be a value between 0 and 100, where a higher score indicates a better guiding effect of the physical signals. This evaluation result will serve as the basis for subsequent adjustments to local heat dissipation bottleneck parameters, thus forming a closed-loop control system.
[0125] Optionally, when bootstrapping effectiveness declines, the steps in S63 to identify adaptive strategies developed by the upper-layer workload scheduler include:
[0126] To obtain the rate, duration, and magnitude of the decline in guiding effectiveness;
[0127] The rate, duration, and magnitude of the decline in guidance effectiveness are compared with preset thresholds to obtain the guidance effectiveness judgment results;
[0128] Based on the effectiveness of the guidance, identify the adaptive strategies that the upper-layer workload scheduler has developed.
[0129] The rate of decline in guidance effectiveness refers to the rate at which guidance effectiveness changes over time, reflecting how fast it declines; the duration refers to the length of time guidance effectiveness remains in a declining state; and the magnitude refers to the maximum extent or total change in guidance effectiveness from its normal or baseline level. These parameters can be obtained by differential calculation of historical guidance effectiveness data, timestamp recording, and comparison of maximum and minimum values. The preset threshold is a set of reference values pre-set to determine the degree of decline in guidance effectiveness. It can be an independent threshold set for each of the rate, duration, and magnitude, or a composite threshold set for a combination of these parameters. Its purpose is to provide a quantitative standard to distinguish different degrees of decline in guidance effectiveness. The guidance effectiveness judgment result is the conclusion drawn after comparing the rate, duration, and magnitude of the decline in guidance effectiveness with the preset threshold. It can be a classification label indicating the degree of decline in guidance effectiveness, such as "slight decline," "moderate decline," or "significant decline," aiming to provide an accurate basis for subsequently identifying the adaptive strategies of the upper-layer workload scheduler. The adaptive strategies developed by upper-layer workload schedulers refer to the resource allocation or task scheduling behavior patterns that upper-layer workload schedulers autonomously adjust in order to maintain their own workload performance after sensing that the processor performance is limited or the heat dissipation conditions change. These strategies may include adjusting task priorities, changing task concurrency, migrating workloads to other processor cores or nodes, or adjusting their internal performance optimization algorithms. The purpose is to cope with changes in the external environment and ensure the performance of the applications they manage.
[0130] In some preferred embodiments, this application is implemented as follows. Assume that the data center's thermal management system continuously monitors the effectiveness of physical signals indicating that the processor is reducing its operating frequency or power consumption in influencing the resource allocation behavior of the upper-layer workload scheduler. When a decrease in guidance effectiveness is detected, the system can activate a data acquisition module that records the current guidance effectiveness value at a fixed frequency. The rate of decline in guidance effectiveness can be calculated from the continuously recorded values. For example, the instantaneous decline rate can be obtained by calculating the ratio of the difference between the guidance effectiveness values at two adjacent time points to the time interval, and the average decline rate over a period of time can also be calculated. Simultaneously, the system can record the time point at which the decline in guidance effectiveness begins and the current time point to obtain the duration. Furthermore, the system can record the difference between the maximum decline value and the initial value during the decline in guidance effectiveness to determine the magnitude of the decline. For example, if guidance effectiveness declines from a high level to a low level over a period of time, and the decline rate reaches a specific value at a certain moment, the system can compare these acquired speed, duration, and magnitude data with preset thresholds. For example, the preset thresholds can be set as: a speed threshold, a duration threshold, and a magnitude threshold. If the current decline rate, duration, and magnitude all exceed their respective preset thresholds, the system determines that the guidance effectiveness has "significantly declined." Based on this assessment of a "significant decrease," the system can identify that the upper-layer workload scheduler may have developed a resource preemption strategy. For example, the scheduler may be attempting to compensate for the performance degradation by frequently assigning high-priority tasks to the current processor or by increasing task concurrency. Based on this identification, the thermal management system can further adjust the parameters of local heat dissipation bottlenecks to more effectively guide the scheduler to adjust its resource allocation strategy.
[0131] Optionally, the steps by which S64 adjusts the parameters of local heat dissipation bottlenecks according to an adaptive strategy include:
[0132] Adjust the flow rate or local pressure of the coolant flowing through the hot spot area;
[0133] The strength of the local heat dissipation bottleneck is changed according to the adjusted flow rate or local pressure of the coolant, so as to dynamically adjust the parameters for manufacturing the local heat dissipation bottleneck.
[0134] Regulating the flow rate or local pressure of the coolant flowing through the hot spot area refers to controlling the flow state of the coolant in a specific area of the liquid cooling circuit to directly affect the heat removal capacity of that area. This can be achieved using actuators such as micropumps, microvalves, or variable flow pumps. Changing the intensity of the local heat dissipation bottleneck refers to adjusting the flow rate or local pressure of the coolant to controllably change the heat dissipation capacity of the hot spot area, thereby affecting the rate of temperature rise and the final stable temperature of that area. The purpose is to precisely control the processor core temperature to rise to a preset internal temperature protection threshold. Dynamically adjusting the parameters that create the local heat dissipation bottleneck refers to modifying the control variables used to create the local heat dissipation bottleneck in real time or periodically, based on the adaptive strategies developed by the upper-level workload scheduler. These variables include the setpoints for coolant flow rate, local pressure, or the duration of the bottleneck. The purpose is to enable the thermal management system to flexibly respond to changes in the scheduler's behavior and avoid ineffective game theory.
[0135] In some preferred embodiments, this solution is implemented as follows. Assume a liquid-cooled server where multiple independent microchannels are integrated on the processor cold plate, each microchannel corresponding to one or a group of processor cores, and each microchannel inlet is equipped with a programmable micro-flow control valve. When the system identifies that the upper-level workload scheduler has developed an adaptive strategy—for example, the scheduler tends to concentrate high-load tasks on processor cores 1 and 2, causing frequent overheating in these two core areas, and traditional frequency reduction signals are insufficient to effectively guide load distribution—the thermal management system can adjust the parameters of the local heat dissipation bottleneck according to this adaptive strategy. Specifically, the system can send instructions to the controller that controls the micro-flow control valve, for example, reducing the coolant flow rate through the microchannels corresponding to processor cores 1 and 2 from 10 ml / min to 8 ml / min. This reduction in flow rate directly changes the heat dissipation capacity of these two hot areas, thereby strengthening the local heat dissipation bottleneck. The system can dynamically adjust the parameters for creating localized heat dissipation bottlenecks based on the actual situation after coolant flow adjustments. For example, it can adjust the trigger speed of the target temperature protection threshold from 1 degree Celsius per second to 2 degrees Celsius per second, or extend the duration of maintaining the localized heat dissipation bottleneck from 5 seconds to 8 seconds. In this way, the system can more precisely control the temperature rise curve of hotspot areas, thereby more effectively driving the upper-level workload scheduler to reallocate tasks to other processor cores with better heat dissipation conditions, in response to its established adaptive strategy.
[0136] Optional, combined Figure 5 As shown, S2 evaluates the cumulative heat load of the critical area of the cold plate in the liquid cooling circuit based on the fluctuation characteristics, and obtains the hardware heat capacity margin of the critical area of the cold plate. The steps include:
[0137] S21, Identify thermal events in the fluctuation characteristics; thermal events include the instantaneous increase in temperature;
[0138] S22, quantify the contribution of thermal events to the thermal stress in critical areas of the cold plate based on the intensity of the thermal event;
[0139] S23, cumulative heating stress contribution, yields the cumulative heat load;
[0140] S24 compares the accumulated heat load with the preset heat damage threshold of the critical area of the cold plate to obtain the hardware heat capacity margin.
[0141] Among these, a thermal event refers to a specific temperature change pattern within the coolant outlet temperature fluctuation characteristics that indicates localized, instantaneous heat changes in critical areas of the cold plate. This can specifically be a rapid rise, fall, or violent oscillation in temperature. The instantaneous temperature rise refers to the maximum temperature difference that occurs rapidly from a reference value to the coolant outlet temperature within a short period. This can be calculated by differentially analyzing temperature values continuously read by a temperature sensor within a very short sampling period, or by setting a temperature change rate threshold. Thermal stress contribution refers to the degree of stress influence generated within the cold plate material by a specific thermal event. Specifically, this can be achieved by mapping temperature change characteristics (such as the instantaneous temperature rise and duration) into a quantitative indicator of fatigue damage or deformation of the cold plate material through a pre-established physical model, simulation data, or experimental results. The preset thermal damage threshold refers to the critical value at which the accumulated thermal stress in the critical areas of the cold plate reaches or exceeds a certain value during long-term operation, at which point the cold plate material may experience irreversible fatigue damage, performance degradation, or structural failure. This threshold can be set based on the mechanical properties of the cold plate material, its fatigue life curve, and the reliability requirements in practical applications.
[0142] In some preferred embodiments, this application is implemented as follows: First, to identify thermal events in the fluctuation characteristics, a high-precision temperature sensor can be used to sample the coolant outlet temperature at high frequency, for example, 100 times per second. Then, digital signal processing techniques, such as applying a moving average filter, are used to smooth the data and calculate the rate of temperature change between consecutive sampling points. When the rate of temperature change exceeds a preset instantaneous rise threshold and the duration exceeds a time window on the order of microseconds, it can be identified as a thermal event, and its instantaneous temperature rise amplitude is recorded, for example, a rise from a reference temperature of 25°C to 30°C instantaneously is recorded as an amplitude of 5°C. Next, to quantify the contribution of the thermal event to the thermal stress of the critical area of the cold plate according to the intensity of the thermal event, a thermal stress model can be pre-established. This model can be obtained through finite element analysis or experimental testing based on the thermal expansion coefficient, elastic modulus, and fatigue life curve of the cold plate material. For example, for a 5°C instantaneous temperature rise, the model might calculate an instantaneous thermal stress of 0.1 MPa on the cold plate material. Based on this stress value and duration, a fatigue damage accumulation algorithm (e.g., a simplified model based on Miner's linear cumulative damage theory) converts it into a thermal stress contribution value, such as 0.001 units of fatigue damage. The system then continuously accumulates these thermal stress contributions to obtain the cumulative heat load. For example, for each identified thermal event, its corresponding thermal stress contribution value is added to the total cumulative heat load. This accumulation process can be a simple summation or a more advanced accumulation algorithm that considers time decay or historical weighting. Finally, the cumulative heat load is compared with a preset thermal damage threshold for critical areas of the cold plate to obtain the hardware thermal capacity margin. The preset thermal damage threshold can be set according to the cold plate's design life and reliability requirements; for example, when the cumulative heat load reaches 100 units of fatigue damage, the cold plate is considered to have reached the end of its design life. By comparing the current cumulative heat load to this threshold—for example, if the current cumulative heat load is 50 units and the threshold is 100 units—the hardware thermal capacity margin can be expressed as 50 units, or as a percentage of 50%. This margin value intuitively reflects the health status and remaining lifespan of the cold plate, providing a basis for decision-making in the data center's thermal management system.
[0143] Optionally, the step of targeted performance intervention on specific computing units within the processor in step S43 includes:
[0144] Get the current temperature of a specific computing unit inside the processor;
[0145] Determine the target performance state of a specific computing unit based on the current temperature;
[0146] Based on the target performance state, the clock frequency or operating voltage of a specific computing unit can be adjusted to achieve targeted performance intervention for that specific computing unit.
[0147] In this context, a specific computing unit refers to a logical functional module within the processor capable of independently performing calculations or executing specific tasks. Specifically, it can be one or more CPU cores, stream processors in a graphics processing unit (GPU), computing arrays in a neural network processing unit (NPU), or other dedicated accelerator modules. Its purpose is to accurately identify and intervene in local hotspots within the processor. Current temperature refers to the real-time thermal state of a specific computing unit at a given moment. This can be measured and acquired in real-time using temperature sensors or thermistor arrays integrated within the processor. Its purpose is to provide real-time and accurate thermal load data for subsequent performance interventions. Target performance state refers to a desired operating performance level set for a specific computing unit. This can be a predefined performance level, a power consumption limit, or a clock frequency range. Its purpose is to maintain the performance output of the computing unit as much as possible while meeting heat dissipation requirements. Targeted performance intervention refers to adjusting the operating parameters of a specific computing unit within the processor to influence its performance and power consumption, thereby controlling local hotspots. This can be achieved by reducing the clock frequency, reducing the operating voltage, or a combination of both.
[0148] In some preferred embodiments, targeted performance intervention for specific computing units within the processor can be implemented as follows: Assuming the processor contains multiple CPU cores, each CPU core is considered a specific computing unit. The system can utilize a digital temperature sensor (DTS) integrated within each CPU core to acquire the current temperature of that CPU core in real time. For example, when the current temperature of a CPU core consistently exceeds a preset warning temperature, the system dynamically calculates the target performance state of that CPU core based on a pre-established temperature-performance mapping table or algorithm. This mapping table can be defined as follows: when the temperature reaches T1, the target performance state is P1 (e.g., a reduction of one performance level); when the temperature reaches T2, the target performance state is P2 (e.g., a reduction of two performance levels). Once the target performance state is determined, the system sends instructions to the specific CPU core through the processor's power management unit (PMU) or clock generation unit (CGU) interface to adjust its clock frequency or operating voltage. For example, if the target performance level is to be reduced by one performance level, the PMU might reduce the CPU core's clock frequency from 3.0GHz to 2.5GHz, while simultaneously reducing the operating voltage from 1.0V to 0.9V. This adjustment is made for a single CPU core and does not affect the normal operation of other non-hotspot CPU cores, thus achieving precise, localized performance intervention in hotspot areas and effectively creating localized heat dissipation bottlenecks.
[0149] A green data center liquid cooling and thermal management system is used to perform green data center liquid cooling and thermal management, combined with Figure 6 As shown, the green data center liquid-cooled thermal management system 1 includes:
[0150] The fluctuation feature acquisition module 11 is used to continuously monitor the coolant outlet temperature corresponding to the liquid cooling circuit in the server and acquire the fluctuation features of the coolant outlet temperature.
[0151] The capacity margin assessment module 12 is used to assess the cumulative heat load of the critical area of the cold plate in the liquid cooling circuit based on the fluctuation characteristics, and to obtain the hardware heat capacity margin of the critical area of the cold plate.
[0152] The intervention mechanism judgment module 13 is used to compare the hardware thermal capacity margin with the preset thermal capacity threshold to determine whether the intervention mechanism is triggered and to obtain the intervention judgment result.
[0153] The heat dissipation bottleneck manufacturing module 14 is used to create a local heat dissipation bottleneck in the hot spot area when the intervention judgment result indicates that the intervention mechanism has been triggered.
[0154] The bottleneck strategy execution module 15 is used to make the processor core temperature in the hot spot area rise to a preset internal temperature protection threshold by using local heat dissipation bottlenecks, thereby driving the corresponding processor to reduce its operating frequency or power consumption.
[0155] The allocation strategy adjustment module 16 is used to control the upper-layer workload scheduler to adjust the resource allocation strategy based on the feedback after driving the corresponding processor to reduce its operating frequency or power consumption.
[0156] The bottleneck recovery management module 17 is used to continuously monitor the recovery progress of the hardware thermal capacity margin. When the hardware thermal capacity margin is fully recovered, it stops creating local heat dissipation bottlenecks.
[0157] The fluctuation characteristic acquisition module is a component used to continuously collect and analyze the coolant outlet temperature data of the liquid cooling circuit within the server. It can be implemented using a microcontroller integrated on the server motherboard or a separate sensor data acquisition unit. Its purpose is to capture the instantaneous changes and trends in coolant temperature in real time, providing basic data for subsequent thermal management decisions. The capacity margin assessment module is a component used to calculate and quantify the cumulative heat load borne by the critical area of the cold plate in the liquid cooling circuit based on temperature fluctuation characteristics, thereby assessing the remaining heat capacity of the hardware in that area. It can be implemented using a software algorithm module running on the Server Management Controller (BMC) or Data Center Infrastructure Management (DCIM) system. The key is to transform physical temperature changes into quantitative indicators of hardware health status, providing a basis for preventative maintenance and intervention. The intervention mechanism judgment module is a decision-making component that compares the assessed hardware thermal capacity margin with a preset safety threshold and determines whether thermal management intervention measures need to be initiated. It can be implemented using embedded logic units or rule-based expert systems, aiming to ensure timely intervention when hardware faces potential risks while avoiding unnecessary performance limitations. The heat dissipation bottleneck creation module is a component used to locally limit heat dissipation in specific hotspot areas inside the server through physical or logical means. It can employ programmable microfluidic valves, local heating elements, or software-based methods. The processor's internal cooling strategy is controlled to precisely increase the temperature of hot spots, triggering the processor's own temperature protection mechanism. The bottleneck strategy execution module monitors the rise in processor core temperature in hot spots after a localized cooling bottleneck is created, ensuring it reaches a preset internal temperature protection threshold. This prompts the processor to automatically reduce its operating frequency or power consumption. It can be implemented using a software agent that interacts with the processor power management unit (PMU) or the operating system power management interface (OSPM). Its purpose is to reduce heat generation and alleviate localized overheating through the processor's adaptive mechanism. The allocation strategy adjustment module is used to adjust the processor's operating frequency or power consumption when it is reduced due to temperature protection. After performance degradation, based on system feedback, the component communicates with the upper-layer workload scheduler to guide it to reallocate computing resources. This can be achieved through software services that interact with the upper-layer scheduler using API interfaces or message queue mechanisms. Its purpose is to optimize resource utilization at a macro level and avoid continuous overload in hotspot areas. The bottleneck recovery management module is a component used to continuously monitor the recovery status of hardware thermal capacity margin and promptly remove local heat dissipation bottlenecks after the hardware risk is eliminated. It can be achieved by using scheduled tasks or event-driven mechanisms to periodically check the hardware status and send removal commands. Its purpose is to ensure that the system can quickly restore normal performance after the risk is eliminated and avoid long-term unnecessary performance limitations.
[0158] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A green data center liquid cooling thermal management method, characterized in that, include: Continuously monitor the coolant outlet temperature corresponding to the liquid cooling circuit inside the server to obtain the fluctuation characteristics of the coolant outlet temperature; Based on the fluctuation characteristics, the cumulative heat load of the critical area of the cold plate in the liquid cooling circuit is evaluated to obtain the hardware heat capacity margin of the critical area of the cold plate. The hardware thermal capacity margin is compared with a preset thermal capacity threshold to determine whether an intervention mechanism is triggered, and an intervention judgment result is obtained. When the intervention judgment result indicates that the intervention mechanism has been triggered, a local heat dissipation bottleneck is created in the hot spot area; By addressing the localized heat dissipation bottleneck, the processor core temperature in the hotspot area rises to a preset internal temperature protection threshold, thereby causing the corresponding processor to reduce its operating frequency or power consumption. Based on the feedback after driving the corresponding processor to reduce its operating frequency or power consumption, the upper-layer workload scheduler is controlled to adjust the resource allocation strategy. Continuously monitor the recovery progress of the hardware thermal capacity margin, and stop creating the local heat dissipation bottleneck when the hardware thermal capacity margin is fully restored.
2. The green data center liquid cooling thermal management method according to claim 1, characterized in that, The step of creating a local heat dissipation bottleneck in the hotspot area when the intervention judgment result indicates that the intervention mechanism has been triggered includes: Monitor the instruction execution flow inside the processor and identify the execution mode of a specific instruction set; the specific instruction set refers to a set of instructions with specific computing characteristics or resource consumption patterns, which are predefined or dynamically identified based on the processor architecture and application load characteristics. The execution mode of a specific instruction set is compared with a dynamic threshold to determine whether a high power density region is formed, thus obtaining the region determination result; When the region determination result indicates the formation of a high power density region, the corresponding high power density region is recorded as a hot spot region, and targeted performance intervention is performed on specific computing units inside the processor to create local heat dissipation bottlenecks; the specific computing unit refers to the logical or physical unit inside the processor that performs a specific computing task, which is determined based on the processor architecture and the identification result of the hot spot region. Monitor the temperature or power density of the specific computing unit, and withdraw the targeted performance intervention after the temperature or power density of the specific computing unit returns to the normal range.
3. The green data center liquid cooling thermal management method according to claim 1, characterized in that, The step of continuously monitoring the coolant outlet temperature corresponding to the liquid cooling circuit within the server and obtaining the fluctuation characteristics of the coolant outlet temperature includes: Continuously monitor the coolant outlet temperature corresponding to the liquid cooling circuit inside the server to obtain the fluctuation characteristics of the coolant outlet temperature; The inlet and outlet pressures of the coolant flowing through the critical area of the cold plate are continuously monitored to obtain the coolant pressure difference in the critical area of the cold plate. Based on the changing trend of the coolant pressure difference, it is determined whether the internal structure of the key area of the cold plate has changed, and the result of the internal change judgment is obtained. When the internal change judgment result indicates a change, the mapping relationship between the fluctuation characteristics and the cumulative heat load is adjusted.
4. The green data center liquid cooling thermal management method according to claim 1, characterized in that, The step of controlling the upper-layer workload scheduler to adjust the resource allocation strategy based on feedback after driving the corresponding processor to reduce its operating frequency or power consumption includes: Continuously monitor the resource allocation behavior corresponding to the physical signals from the upper-layer workload scheduler that reduce the processor's operating frequency or power consumption; Based on the resource allocation behavior, the effectiveness of the physical signals in guiding the behavior of the upper-layer workload scheduler is evaluated to obtain the guidance effectiveness; When bootstrapping effectiveness declines, identify the adaptive strategies that upper-layer workload schedulers have developed; Based on the adaptive strategy, adjust the parameters of the local heat dissipation bottleneck; Based on the adjusted parameters of the local heat dissipation bottleneck, a local heat dissipation bottleneck is created, causing the processor core temperature in the hot spot area to rise to a preset internal temperature protection threshold, thereby driving the corresponding processor to reduce its operating frequency or power consumption. The physical signals that cause the processor to reduce its operating frequency or power consumption are used as feedback to control the upper-layer workload scheduler to adjust the resource allocation strategy.
5. A green data center liquid cooling thermal management method according to claim 4, characterized in that, The step of evaluating the effectiveness of the physical signals in guiding the behavior of the upper-layer workload scheduler based on the resource allocation behavior, and obtaining the guiding effectiveness, includes: Monitor the adjustment behavior of the upper-layer workload scheduler in allocating computing resources and generate a sequence of adjustment behaviors; Based on the adjusted behavior sequence, analyze and identify the response pattern of the upper-layer workload scheduler to the physical signal; The response pattern is compared with a preset response feature to determine whether the response pattern has changed, and a response change determination result is obtained. When the response change judgment result indicates a change, the effectiveness evaluation rules for the physical signal in guiding the behavior of the upper-layer workload scheduler are redefined based on the changed response pattern. Based on the aforementioned effectiveness evaluation rules, evaluate the effectiveness of the physical signals in guiding the behavior of the upper-layer workload scheduler.
6. The green data center liquid cooling thermal management method according to claim 4, characterized in that, The step of identifying adaptive strategies developed by the upper-layer workload scheduler when bootstrapping effectiveness decreases includes: To obtain the rate, duration, and magnitude of the decline in guiding effectiveness; The rate, duration, and magnitude of the decline in guidance effectiveness are compared with preset thresholds to obtain the guidance effectiveness judgment results; Based on the effectiveness of the guidance, identify the adaptive strategies that the upper-layer workload scheduler has developed.
7. A green data center liquid cooling thermal management method according to claim 4, characterized in that, The step of adjusting the parameters of the local heat dissipation bottleneck according to the adaptive strategy includes: Adjust the flow rate or local pressure of the coolant flowing through the hot spot area; The strength of the local heat dissipation bottleneck is changed according to the adjusted flow rate or local pressure of the coolant, so as to dynamically adjust the parameters for manufacturing the local heat dissipation bottleneck.
8. A green data center liquid cooling thermal management method according to claim 1, characterized in that, The step of evaluating the cumulative heat load of the critical area of the cold plate in the liquid cooling circuit based on the fluctuation characteristics, and obtaining the hardware heat capacity margin of the critical area of the cold plate, includes: Identify thermal events within the fluctuation characteristics; the thermal events include the instantaneous increase in temperature. Based on the intensity of the thermal event, quantify the contribution of the thermal event to the thermal stress in the critical area of the cold plate; The cumulative thermal load is obtained by summing the thermal stress contributions. The accumulated heat load is compared with the preset thermal damage threshold of the critical area of the cold plate to obtain the hardware thermal capacity margin.
9. A green data center liquid cooling thermal management method according to claim 2, characterized in that, The steps of targeted performance intervention on specific computing units within the processor include: Get the current temperature of a specific computing unit inside the processor; Based on the current temperature, determine the target performance state of the specific computing unit; Based on the target performance state, the clock frequency of the specific computing unit or the operating voltage of the specific computing unit is adjusted to achieve targeted performance intervention of the specific computing unit.
10. A green data center liquid cooling and thermal management system, used to perform green data center liquid cooling and thermal management, characterized in that, include: The fluctuation feature acquisition module is used to continuously monitor the coolant outlet temperature corresponding to the liquid cooling circuit in the server and acquire the fluctuation feature of the coolant outlet temperature. The capacity margin assessment module is used to assess the cumulative heat load of the critical area of the cold plate in the liquid cooling circuit based on the fluctuation characteristics, and to obtain the hardware heat capacity margin of the critical area of the cold plate. The intervention mechanism judgment module is used to compare the hardware thermal capacity margin with a preset thermal capacity threshold, determine whether to trigger the intervention mechanism, and obtain the intervention judgment result. The heat dissipation bottleneck manufacturing module is used to create local heat dissipation bottlenecks in hot spots when the intervention judgment result indicates that the intervention mechanism has been triggered. The bottleneck strategy execution module is used to raise the processor core temperature in the hot spot area to a preset internal temperature protection threshold through the local heat dissipation bottleneck, thereby driving the corresponding processor to reduce its operating frequency or power consumption. The allocation strategy adjustment module is used to control the upper-layer workload scheduler to adjust the resource allocation strategy based on the feedback after driving the corresponding processor to reduce its operating frequency or power consumption. The bottleneck recovery management module is used to continuously monitor the recovery progress of the hardware thermal capacity margin. When the hardware thermal capacity margin is fully recovered, the creation of the local heat dissipation bottleneck is stopped.
Citation Information
Patent Citations
Cold plate temperature control method and device, electronic equipment and computer readable medium
CN112506254A
Inverter cooling fin temperature control method and system
CN119847233A