Temperature regulation and control method for high-performance computer system
By dividing the heat zones in a high-performance computing system and building a dynamic heat capacity model, combining distributed optimization and dual-ring control, the problems of dynamic load fluctuations and multi-heat source coupling are solved, and the precise temperature regulation and system stability are achieved.
Patent Information
- Application Number
- CN202510725080.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-08-15
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art is difficult to adapt to the transient thermal shock caused by dynamic load fluctuations and multi-heat source coupling in high-performance computing systems, resulting in accumulation of temperature prediction deviations, local heat accumulation and control accuracy reduction, and lack of a dynamic compensation mechanism for model prediction errors, affecting system stability and energy efficiency.
The system is divided into multiple independent heat zones, temperature, power consumption and process data are collected in real time, dynamic heat capacity model is constructed and heat capacity parameters and heat dissipation coefficients are updated online, combined with a distributed optimization model of thermal coupling effect, and the strategy deviation is corrected through dual-ring control, and the data update cycle is triggered under burst processes or temperature overlimit events.
Real-time tracking of sudden processor power consumption is achieved, temperature stability and noise suppression is balanced, heat accumulation is avoided, stable system operation and heat dissipation efficiency is ensured, control accuracy and energy efficiency are improved.
Smart Images

Figure CN120491789A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer heat dissipation control, and in particular to a temperature control method for a high-performance computer system. Background Art
[0002] As the computing power density of high-performance computing systems continues to increase, thermal management faces severe challenges in dynamic load fluctuations, coupled multiple heat sources, and energy efficiency balancing. Traditional heat dissipation control methods, which often rely on static thermodynamic models and fixed policy weights, are unable to adapt to the transient thermal shock caused by dynamic processor loads and sudden processes, leading to the accumulation of temperature prediction errors.
[0003] At the same time, existing technologies often ignore the coupling effect of thermal diffusion between hot zones when addressing multi-objective optimization, creating a conflict between local heat dissipation strategies and the need for global thermal balance, leading to heat accumulation in adjacent areas. Furthermore, control robustness is insufficient under parameter mismatch or external disturbances, and there is a lack of dynamic compensation mechanisms for model prediction errors, which can easily lead to thermal oscillations and even system frequency reduction protection under extreme operating conditions.
[0004] The above problems cause existing solutions to face challenges such as decreased control accuracy, worsened energy efficiency and insufficient system stability in complex computing scenarios. Summary of the Invention
[0005] In response to the shortcomings of the existing technology, the present invention provides a temperature control method for high-performance computer systems, which solves the problems in the existing technology that static models are difficult to adapt to dynamic load fluctuations, the coupling effect of multiple heat sources leads to local heat accumulation, and the heat dissipation efficiency and noise suppression goals are difficult to coordinately optimize.
[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: A high-performance computer system temperature control method, comprising the following steps: Initialize system parameters and divide the system into multiple independent thermal zones; Based on the divided thermal zones, real-time collection of temperature, power consumption and process data of each area is carried out; Build a dynamic heat capacity model based on the collected data, and update the heat capacity parameters and heat dissipation coefficient based on the real-time relationship between temperature change rate and power consumption; Dynamically adjust the multi-objective optimization weights based on the updated heat capacity parameters and heat dissipation coefficients, establish a distributed optimization model that includes thermal coupling effects, and solve the heat dissipation strategy; The heat dissipation strategy is executed and the strategy deviation is corrected through closed-loop control, while a data update cycle is triggered according to temperature fluctuations or process changes.
[0007] Preferably, the initialization system parameters include: Calculate the initial heat capacity value based on the density, specific heat capacity and core volume of the processor silicon wafer material; Determine the initial heat dissipation coefficient based on the radiator fin surface area and convection coefficient; When the system is divided into a plurality of independent hot zones, at least the central processing unit core, the memory channel and the network communication chip are allocated to different hot zones.
[0008] Preferably, the process data includes the process name and resource occupancy rate, and the burst process is marked in the following manner: If the real-time resource usage of a process exceeds a preset threshold, it is marked as a bursty process.
[0009] Preferably, the step of constructing a dynamic heat capacity model based on the collected data and updating the heat capacity parameters and the heat dissipation coefficient based on the relationship between the real-time temperature change rate and the power consumption includes: The thermodynamic differential equations for each thermal zone are established. The equations express the relationship between the heat capacity parameter and the temperature change rate, real-time power consumption, heat dissipation power, and ambient temperature. Specifically, in, Hot Zone The dynamic heat capacity, is the temperature change rate, Hot Zone Real-time temperature, Hot Zone Real-time power consumption, Hot Zone The heat dissipation coefficient, is the ambient temperature, Hot Zone Heat dissipation power; Based on real-time temperature change rate and power consumption , using the recursive least squares method to and Perform online identification, specifically: Define the parameter vector and observation vector , update the parameters iteratively by the following formula: in, is the Kalman gain matrix, Indicates the current moment.
[0010] Preferably, the step of dynamically adjusting the multi-objective optimization weights according to the updated heat capacity parameters and heat dissipation coefficients includes: According to the heat capacity parameters Rate of change Adjusting temperature stability weight , where the greater the rate of change, The higher the weight ratio; According to the hot zone temperature and preset safety thresholds The difference between the two values adjusts the noise suppression weight ,when near hour, The weight ratio is reduced; in, The adjustment satisfies the relationship , The adjustment satisfies the relationship , is the regulating factor, is the attenuation coefficient, and is the initial weight.
[0011] Preferably, the step of establishing a distributed optimization model including thermal coupling effects and solving a heat dissipation strategy includes: Heat diffusion constraints are introduced into the objective function, and the constraints are adjacent hot zones. The temperature gradient of the current hot zone The influence of is expressed as: in, Indicates hot zone Real-time temperature, Represents Adjacent hot zones Real-time temperature, Indicates hot zone and The thermal diffusion coefficient between for The set of adjacent hot zones; The improved alternating direction multiplier method is used to solve the optimization model. The heat dissipation strategy satisfies the thermal coupling constraints by alternately updating local variables and global consistency variables.
[0012] Preferably, the step of correcting the strategy deviation through closed-loop control includes: Generate an initial heat dissipation strategy based on current heat capacity parameters; The optimization objective is reversely corrected based on the parameter estimation error and the solution is re-solved.
[0013] Preferably, the conditions for triggering the data update cycle include at least one of the following: A burst of newly marked processes is detected; The temperature fluctuation in any hot zone exceeds the preset safety threshold.
[0014] The present invention also provides a temperature control device for implementing the method, comprising: Parameter initialization module, used to divide the hot zone and set the initial parameters; Data acquisition module, used to obtain temperature, power consumption and process data; Model building module for dynamically updating heat capacity parameters and heat dissipation coefficients; Optimization solution module, used to establish and solve distributed multi-objective optimization models; Dual-loop control module for generating and correcting cooling strategies; The instruction execution module is used to adjust the fan speed and cooling medium flow.
[0015] The present invention also provides a high-performance computer temperature control system, comprising: The control device; The radiator structure includes independent heat dissipation channels corresponding to the hot zones; A cooling medium supply unit, used to inject low-temperature medium into the cooling radiator; Distributed sensor array, used to collect temperature and power consumption data of each hot zone.
[0016] The present invention provides a temperature control method for a high-performance computer system. It has the following beneficial effects: 1. This invention, through the synergistic mechanism of a dynamic heat capacity model and online parameter identification, can track sudden changes in processor power consumption and heat dissipation characteristics in real time, overcoming the prediction lag caused by thermal inertia in traditional static models. Especially when sudden processes are frequently started or the load fluctuates drastically, the dynamic update of heat capacity parameters and heat dissipation coefficients ensures the accuracy of the thermodynamic model and provides a reliable basis for optimization decisions.
[0017] 2. This invention utilizes adaptive weighting rules based on the thermal zone temperature deviation and the rate of change of heat capacity, combined with a distributed optimization model that utilizes thermal diffusion coupling constraints, to effectively balance the conflicts between temperature stability, noise suppression, and thermal balance. By improving the ADMM algorithm to coordinate local and global optimization objectives, this approach avoids heat accumulation in adjacent areas caused by excessive local heat dissipation, thereby improving system-level heat dissipation efficiency.
[0018] 3. This invention utilizes a dual-loop control mechanism and an event-triggered data update loop to correct model deviations in real time during policy execution and rapidly switch control modes in response to unexpected processes or temperature violations. Through residual error detection and parameter rollback mechanisms, basic heat dissipation functions can be maintained in the event of sensor anomalies or model mismatches, preventing the risk of thermal runaway and ensuring continuous and stable system operation. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 Schematic diagram of the method flow of the present invention; Figure 2 Schematic diagram of the device structure of the present invention; Figure 3 Schematic diagram of the system structure of the present invention.
[0020] Among them, 10. Temperature control device; 11. Parameter initialization module; 12. Data acquisition module; 13. Model construction module; 14. Optimization solution module; 15. Dual-loop control module; 16. Instruction execution module; 20. Radiator structure; 30. Cooling medium supply unit; 40. Distributed sensor array. DETAILED DESCRIPTION
[0021] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the present specification. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0022] Please see the attached Figure 1 The present invention provides a temperature control method for a high-performance computer system, which realizes refined temperature control under complex load scenarios through a collaborative mechanism of dynamic heat capacity modeling, multi-objective optimization and closed-loop correction.
[0023] like Figure 1 As shown, the temperature control method of the high performance computer system may include the following steps: S1, initialize system parameters and divide the system into multiple independent hot zones; S2. Real-time collection of temperature, power consumption and process data of each area based on the divided thermal zones; S3. Build a dynamic heat capacity model based on the collected data, and update the heat capacity parameters and heat dissipation coefficient based on the relationship between the real-time temperature change rate and power consumption; S4. Dynamically adjust the multi-objective optimization weights based on the updated heat capacity parameters and heat dissipation coefficients, establish a distributed optimization model including thermal coupling effects, and solve the heat dissipation strategy; S5. Execute the cooling strategy and correct the strategy deviation through closed-loop control, while triggering the data update cycle according to temperature fluctuations or process changes.
[0024] The following is a detailed description of each step in the method of the present invention, which comprehensively explains the specific implementation principles, technical details and processes of each step.
[0025] In step S1, in this embodiment, the system parameters are initialized and the system is divided into multiple independent thermal zones. Physical property modeling and functional module decoupling are used to provide a basic framework for subsequent dynamic temperature control. The specific implementation is as follows: In this embodiment, the heat capacity parameters are pre-calibrated based on the inherent thermophysical properties of the processor silicon material. Specifically, according to the density of the silicon material Specific heat capacity and calculate the volume of the core , use the following formula to calculate the initial dynamic heat capacity: Among them, the volume The thickness of the silicon wafer is determined by the package size of the processor core and the thickness of the silicon wafer. Preferably, the thickness of the silicon wafer is obtained according to chip manufacturing process parameters, such as the thickness after wafer thinning using a photolithography process.
[0026] The initialization of the heat dissipation coefficient is set by combining the radiator structure parameters and the environmental convection conditions. and air convection coefficient , use the following formula to determine the initial heat dissipation coefficient: Preferably, the effective surface area Through 3D modeling and fluid dynamics simulation extraction of the radiator, we ensure that the main heat conduction paths are covered.
[0027] In this embodiment, a high-performance computer system is divided into multiple independent thermal zones based on the system hardware layout and thermal field distribution characteristics. Each thermal zone corresponds to at least one high-power computing unit or communication module. Preferably, the physical cores of the central processing unit (CPU) are grouped into independent thermal zones according to the non-uniform memory access (NUMA) architecture, and the memory controller channels and network communication chips are separately zoned. The basis for dividing the thermal zone boundaries includes, but is not limited to, the following conditions: The physical position spacing is greater than the preset threshold to avoid thermal diffusion coupling effects; The power consumption characteristics of functional modules vary significantly, requiring independent control strategies; The radiator structure supports independent control of each partition, such as a multi-fan array or split-path liquid cooling pipes.
[0028] Preferably, the hot zone division is completed through firmware configuration during the system startup phase, and the division results are stored in non-volatile memory for subsequent process calls. The divided hot zone numbers are mapped one-to-one with the sensor network to ensure physical consistency of data collection and control.
[0029] Initialized heat capacity parameters and heat dissipation coefficient Assign parameters to each independent hot zone according to their hot zone number to form an initial parameter matrix. Preferably, the parameter matrix is loaded into a shared memory area via a dynamic link library for access by the real-time control process. For heterogeneous computing units (such as a hybrid CPU and GPU architecture), a differentiated parameter initialization strategy is adopted. For example, the initial thermal capacity of the GPU hot zone is scaled proportionally to the size of the stream processor.
[0030] This step lays the data and structural foundation for subsequent dynamic modeling, optimization solution and closed-loop control through physical parameter calibration and modular hot zone division.
[0031] Regarding step S2, in this embodiment, the step of collecting temperature, power consumption, and process data of each area in real time based on the divided hot zones provides high-precision input for dynamic modeling and optimization control through multi-source heterogeneous data fusion and intelligent labeling mechanism. The specific implementation method is as follows: In this embodiment, temperature data acquisition is to deploy a distributed temperature sensor network in each independent thermal zone, preferably, using an embedded thermistor array or infrared thermal imaging unit to obtain the real-time temperature value of the thermal zone at a preset sampling frequency. ( is the hot zone number).
[0032] The temperature sensing network is strictly aligned with the thermal zone division results to ensure data spatial consistency. Preferably, the sampling frequency is dynamically adjusted according to the thermal inertia parameters of the hot zone, and a lower sampling rate is used in areas with larger heat capacity values to reduce system overhead.
[0033] In this embodiment, the power consumption data monitoring is to capture the power consumption data of each hot zone in real time by the processor's built-in performance counter unit. Specifically, for CPU core hotspots, dynamic power consumption is measured using a current sensor based on a voltage-frequency regulator. For memory and communication chip hotspots, the energy consumption register values of the power management module are preferably read through the bus interface and converted into a hotspot-level power consumption distribution. Power consumption data is synchronized with temperature acquisition timing to ensure timestamp alignment.
[0034] In this embodiment, process data analysis and marking is performed by intercepting operating system-level process scheduling information and extracting the mapping relationship between each process's resource usage and hotspots. Specifically, the kernel module is used to hook process migration events and record the switching trajectory of processes between hotspots. For processes that continuously run in a specific hotspot, the following rules are preferably used to determine whether they are burst processes: Dynamic threshold calculation: Set dynamic thresholds based on the statistical characteristics of historical resource utilization sequences ,in is the mean occupancy rate within the sliding time window, is the standard deviation, is the adjustment coefficient; Real-time marking logic: If the resource usage of a process in the current sampling period exceeds , it is marked as a burst process and its process identifier and associated hot zone number are recorded.
[0035] Preferably, the length of the sliding time window is adaptively adjusted according to the load fluctuation characteristics of the hot zone, and a shorter window is used in the high-fluctuation hot zone to enhance sensitivity. The marked burst process list is passed to the optimization control module through shared memory to trigger subsequent policy updates.
[0036] In this embodiment, the collected raw data is also subjected to noise filtering and outlier correction. Preferably, the temperature data is smoothed using a Kalman filter algorithm, whose state equation is constructed based on a thermodynamic differential model, and the observation equation is associated with the sensor accuracy parameter. For power consumption data, transient glitches are detected through differential verification of adjacent sampling points, and anomalous data points are replaced using linear interpolation. The preprocessed data is stored in a circular buffer for use by the model construction and parameter update modules.
[0037] This step uses high-precision sensor networks, intelligent process labeling, and robust data preprocessing to provide low-noise, high-quality input data for subsequent dynamic thermal capacity modeling and multi-objective optimization, ensuring the response speed and decision-making reliability of the temperature control system under complex load scenarios.
[0038] Regarding step S3, in this embodiment, the steps of constructing a dynamic heat capacity model based on the collected data and updating the heat capacity parameters and heat dissipation coefficient based on the relationship between the real-time temperature change rate and power consumption are achieved through a collaborative mechanism of thermodynamic modeling and online parameter identification. The specific implementation method is as follows: In this embodiment, the thermodynamic differential equations of each hot zone are first established based on the principle of energy conservation. , the relationship between the temperature change rate and energy input is expressed as: in: Indicates hot zone The dynamic heat capacity (unit: J / ℃) reflects the ability of the hot zone to store heat; is the real-time temperature change rate (unit: °C / s), which is calculated by the difference of temperature sensor data; The real-time power consumption of the hot zone processor (unit: W), directly read by the hardware performance counter; is the heat dissipation coefficient of the hot zone (unit: W / ℃), which represents the passive heat dissipation efficiency; is the ambient temperature (unit: °C), measured by an independent environmental sensor; Active cooling power (unit: W) is calculated by reverse calculation based on the cooling device control signal, such as the mapping relationship between fan speed and power curve.
[0039] In this embodiment, based on the model construction, the recursive least squares method is used to and Dynamic update. Specifically, rewrite the model into a linear regression form: Define the parameter vector With the observation vector , update the parameters iteratively through the recursive formula: in, is the Kalman gain matrix, Indicates the current moment.
[0040] Kalman gain matrix Calculated by the following formula: in, is the covariance matrix, initialized to a diagonal matrix ( is a constant greater than zero, preferably 10 3 ~10 5 ), and updated in each iteration to: Forgetting Factor (satisfy ) is used to adjust the weight of historical data: When the temperature change rate When the threshold is exceeded, the to below 0.9 to accelerate parameter tracking; when the temperature stabilizes, restore to above 0.98 to improve estimation stability.
[0041] In this embodiment, in order to improve the parameter identification accuracy, the original temperature and power consumption data are also pre-processed. , using sliding window weighted average filtering, the window length is related to the heat capacity of the hot zone Positive correlation, specifically Sampling points ( In J / ℃). For power consumption data ,Transient noise is detected by differential verification of adjacent sampling points.,If the power consumption jump of two consecutive sampling points exceeds 3 times the,standard deviation of the historical mean, it is determined to be an outlier and,replaced with the linear interpolation result.
[0042] In this embodiment, during the parameter update process, the residual is calculated in real time. , and calculates its root mean square error (RMSE). If the RMSE exceeds a preset threshold (for example, 10% of the maximum power consumption of a hot zone) for five consecutive cycles, it is determined to be a model mismatch, triggering the following recovery process: Reset the covariance matrix is the initial diagonal matrix, restoring the fast convergence capability of the parameter identifier; Pause active cooling control for three sampling cycles to collect non-interference data through the natural cooling process to recalibrate the model; After the control is restored, the initial parameters are calculated in batches using the data from the first three cycles, and then the system switches to the recursive mode.
[0043] This step achieves dynamic tracking and adaptive correction of thermodynamic parameters through the deep integration of physical models and data-driven methods, providing a high-confidence model foundation for subsequent multi-objective optimization and ensuring the accuracy and stability of the heat dissipation strategy under complex working conditions.
[0044] Regarding step S4, in this embodiment, the step of dynamically adjusting the multi-objective optimization weights and establishing a distributed optimization model including thermal coupling effects is achieved through weight adaptive mapping and collaborative optimization mechanism. The specific implementation method is as follows: In this embodiment, based on the heat capacity parameter updated in step S3 and heat dissipation coefficient , establish the temperature stability weight With noise suppression weight Real-time adjustment rules.
[0045] For the temperature stability weight, the rate of change of the heat capacity parameter To perform linear regulation: in, is the initial weight reference value, is a regulating factor used to control the sensitivity of the change rate to the weight. The sliding window difference calculation of the heat capacity parameter sequence is performed, and the window length matches the thermal inertia parameters of the hot zone to avoid high-frequency noise interference.
[0046] For noise suppression weight, according to the real-time temperature of the hot zone and preset safety thresholds The degree of deviation is adjusted exponentially by decay: in, is the initial weight reference value, is the attenuation coefficient, which is used to control the influence of temperature deviation on the weight. Close to or exceed hour, The significant reduction makes the optimization model prioritize temperature safety rather than noise suppression goals.
[0047] In this embodiment, the dual objectives of minimizing temperature fluctuations and heat dissipation equipment noise are used to construct an optimization model that includes the thermal zone coupling effect. The objective function is defined as: in, Indicates the control variable of the cooling device (such as fan speed or liquid cooling pump power) corresponding to the noise function, preferably, a quadratic function Approximately describe the relationship between noise and control quantity.
[0048] Furthermore, the thermal coupling effect constraint between the thermal zones is introduced into the objective function. According to the heat conduction principle, the adjacent thermal zones The temperature gradient of the current hot zone The effect of is described by the partial differential equation: in, for The set of adjacent hot zones, Indicates hot zone Real-time temperature, Represents Adjacent hot zones Real-time temperature, Indicates hot zone and The thermal diffusivity between the hot zones is calculated from the distance between them, the thermal conductivity of the materials, and the contact area. Constraints ensure that the optimization strategy reduces local temperatures without causing heat accumulation in adjacent areas.
[0049] In this embodiment, the improved alternating direction method of multipliers (ADMM) is used to solve the distributed optimization model. The global optimization problem is decomposed into two sub-problems: Local variable update: Each hot zone solves the control variable independently , the goal is to minimize the local objective function , while satisfying the heat capacity model constraints; Global consistency update: The temperature prediction values of each hot zone are coordinated through Lagrange multipliers to satisfy the heat diffusion equation coupling relationship.
[0050] Preferably, the improved ADMM algorithm introduces a dynamic relaxation factor into the standard ADMM framework to accelerate the convergence of thermal coupling constraints. The relaxation factor is calculated based on the temperature difference between adjacent hot zones. Adaptive adjustment: When the difference is larger, the relaxation factor is smaller to avoid oscillation.
[0051] The iterative termination conditions of the solution process include: Set the double residual threshold as the iteration termination condition. Define the original residual and the dual residual ,in is the penalty parameter, is a global consistency variable. and When the iteration is terminated, the final cooling strategy is output .
[0052] Through this process, dynamically adjusted weight coefficients and thermal coupling constraints work together on the optimization model to ensure that the cooling strategy strikes a balance between temperature stability, noise suppression, and heat diffusion equilibrium. At the same time, the distributed solution mechanism supports the real-time control requirements of large-scale systems.
[0053] Regarding step S5, in this embodiment, executing the heat dissipation strategy and correcting the strategy deviation through closed-loop control, while triggering the data update cycle, is achieved through dual-loop control and dynamic triggering mechanism. The specific implementation is as follows: In this embodiment, the optimal heat dissipation strategy obtained based on step S4 is , the control variable (such as fan speed, liquid cooling pump power) is converted into a drive signal and sent to the cooling device.
[0054] Preferably, a pulse width modulation (PWM) signal is used to control the fan speed, and its duty cycle and target speed The mapping relationship is realized through a preset linear or nonlinear curve. For liquid cooling system, the control quantity It is converted into voltage or current instruction of the pump and output to the actuator through the digital-to-analog conversion module.
[0055] During the strategy execution process, real-time monitoring of hot zone temperature and the model prediction value Deviation . If there is a deviation When the preset threshold is exceeded (e.g. 5% of the hot zone safety threshold), the closed-loop correction mechanism is triggered: Parameter error back propagation: according to the deviation Reverse correction of the heat capacity parameter in step S3 and heat dissipation coefficient Specifically, the gradient descent method is used to update the parameters:
[0056] in, is the learning rate, , partial derivatives Calculated using the chain rule combined with a thermodynamic model.
[0057] Optimization target redefinition: Reconstruct the multi-objective optimization model in step S4 based on the revised parameters, and call the ADMM solver to calculate the updated cooling strategy .
[0058] During the closed-loop control process, the following two types of events are dynamically detected to trigger a data update cycle to restart the complete process starting from step S2: Burst process flag event: When a new process is marked as a burst process, the current control cycle is immediately interrupted and the process jumps to step S2 to recollect temperature, power consumption, and process data. Preferably, the operating system kernel's event notification mechanism captures the process flag status changes in real time, ensuring that the response delay is less than a preset time window (e.g., 50 milliseconds).
[0059] Temperature safety limit event: If any hot zone temperature Exceeding the preset safety threshold If this condition persists for at least two sampling cycles, it is considered an emergency thermal runaway risk and a data update cycle is triggered. During the cycle restart phase, the cooling system switches to maximum cooling power mode until the temperature returns to a safe range.
[0060] When the parameter error After multiple corrections, it still fails to converge (for example, after 3 consecutive iterations If the drop rate is less than 1%, it is determined to be a model mismatch or hardware failure. Perform the following recovery operations: 1. Cooling strategy rollback: control the amount Restore the historical strategy to the most recent stable state to avoid control oscillations; 2. Data link verification: Check the communication integrity between the sensor data acquisition channel and the actuator control channel, and reset abnormal hardware interfaces; 3. Thermodynamic model reset: clear the parameter estimation results in step S3 and reinitialize the covariance matrix And trigger batch parameter identification.
[0061] Through the above mechanism, the closed-loop control system can dynamically compensate for model prediction errors and external disturbances, while ensuring system safety under extreme working conditions through an event triggering mechanism, thereby achieving robustness and adaptability of the cooling strategy.
[0062] In general, the present invention divides the system into multiple independent thermal zones and initializes thermal capacity and heat dissipation parameters; collects temperature, power consumption and process data in real time, identifies sudden processes and constructs a dynamic thermal capacity model; uses recursive least squares method to update thermodynamic parameters online, and dynamically adjusts optimization weights based on the degree of deviation between the hot zone temperature and the safety threshold; establishes a distributed optimization model that includes thermal diffusion coupling constraints, and solves the heat dissipation strategy through an improved alternating direction multiplier method; when executing the strategy, corrects model deviations through dual-loop control, and triggers a data update cycle based on sudden process markers or temperature limit-exceeding events, ensuring temperature stability and system safety under complex load and thermal coupling scenarios.
[0063] Please see the attached Figure 2 The present invention also provides a temperature control device 10 for implementing the above method. The device comprises a parameter initialization module 11, a data acquisition module 12, a model construction module 13, an optimization solution module 14, a dual-loop control module 15, and an instruction execution module 16. Each module operates in coordination to implement closed-loop temperature control logic: Parameter initialization module 11: divide the hot zone boundary based on the system physical layout and historical thermal imaging data. Preferably, use three-dimensional heat conduction simulation to determine the initial heat capacity of the hot zone and heat dissipation coefficient and storing the parameter configuration table via non-volatile memory; Data acquisition module 12: Integrates multi-source heterogeneous data interfaces, including an embedded thermistor array connected to the I2C bus, a power consumption metering unit of the PCIe interface, and an operating system kernel-level process monitoring agent to achieve temperature , power consumption And millisecond-level synchronous collection of process data, preferably, a data transmission mechanism combining hardware interrupts and DMA is used to reduce CPU load; Model building module 13: built-in recursive least squares algorithm core, real-time input from data acquisition module, dynamic update of heat capacity parameters and heat dissipation coefficient , and through the covariance matrix Maintain confidence in parameter estimates; Optimization solution module 14: uses an improved ADMM algorithm to solve the multi-objective optimization problem, wherein local optimizers are deployed in computing nodes corresponding to each hot zone, and a global coordinator synchronizes Lagrange multipliers through a high-speed interconnection network. Preferably, sparse matrix compression technology is introduced for heat diffusion coupling constraints to reduce communication overhead; Dual-loop control module 15: integrated PID controller and event trigger logic, the inner loop is based on temperature deviation Adjust the cooling strategy in real time. The outer loop triggers model reconstruction through residual analysis. Preferably, a hardware-accelerated gradient calculation unit is used to implement parameter back propagation. Instruction execution module 16: supports multi-modal output of PWM, analog voltage and RS485 protocol, and is adapted to heterogeneous heat dissipation devices such as fans, liquid cooling pumps and semiconductor refrigeration chips. Preferably, a dead zone compensation circuit is configured to eliminate the influence of actuator response lag.
[0064] The device of this embodiment can be used to execute the above method embodiment, and its principles and technical effects are similar, so they will not be repeated here.
[0065] Please see the attached Figure 3 The present invention also provides a high-performance computer temperature control system. The system consists of a control device 10, a cooling structure 20, a cooling medium supply unit 30, and a distributed sensor array 40. Specifically, Radiator structure 20: A microchannel radiator design is used, corresponding to each hot zone one by one. Each independent heat dissipation channel is equipped with a micro solenoid valve and a flow sensor. Preferably, the radiator channel topology is optimized based on the thermal coupling strength of the hot zone, and cross-connected channels are used in high-coupling areas to enhance heat dissipation balance. Cooling medium supply unit 30: includes a multi-stage centrifugal pump, a liquid storage tank and a temperature control device, dynamically adjusts the cooling medium flow rate and inlet temperature according to the demand of the hot zone. Preferably, a branch closed-loop control strategy is adopted, and a pressure-flow decoupling algorithm is used to eliminate mutual interference when multiple channels are connected in parallel; Distributed sensor array 40: deploys high-precision thin-film platinum resistance temperature sensors and Hall effect power consumption metering chips, and transmits data via shielded twisted pair cables or optical fibers. Preferably, the sensor nodes have a built-in self-test function and periodically output diagnostic signals to detect disconnection or drift faults.
[0066] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is determined by the appended claims and their equivalents.
Claims
1. A temperature control method for a high-performance computer system, characterized in that: The following steps are involved: Initialize system parameters and divide the system into multiple independent thermal zones; Based on the divided thermal zones, real-time collection of temperature, power consumption and process data of each area is carried out; Build a dynamic heat capacity model based on the collected data, and update the heat capacity parameters and heat dissipation coefficient based on the real-time relationship between temperature change rate and power consumption; Dynamically adjust the multi-objective optimization weights based on the updated heat capacity parameters and heat dissipation coefficients, establish a distributed optimization model that includes thermal coupling effects, and solve the heat dissipation strategy; The heat dissipation strategy is executed and the strategy deviation is corrected through closed-loop control, while a data update cycle is triggered according to temperature fluctuations or process changes.
2. The temperature control method for a high performance computer system according to claim 1, wherein: The initialization system parameters include: Calculate the initial heat capacity value based on the density, specific heat capacity and core volume of the processor silicon wafer material; Determine the initial heat dissipation coefficient based on the radiator fin surface area and convection coefficient; When the system is divided into a plurality of independent hot zones, at least the central processing unit core, the memory channel and the network communication chip are allocated to different hot zones.
3. The temperature control method for a high performance computer system according to claim 1, wherein: The process data includes the process name and resource usage, and the burst process is marked in the following ways: If the real-time resource usage of a process exceeds a preset threshold, it is marked as a bursty process.
4. The temperature control method of a high performance computer system according to claim 1, wherein: The steps of constructing a dynamic heat capacity model based on the collected data and updating the heat capacity parameters and the heat dissipation coefficient based on the relationship between the real-time temperature change rate and the power consumption include: The thermodynamic differential equations for each thermal zone are established. The equations express the relationship between the heat capacity parameter and the temperature change rate, real-time power consumption, heat dissipation power, and ambient temperature. Specifically, in, Hot Zone The dynamic heat capacity, is the temperature change rate, Hot Zone Real-time temperature, Hot Zone Real-time power consumption, Hot Zone The heat dissipation coefficient, is the ambient temperature, Hot Zone Heat dissipation power; Based on real-time temperature change rate and power consumption , using the recursive least squares method to and Perform online identification, specifically: Define the parameter vector and observation vector , update the parameters iteratively by the following formula: in, is the Kalman gain matrix, Indicates the current moment.
5. The temperature control method for a high performance computer system according to claim 1, wherein: The step of dynamically adjusting the multi-objective optimization weights according to the updated heat capacity parameters and heat dissipation coefficients includes: According to the heat capacity parameters Rate of change Adjusting temperature stability weight , where the greater the rate of change, The higher the weight ratio; According to the hot zone temperature and preset safety thresholds The difference between the two values adjusts the noise suppression weight ,when near hour, The weight ratio is reduced; in, The adjustment satisfies the relationship , The adjustment satisfies the relationship , is the regulating factor, is the attenuation coefficient, and is the initial weight.
6. The temperature control method for a high performance computer system according to claim 5, wherein: The steps of establishing a distributed optimization model including thermal coupling effects and solving a heat dissipation strategy include: Heat diffusion constraints are introduced into the objective function, and the constraints are adjacent hot zones. The temperature gradient of the current hot zone The influence of is expressed as: in, Indicates hot zone The real-time temperature of Adjacent hot zones Real-time temperature, Indicates hot zone and The thermal diffusion coefficient between for The set of adjacent hot zones; The improved alternating direction multiplier method is used to solve the optimization model. The heat dissipation strategy satisfies the thermal coupling constraints by alternately updating local variables and global consistency variables.
7. The temperature control method for a high performance computer system according to claim 1, wherein: The step of correcting the strategy deviation through closed-loop control includes: Generate an initial heat dissipation strategy based on current heat capacity parameters; The optimization objective is reversely corrected based on the parameter estimation error and the solution is re-solved.
8. The temperature control method for a high performance computer system according to claim 7, wherein: The conditions for triggering the data update cycle include at least one of the following: A burst of newly marked processes is detected; The temperature fluctuation in any hot zone exceeds the preset safety threshold.
9. A temperature control device for implementing the method according to any one of claims 1 to 8, characterized in that: include: Parameter initialization module, used to divide the hot zone and set the initial parameters; Data acquisition module, used to obtain temperature, power consumption and process data; Model building module for dynamically updating heat capacity parameters and heat dissipation coefficients; Optimization solution module, used to establish and solve distributed multi-objective optimization models; Dual-loop control module for generating and correcting cooling strategies; The instruction execution module is used to adjust the fan speed and cooling medium flow.
10. A high performance computer temperature control system, characterized in that: include: The control device according to claim 9; The radiator structure includes independent heat dissipation channels corresponding to the hot zones; A cooling medium supply unit, used to inject low-temperature medium into the cooling radiator; Distributed sensor array, used to collect temperature and power consumption data of each hot zone.
Citation Information
Cited By
Cooperative control method and device for light transmittance and heat dissipation of front light-emitting LED transparent screen
CN121053902A
Corrugated roller temperature control method and system
CN121326046A
Multi-mode power coordination control method and system for electric tractor
CN122219285A