A variable temperature cooling energy consumption control method for a data center cooling system

By constructing a demand database and a dual-channel predictive architecture, a variable-temperature cooling energy consumption control method was developed, which solved the thermal inertia lag problem of data center cooling systems, achieved real-time and accurate matching of cooling systems, reduced energy consumption, and improved system security.

CN122269650APending Publication Date: 2026-06-23BEIJING ZHONGDAO GREEN ENERGY TECH DEV CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING ZHONGDAO GREEN ENERGY TECH DEV CO LTD
Filing Date
2026-03-30
Publication Date
2026-06-23

AI Technical Summary

Technical Problem

Existing data center cooling systems suffer from thermal inertia, causing the supply of cooling capacity to lag behind the generation of heat, resulting in localized hot spots, chip temperature overshoot, and wasted energy.

Method used

A variable-temperature cooling energy consumption control method is adopted. By constructing a demand database, a dual-channel prediction architecture, and an adaptive model prediction control framework, the system can achieve real-time and accurate matching of the cooling system. Combined with a hierarchical safety interlock mechanism, the cooling parameters are dynamically adjusted to solve the thermal inertia problem.

Benefits of technology

It enables proactive prediction of IT load, reduces cooling system response latency, lowers cooling energy consumption, and improves the system's energy efficiency and security, especially effectively preventing equipment overheating in scenarios with sudden load increases such as AI training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122269650A_ABST
    Figure CN122269650A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of data center cooling control, and discloses a variable-temperature cooling energy consumption control method of a data center cooling system. A double-channel prediction framework is constructed to predict steady-state loads and burst loads respectively, and a hardware performance counter and other fine-grained indexes are introduced to capture the precursor characteristics of load sudden increase, thereby realizing pre-position prediction of IT loads. A thermal dynamics model is embedded into a model prediction control framework, online identification of cooling system thermal capacity and thermal resistance parameters is carried out, explicit modeling of system thermal inertia is realized, an adaptive mechanism based on prediction confidence is introduced, the prediction time domain and temperature change rate are dynamically adjusted, the system robustness is enhanced at low confidence, a three-level protection system from early warning to emergency protection is constructed through a hierarchical safety interlocking mechanism, the system can be rapidly intervened and the system safety can be ensured when abnormal working conditions occur, and bottom security guarantee is provided for variable-temperature regulation and control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data center cooling control technology, and in particular to a method for controlling the energy consumption of variable temperature cooling in a data center cooling system. Background Technology

[0002] With the rapid development of artificial intelligence, cloud computing, and big data technologies, the scale and computing density of data centers are constantly increasing. Among them, the power fluctuation of AI training clusters is large and the rate of change is fast, with the power density of a single rack climbing from the traditional 5-8kW to over 30-60kW. As an important component of data centers, the cooling system typically accounts for 30%-40% of the total energy consumption of a data center, indicating huge potential for energy saving.

[0003] Currently, mainstream data center cooling systems include water-cooled chiller systems, immersion liquid cooling systems, and indirect evaporative cooling systems. These systems generally employ constant temperature control or simple PID feedback control, meaning a fixed chilled water / coolant supply temperature is set, and adjustment is initiated only when sensors detect a temperature rise. However, due to the inherent thermal inertia of cooling systems—water systems have minute-level heat transfer delays, liquid cooling systems have high heat capacity, and air-cooled systems have air convection inertia—the supply of cooling always lags behind the generation of heat. This lag effect is particularly pronounced in scenarios with sudden load increases, such as AI training, often leading to localized hotspots, chip temperature overshoot, and even triggering equipment frequency reduction or shutdown protection. To avoid such risks, maintenance personnel are forced to adopt a conservative strategy of erring on the side of cooler rather than hotter, setting the supply water temperature at excessively low levels for extended periods, resulting in wasted cooling capacity and decreased energy efficiency.

[0004] Therefore, how to achieve real-time and accurate matching between cooling supply and dynamic heat load, and fundamentally solve the lag problem caused by the thermal inertia of the cooling system, has become a technical problem that urgently needs to be solved in this field. Summary of the Invention

[0005] The technical problem to be solved by this invention is that the existing technology has the disadvantage of lag caused by the thermal inertia of the cooling system. To address this, we propose a variable temperature cooling energy consumption control method for data center cooling systems.

[0006] To achieve the above objectives, this application adopts the following technical solution: a method for controlling the variable temperature cooling energy consumption of a data center cooling system, comprising: analyzing the cooling demand distribution of the data center and constructing a demand database containing IT equipment temperature safety thresholds, cooling system equipment safety thresholds, and energy efficiency targets; acquiring multi-dimensional real-time operating data of the data center and constructing a multi-source feature dataset that has undergone time-series alignment and validity verification; based on the multi-source feature dataset, performing steady-state load prediction and burst load prediction respectively through a dual-channel prediction architecture, and outputting a load prediction curve with prediction confidence; based on the load prediction curve and the safety constraints in the demand database, employing an adaptive model predictive control framework embedded with a thermodynamic model to solve for the optimal variable temperature control command sequence; executing cooling system parameter adjustments according to the optimal variable temperature control command sequence, and performing closed-loop correction based on real-time feedback to generate a closed-loop correction deviation; when the monitored parameters trigger a safety threshold, activating a graded safety interlocking mechanism and outputting the interlocking protection status.

[0007] Furthermore, the analysis of the cooling demand distribution of the data center includes: spatially partitioning the data processing equipment according to the temperature distribution gradient of the data center space to obtain M temperature space intervals; performing load tracking on the data processing equipment within the M temperature space intervals to obtain the load of each temperature space interval; establishing a relationship between load and spatial heat accumulation calculation based on the temperature characteristics of the M temperature space intervals and the load, and obtaining the heat accumulation value of each temperature space interval through heat accumulation calculation; determining the cooling demand of each temperature space interval based on the heat accumulation value, and analyzing the correspondence between the M temperature space intervals, load, and cooling demand to construct the demand database.

[0008] Furthermore, the construction of the demand database also includes: fitting the cooling penalty probability of data processing devices in M ​​temperature space intervals based on historical tracking samples, converting it into cooling penalty coefficients for each temperature space region; and setting demand weights for the demand database using the cooling penalty coefficients for each temperature space region.

[0009] Furthermore, the dual-channel prediction architecture is used to perform steady-state load prediction and burst load prediction respectively, including: constructing a steady-state load prediction channel based on historical IT load operation data, and using a long short-term memory network to output the load trend baseline and steady-state prediction confidence interval for the next 1-24 hours; constructing a burst load prediction channel based on server hardware performance counter indicators, and using a temporal convolutional network to output the probability of load mutation, estimated burst increment, and burst prediction confidence for the next 15-60 minutes; and using a Bayesian fusion method to dynamically weight the prediction confidence of the two channels to generate the final IT load prediction curve and load mutation probability heatmap, while outputting the prediction confidence at each prediction time.

[0010] Furthermore, the server hardware performance counter metrics include one or more of the following: CPU instruction reordering buffer utilization, cache miss rate, memory bandwidth utilization, PCIe link throughput, GPU utilization, video memory bandwidth, gradient synchronization frequency, and NVLink throughput.

[0011] Furthermore, the adaptive model predictive control framework employing an embedded thermodynamic model to solve for the optimal temperature regulation command sequence in a rolling manner includes:

[0012] A simplified lumped-parameter thermodynamic model of the cooling system is established, treating the core components of the chiller, plate heat exchanger, cooling tower, circulation pipeline, and terminal heat exchange equipment as independent lumped heat capacity nodes. The heat capacity parameters of each node and the thermal resistance parameters between adjacent nodes are estimated using a system identification method, forming a thermodynamic model parameter library. Based on the predicted confidence level, the prediction time domain length and temperature change rate constraints of the model predictive control are dynamically adjusted. When the prediction confidence level is high, the prediction time domain is extended and a faster temperature change rate is allowed; when the prediction confidence level is low, the prediction time domain is shortened and the temperature change rate is tightened. In each control cycle, based on the current state, load prediction curve, and thermodynamic model, and with the chip temperature not exceeding the safety constraint in the future time domain as a rigid condition, a mixed-integer nonlinear programming algorithm is used to solve for the globally optimal control sequence. The instruction sequence of the first control cycle is then issued and executed as the optimal temperature regulation instruction sequence.

[0013] Furthermore, the simplified lumped parameter thermodynamic model of the cooling system includes N lumped heat capacity nodes. The temperature recursion relationship of each node is determined by the node heat capacity, the thermal resistance between adjacent nodes, the node heat source power, and the discrete time step. The heat capacity parameters and thermal resistance parameters are estimated and dynamically updated online by recursive least squares method or extended Kalman filter.

[0014] Furthermore, the step of adjusting the cooling system parameters according to the optimal variable temperature control command sequence and performing closed-loop correction based on real-time feedback includes: adjusting slowly varying parameters such as the water supply temperature setpoint and the number of operating devices according to the model predictive control command cycle; performing PID closed-loop control on a minute-level cycle for rapidly varying parameters such as pump frequency, fan frequency, and valve opening; triggering PID fine-tuning when the operating deviation exceeds the set threshold, generating a closed-loop correction deviation amount and adding it to the execution command; and immediately triggering model predictive control re-optimization and updating the control parameters when the actual IT load suddenly increases by more than 20% of the predicted value or the chip temperature approaches the safety threshold.

[0015] Furthermore, the activation of the graded safety interlock mechanism includes: triggering a level one warning when the monitored parameter reaches 90% of the safety threshold, suspending the optimization and adjustment of the temperature parameters; triggering a level two over-limit protection when the monitored parameter exceeds the safety warning value, switching to the safety priority mode and forcibly reducing the cooling temperature; and triggering a level three emergency interlock when there are extreme safety risks such as a major leakage in the cooling system, the temperature of IT equipment soaring to the shutdown threshold, or the fire alarm signal being triggered, locking the temperature control system and switching to emergency manual mode.

[0016] Furthermore, a variable-temperature cooling energy consumption control system for a data center cooling system is proposed. The system includes: a demand database construction module for analyzing the cooling demand distribution of the data center and constructing a demand database containing IT equipment temperature safety thresholds, cooling system equipment safety thresholds, and energy efficiency targets; a multi-dimensional data acquisition module for acquiring multi-dimensional real-time operational data of the data center and constructing a multi-source feature dataset that has undergone time-series alignment and validity verification; a load and weather forecasting module for performing steady-state load forecasting and burst load forecasting respectively based on the multi-source feature dataset using a dual-channel forecasting architecture, outputting a load forecast curve with prediction confidence; and a multi-objective optimization decision module. The system comprises four modules: a load forecasting module and an adaptive model predictive control framework with an embedded thermodynamic model, used to solve for the optimal temperature regulation command sequence based on the load forecasting curve and safety constraints in the demand database; a closed-loop execution control module, used to adjust cooling system parameters according to the optimal temperature regulation command sequence and perform closed-loop correction based on real-time feedback, generating a closed-loop correction deviation; a safety interlock and emergency control module, used to activate a graded safety interlock mechanism and output the interlock protection status when the monitored parameters trigger a safety threshold; and an effect verification and iterative optimization module, used to retrain the prediction model based on historical operating data and update the thermodynamic model parameters, outputting iterative optimization parameters.

[0017] The technical effects and advantages of this invention are as follows: This invention constructs a dual-channel prediction architecture to predict steady-state loads and sudden loads separately. It also introduces fine-grained indicators such as hardware performance counters to capture precursory characteristics of load surges, achieving proactive prediction of IT loads. This provides the cooling system with differentiated time margins ranging from 15 minutes to 24 hours, fundamentally solving the lag problem of "heating first, cooling later" in traditional control. Furthermore, this invention embeds a thermodynamic model into the model predictive control framework. By identifying the cooling system's heat capacity and thermal resistance parameters online, it achieves explicit modeling of the system's thermal inertia. This allows the optimizer to proactively plan the liquid supply temperature trajectory and actively utilize the system's thermal inertia to store cooling capacity, avoiding response delays caused by thermal inertia. Finally, this invention introduces an adaptive mechanism based on prediction confidence, dynamically adjusting the prediction time domain and temperature change rate. At high confidence levels, it fully exploits energy-saving potential, while at low confidence levels, it enhances system robustness, achieving a dynamic balance between energy-saving effects and operational safety. This invention constructs a three-tiered protection system from early warning to emergency protection through a hierarchical safety interlock mechanism. This system can quickly intervene and ensure system safety when abnormal operating conditions occur, providing a safety net for temperature regulation and addressing security concerns regarding intelligent control algorithms in practical deployments. Significant results have been achieved in multiple data center applications: the cabinet intake temperature fluctuation of water-cooled chiller systems has been reduced from ±3℃ to within ±1℃, cooling energy consumption has decreased by 22%, and the utilization time of natural cold sources has increased by 62%; the chip temperature fluctuation of immersion liquid cooling systems has decreased by 40%, with the annual CLF remaining stable below 0.08, and PUE reaching advanced levels of 1.13 and 1.08 respectively, fully verifying the technical effectiveness and practical value of this invention. Attached Figure Description

[0018] The disclosure of this invention is illustrated with reference to the accompanying drawings. It should be understood that the drawings are for illustrative purposes only and are not intended to limit the scope of protection of this invention. In the drawings, the same reference numerals are used to refer to the same parts:

[0019] Figure 1 This is an overall framework diagram of the variable temperature cooling energy consumption control system of the present invention; Figure 2 This is a flowchart of the method of the present invention; Figure 3 This is the logic diagram of the adaptive MPC based on prediction confidence of the present invention. Detailed Implementation

[0020] It is readily understood that, based on the technical solution of this invention, those skilled in the art can propose various interchangeable structural methods and implementations without altering the essential spirit of the invention. Therefore, the following detailed embodiments and accompanying drawings are merely illustrative examples of the technical solution of this invention and should not be considered as the entirety of the invention or as limitations or restrictions on the technical solution of this invention.

[0021] Reference Figure 1 - Figure 3 As shown, the present invention provides a variable-temperature cooling energy consumption control method for data center cooling systems, aiming to solve the core technical problem in existing cooling control technologies where the time scale mismatch between the dynamic fluctuations of IT equipment load and the thermal inertia of the cooling system leads to a lag in cooling supply compared to heat generation, resulting in local hot spots or over-cooling and energy waste. The method specifically includes the following steps:

[0022] Step S1: Pre-configure the reference constraints and hierarchical system for variable temperature control.

[0023] Before the system is put into operation, the rigid safety constraint thresholds, energy efficiency targets, and control condition classification benchmarks must be pre-configured. All configuration parameters must be set by qualified operation and maintenance engineers on the local or cloud management platform and take effect after double confirmation to ensure compliance with relevant national and industry standards. Specifically, this includes: S11: Configuring rigid safety constraint thresholds, divided into two categories: IT equipment safety constraints and cooling system equipment safety constraints. IT equipment and data center environment safety constraints: Configuring the main data center ambient temperature threshold to be 18℃-30℃ and the relative humidity threshold to be 35%-75% during startup, with no condensation allowed throughout; for single-phase immersion liquid cooling systems, configuring the temperature difference threshold at the same horizontal cross-section in the liquid phase zone to be ≤5℃; for phase change immersion liquid cooling systems, configuring the gas phase zone pressure control threshold to be atmospheric pressure ±25kPa, configuring the junction temperature of server core components not to exceed the rated maximum allowable temperature of the components, and reserving a 5℃ safety redundancy; Cooling system equipment safety constraints: Configuring the total cooling capacity of the cooling system to reserve a cooling margin of not less than 20%, configuring the water quality of the water cooling system to ensure the cooling medium temperature is at least 2℃ higher than the dew point temperature of the data center environment to prevent condensation protection thresholds; configuring the rated operating parameter boundaries of equipment such as chillers, pumps, and fans, including the highest / lowest operating frequency, maximum allowable start / stop times / hour, minimum load rate, etc., to avoid equipment exceeding its rated operating capacity or frequent start / stop causing lifespan degradation.

[0024] S12: Configure Energy Efficiency Target Thresholds: Configure the target PUE value for data centers. Existing data centers must meet the corresponding local standard PUE limit requirements, while newly built data centers must meet the admission requirements. The optimization objective is to achieve the advanced value requirements: Level 1 energy efficiency target PUE ≤ 1.20, and Level 2 energy efficiency target PUE ≤ 1.30. This target value will serve as the basis for the soft constraint weights in the optimization objective function.

[0025] S13: Configure control condition classification benchmarks, including meteorological conditions and load conditions, to provide a classification basis for differentiated temperature control: Based on the dry bulb temperature outside the data center, it is divided into 5 standard control conditions: high temperature condition (outdoor dry bulb temperature ≥ 30℃), medium-high temperature condition (20℃ ≤ outdoor dry bulb temperature < 30℃), medium temperature condition (10℃ ≤ outdoor dry bulb temperature < 20℃), low temperature condition (0℃ ≤ outdoor dry bulb temperature < 10℃), and severe cold condition (outdoor dry bulb temperature < 0℃); Based on the IT equipment load utilization rate, it is divided into 4 load levels: low load range (IT load rate < 30%), medium-low load range (30% ≤ IT load rate < 50%), medium-high load range (50% ≤ IT load rate < 75%), and full load range (IT load rate ≥ 75%).

[0026] S14: Pre-built energy efficiency MAP library for core equipment of the cooling system. Through factory test data and on-site trial operation data, the COP curves of chiller units under different chilled water supply temperature, cooling water inlet temperature and load rate are collected; the approximation curves of cooling towers under different ambient temperature and humidity and fan frequency are collected; and the efficiency curves of circulating pumps under different flow rates and heads are collected. All data are stored in the form of two-dimensional or three-dimensional lookup tables to provide basic data for subsequent energy efficiency optimization.

[0027] Step S2: Multi-dimensional Real-time Data Acquisition and Preprocessing. Through a multi-dimensional data acquisition module, real-time acquisition, verification, and preprocessing of data across the entire data chain are completed, providing reliable data support for regulatory decisions. Specifically, this includes:

[0028] S21: Real-time acquisition of four major categories of data according to preset measurement points and acquisition frequency: Total real-time power of IT equipment, IT power consumption of the entire rack, temperature of server core components, and IT load utilization rate. Measurement points are located at the transformer low-voltage side main input point, UPS output end, rack head cabinet inlet end, rack intelligent PDU, and server out-of-band management system. Outdoor dry-bulb temperature, wet-bulb temperature, relative humidity, and atmospheric pressure are collected. Measurement points are located 1m away from the windward side of the data center building, away from the cooling equipment exhaust area. An equilateral triangle layout of three measurement points is used to take the average value to avoid interference from local heat island effects. Inlet / outlet medium temperature, operating power, and operating status of cold source equipment are collected; supply and return water temperature difference, flow rate, pump operating frequency, and valve opening on the distribution side; inlet / outlet medium temperature, return air temperature, and equipment operating power on the terminal side. Measurement points cover all equipment inlet / outlet, key pipeline nodes, and cavity monitoring points involved in temperature regulation. Hot and cold aisle temperature, rack inlet / outlet air temperature, computer room relative humidity, and condensation monitoring sensor data are collected at multiple locations.

[0029] S22: The accuracy of power metering instruments shall not be lower than Class 1.0, the accuracy of temperature measuring instruments shall be ±0.5℃, and the accuracy of relative humidity measuring instruments shall be ±5%; the acquisition frequency of real-time control parameters shall not be lower than once per hour, and the acquisition frequency of key control parameters (such as total IT power consumption, liquid supply temperature, and chip temperature) shall be increased to once per minute to capture rapid changes in dynamic load.

[0030] S23: The collected raw data is filtered for outliers. Abnormal data with jumps are removed by the 3σ criterion. Fluctuating data is filtered by moving average. The window length is set according to the dynamic characteristics of the parameters. For example, power data is filtered by a 5-point moving average. At the same time, data cross-validation is performed, including thermodynamic consistency verification of supply and return water temperature difference, flow rate and heat exchange, and balance verification of total heat generation of IT and total heat exchange of cooling system. When the data error exceeds ±5%, a data anomaly alarm is triggered, automatic temperature control is suspended, manual intervention mode is switched, and the abnormal event is recorded to the system log.

[0031] Step S3: Short-term load and weather trend forecasting.

[0032] The load and weather forecasting module enables short-term prediction of IT load and weather parameters, allowing for proactive regulation. Specifically, this includes: S31: Short-term IT load forecasting. This is further broken down into the following steps: S31-1: Based on a multi-source feature dataset, extract fine-grained features such as historical IT power consumption sequences, business type labels (e.g., online transactions, offline training, data backup), and server hardware performance counters (e.g., CPU / GPU instruction reordering buffer occupancy, L2 / L3 cache miss rate, memory bandwidth occupancy, PCIe link throughput, NVLink bidirectional throughput). All features are normalized and then used to construct a multi-dimensional time-series feature matrix.

[0033] S31-2: A steady-state prediction model is constructed using a Long Short-Term Memory (LSTM) network. First, historical power consumption sequences are cleaned and normalized to remove noise and outliers. Then, an LSTM network structure is constructed. The input layer receives historical power consumption data over a continuous time window (e.g., 24 hours), and learns the long-short-term dependencies in the time series through forgetting, input, and output gate mechanisms. The hidden layers are set to three layers with 128 neurons each to enhance the model's expressive power. The output layer uses a fully connected network to generate hourly predictions for the next 24 hours. During model training, the Adam optimizer and mean squared error loss function are used to iteratively optimize the network weights using historical data from the past 180 days. Finally, the baseline load trend and steady-state prediction confidence interval for the next 1-24 hours are output.

[0034] S31-3: A burst prediction model is constructed using a Temporal Convolutional Network (TCN). First, fine-grained metrics such as hardware performance counters are temporally aligned and feature-engineered to construct a multi-dimensional input feature matrix containing data from the past 60 minutes. The TCN model employs causal convolution to ensure no leakage of future information, and expands the receptive field through dilated convolutions (with dilation coefficients set to 1, 2, 4, and 8) to capture long-term dependencies. Residual connections help alleviate the vanishing gradient problem. The input layer receives micro-architectural features from 60 time steps, which are processed through four convolutional layers and a ReLU activation function. The output layer generates the probability of load mutations and the estimated burst value for the next 15-60 minutes. During model training, a focus loss function is used to enhance attention to a few mutation samples, improving the early warning capability for burst loads.

[0035] S31-4: Using the Bayesian fusion method, the prediction confidence of the two channels is dynamically weighted to generate the final IT load prediction curve (including hourly values ​​for the next 1-24 hours and 15-minute granular interpolation) and load mutation probability heatmap. At the same time, the prediction confidence index (a value between 0 and 1, representing the reliability of the prediction result) for each prediction time is output.

[0036] S32: Connects to the local meteorological department's hourly weather forecast API, combines real-time data from local outdoor measuring points, and uses linear regression or Kalman filter fusion algorithms to construct a short-term meteorological parameter prediction model. It outputs the trend of outdoor dry-bulb and wet-bulb temperature changes over the next 1-24 hours, predicts the availability of natural cold sources, and provides forward-looking guidance for cooling mode switching and temperature setting.

[0037] Step S4: Multi-objective energy efficiency optimization decision-making under safety constraints. Based on the prediction data provided in Step S3, a predictive time-domain rolling optimization framework is constructed using the multi-objective optimization decision-making module to achieve forward-looking planning of cooling capacity supply. Furthermore, a simplified thermodynamic model of the cooling system is established, explicitly embedding the system's thermal inertia into the optimization problem to achieve dynamic planning of the liquid supply temperature trajectory. Specifically, this includes: S41: Constructing the optimization objective function, with minimizing the total energy consumption of the entire cooling system as the core optimization objective. The objective function is: ,in, The total energy consumption of the entire cooling system. This represents the total energy consumption of the chiller unit / refrigeration compressor. This represents the total energy consumption of the cooling tower, dry cooler, and evaporative air cooler. Total energy consumption of various circulating pump distribution equipment. This refers to the total energy consumption of terminal air conditioning and liquid cooling circulation equipment. For valves, auxiliary equipment, and other energy consumption.

[0038] S42: Clearly define the constraints for optimization. All solution processes must meet the following constraints to ensure safe and compliant control: Rigid safety constraints: The temperature safety thresholds for IT equipment and cooling system equipment pre-configured in step S11 cannot be exceeded; Thermodynamic balance constraints: The total cooling capacity of the cooling system ≥ the total heat generation of IT equipment + cooling loss, and a cooling margin of not less than 20% is reserved; Equipment operation constraints: The operating frequency, load rate, and start-stop frequency of equipment such as chillers, pumps, and fans must meet the rated requirements of the equipment to avoid frequent start-stops; Temperature change rate constraints: The adjustment rate of chilled water / coolant temperature ≤ 1℃ / 10 minutes to avoid rapid temperature fluctuations that could lead to instability in the computer room environment; Condensation protection constraints: The temperature of the cooling medium must be at least 2℃ higher than the dew point temperature of the computer room environment.

[0039] S43: A rolling optimization solution is performed using a Model Predictive Control (MPC) framework, further refined into the following steps: S43-1: First, the cooling system is physically simplified based on the lumped parameter method. Core components such as the chiller, plate heat exchanger, cooling tower, circulation pipeline, and terminal heat exchange equipment are considered as independent lumped heat capacity nodes, with each node represented by its average temperature. Second, differential equations for each node are established based on the law of conservation of energy. To adapt to computer solving, these equations are discretized, resulting in the following recursive relationship: ,in, Let be the temperature of the i-th node at time k. The discrete time step is 60 seconds. Let i be the heat capacity of node i (unit: J / K). The thermal resistance between adjacent nodes i and j (unit: K / W). Let be the heat source power at node i (e.g., power consumption of IT equipment, heat dissipation, etc.). For a typical water-cooled chiller system, take N=6 nodes: chip node (i=1), cold plate node (i=2), liquid supply pipeline node (i=3), liquid return pipeline node (i=4), heat exchanger node (i=5), and ambient air node (i=6), with corresponding heat capacity parameters as follows: The thermal resistance between adjacent nodes is A total of 11 physical parameters were identified. Then, using a system identification method, the unknown parameters were estimated online using historical operating data collected in step S2, and the model parameters were updated every 15 minutes using recursive least squares. Finally, the identified parameters were integrated into a thermodynamic model parameter library, enabling the model to accurately predict the dynamic response relationship between the liquid supply temperature and the chip temperature, providing a mathematical expression of the system's dynamic characteristics for the MPC optimizer.

[0040] S43-2: Based on the prediction confidence level output in step S3, dynamically adjust the prediction time domain length and temperature change rate constraint of MPC. Set the basic prediction time domain Pbase=4 (corresponding to 60 minutes). When the prediction confidence level > 0.8, extend the prediction time domain to P=6 steps (90 minutes) and allow a faster temperature change rate (e.g., ≤1℃ / 5 minutes) to fully utilize the system's thermal inertia reserve cooling capacity. When the prediction confidence level < 0.5, shorten the prediction time domain to P=2 steps (30 minutes) and tighten the temperature change rate (e.g., ≤1℃ / 15 minutes) to enhance system robustness. Linear interpolation is used to adjust parameters in the intermediate confidence level range.

[0041] S43-3: In each control cycle (15 minutes), based on the current state, load forecast curve, meteorological forecast, and thermodynamic model, and with the rigid condition that the chip temperature (or equivalent rack inlet air temperature) in the future time domain does not exceed the safety constraint, a mixed-integer nonlinear programming (MINLP) algorithm is used to solve for the globally optimal control sequence. Decision variables include the chilled water / coolant supply temperature setpoint trajectory (values ​​for the next P time domains), equipment start-up / shutdown combinations and load allocation (0-1 variables), and the switching ratio between natural cooling and mechanical refrigeration (continuous variables). The solver uses an open-source or commercial solver (such as IPOPT or Gurobi), with a maximum solution time limit of 60 seconds. If the timeout occurs, the feasible solution from the previous cycle is used as a backup.

[0042] S43-4: The instruction sequence of the first control cycle obtained by solving (including water supply temperature setpoint, equipment frequency, valve opening, etc.) is issued and executed as the optimal temperature control instruction sequence, and the remaining time domain instructions are stored in the cache as backups for use in the initialization of the next cycle.

[0043] Step S5: Closed-Loop Execution and Real-Time Feedback Correction. Through the closed-loop execution control module, based on the optimized decision output control parameters, hierarchical execution and closed-loop correction are completed. This ensures that the actual operating effect meets expectations when dealing with prediction deviations and instantaneous disturbances, serving as a supplementary defense to pre-emptive control. Specifically, this includes: S51: Setting differentiated adjustment cycles for different types of control parameters to avoid frequent adjustments. For large-scale parameters such as water / liquid supply temperature setpoints and the number of operating devices, updates are performed once per MPC command cycle. For small-scale parameters such as pump frequency, fan frequency, and valve opening, the adjustment cycle is ≥1 minute. Precise differential pressure, flow rate, and temperature closed-loop control is achieved through a PID controller. PID parameters are pre-tuned according to equipment characteristics and stored in the controller.

[0044] S52: Based on real-time collected system operation data and IT equipment temperature data, compare the deviation between the control target value and the actual operating value; when the deviation exceeds the set threshold, trigger PID closed-loop fine-tuning, generate closed-loop correction deviation amount and add it to the execution instruction; when the actual IT load suddenly increases by more than 20% of the predicted value or the chip temperature approaches the safety threshold, immediately trigger MPC re-optimization, use the current state to update the prediction and resolve the control sequence, realize the dual protection of prediction and correction, and avoid temperature exceeding the standard.

[0045] S53: Based on the temperature setpoint adjusted by temperature variation, synchronous linkage between air and water control is implemented. For example, when the water supply temperature rises, the terminal fan speed is reduced and the water pump frequency is increased according to the preset coordination curve to maintain heat exchange, achieving coordinated energy saving across the entire chain from cold source to distribution and terminal. The coordination curve is obtained through prior debugging or simulation optimization.

[0046] Step S6: Hierarchical safety interlocking and emergency control. Through the safety interlocking and emergency control module, full-process safety protection is achieved, avoiding safety accidents caused by control errors, equipment failures, and sudden changes in operating conditions. In particular, it provides fallback protection for predicted failure scenarios that may occur in Step S5.

[0047] Specifically, this includes: S61: A graded safety interlock protection mechanism, establishing a three-level interlock system, deeply integrated with the temperature control system: Level 1 Early Warning Mechanism: When the monitored parameter reaches 90% of the safety threshold, an early warning is triggered, the optimization and adjustment of the temperature parameters are suspended, the current operating parameters are maintained, and an early warning message is pushed to the maintenance personnel; Level 2 Over-limit Protection Mechanism: When the monitored parameter exceeds the safety warning value (i.e., reaches 95% of the threshold), automatic temperature control is automatically suspended, switched to safety priority mode, and the cooling temperature is immediately reduced by 1℃ at a preset step size until the parameter falls back, the output of the cooling source equipment is increased, and the parameter is quickly pulled back to the safe range; the automatic temperature control mode can be gradually restored only after the parameter has recovered to the safe range and stabilized for more than 30 minutes; Level 3 Emergency Interlock Mechanism: In the event of extreme safety risk scenarios such as a major leakage in the cooling system, the temperature of IT equipment soaring to the downtime threshold, or the triggering of a fire alarm signal, the temperature control system is immediately locked, switched to emergency manual mode, and preset emergency protection actions are triggered to prioritize the safety of IT equipment and personnel.

[0048] S62: Typical Abnormal Operating Condition Emergency Control Strategy: When the IT load suddenly increases by more than 20% in a short period of time, this is a typical scenario where the predictive model may fail. Immediately suspend optimization and rapidly lower the cooling temperature based on the sudden increase in load, increase the output of the cooling source equipment, and ensure sufficient cooling supply. After the load stabilizes, restart the optimization algorithm to adjust the variable temperature parameters. When the outdoor temperature and humidity suddenly rise sharply, immediately adjust the cooling source side operation strategy, prioritize heat exchange efficiency, and appropriately tighten the water supply temperature adjustment range to avoid insufficient cooling. When the outdoor temperature suddenly drops sharply, prioritize triggering antifreeze protection and adjust the operating parameters of the cooling tower and dry cooler (such as increasing the fan speed). To prevent pipes and heat exchange equipment from freezing and cracking, the cooling temperature is gradually increased to optimize energy consumption. When a cooling device fails and is taken out of service, redundant equipment is immediately started, and the variable temperature parameters are adjusted to appropriately lower the cooling temperature to ensure that the total cooling capacity meets the demand and reserves a safety margin. Before the faulty equipment is repaired, the upper limit of the temperature adjustment is locked to avoid cooling risks. The real-time dew point temperature of the computer room is monitored. When the difference between the cooling medium temperature and the dew point temperature is less than 2°C, the temperature increase control is immediately stopped, and the cooling temperature setpoint is lowered if necessary. At the same time, the fresh air and dehumidification strategies of the computer room are adjusted to reduce the ambient dew point temperature. The temperature increase operation shall not be resumed until the risk of condensation is eliminated. All interlock actions generate interlock protection status records, including trigger time, trigger parameters, action content, and recovery time, and are fed back to the monitoring system for archiving.

[0049] Step S7: Verify the effect of regulation and iteratively optimize the strategy.

[0050] Through the effect verification and iterative optimization module, the energy-saving effect is quantitatively verified, data is archived, and strategies are iterated. In particular, continuous optimization of the prediction model improves prediction accuracy and fundamentally strengthens the ability to address lag issues. Specifically, this includes: S71: Quantitative verification of energy-saving effect, using industry-standard unified evaluation indicators to quantify the energy-saving effect of variable temperature control. Core evaluation indicators include: Power Usage Effectiveness (PUE), cooling system energy consumption ratio, and natural cold source utilization ratio. The statistical scope strictly adheres to national standards, covering all power consumption of IT equipment, air conditioning and refrigeration equipment, power supply and distribution systems, and auxiliary facilities, with no omissions or duplicates. The verification method uses a concurrent comparison method, comparing the energy consumption differences between the variable temperature control mode and the traditional constant temperature control mode under the same IT load and weather conditions to objectively verify the energy-saving effect. The comparison period is at least 30 consecutive days to eliminate the impact of short-term fluctuations.

[0051] S72: Establish a complete data archive for temperature control, including real-time collected raw data, control parameters output by the algorithm, actual operating effects after control, early warning and abnormal event records, and energy consumption statistics (daily / weekly / monthly reports). The data storage period shall be no less than 3 years, and a hybrid storage method of time-series database (such as InfluxDB) and relational database (such as MySQL) shall be used to meet the needs of algorithm iteration, effect review, and compliance audit.

[0052] S73: Based on historical operational data and control effects, the load forecasting model and energy efficiency optimization model are retrained and their parameters optimized monthly to improve forecast accuracy and optimization precision. Specifically, this includes: automatically triggering a model retraining task on the 1st of each month, retraining the LSTM and TCN models using historical data from the past three months, and selecting the optimal hyperparameters through cross-validation; updating equipment curves in the energy efficiency MAP library; and immediately updating the equipment energy efficiency MAP library and thermodynamic model parameters and recalibrating the model when significant changes occur in the data center, such as IT equipment mounting / removal, cooling equipment upgrades, or server room layout adjustments. All optimized parameters are stored in the knowledge base as iterative optimization parameters for the next stage of control, and version numbers are recorded for backtracking.

[0053] This method is based on a variable-temperature cooling energy consumption control system for a data center cooling system. The system includes: a demand database construction module, a multi-dimensional data acquisition module, a load and weather forecasting module, a multi-objective optimization decision-making module, a closed-loop execution control module, a safety interlock and emergency control module, and an effect verification and iterative optimization module. The modules interact with each other via standard industrial communication protocols (such as Modbus TCP / IP, BACnet, and OPCUA), specifically including:

[0054] Demand Database Construction Module: Used to analyze the distribution of cooling demand in data centers and build a demand database that includes temperature safety thresholds for IT equipment, safety thresholds for cooling system equipment, and energy efficiency targets.

[0055] Specifically, the demand database construction module spatially partitions the data processing equipment according to the temperature distribution gradient of the data center space, obtaining M temperature space intervals; it performs load tracking on the data processing equipment within the M temperature space intervals to obtain the load of each temperature space interval; based on the temperature characteristics and load of the M temperature space intervals, it establishes a relationship between load and spatial heat accumulation calculation, and obtains the heat accumulation value of each temperature space interval through heat accumulation calculation; it determines the cooling demand of each temperature space interval based on the heat accumulation value, and analyzes the correspondence between the M temperature space intervals, load, and cooling demand to construct the demand database.

[0056] Multi-dimensional data acquisition module: used to acquire multi-dimensional real-time operational data from the data center and construct a multi-source feature dataset that has been time-aligned and validated.

[0057] Specifically, the multi-dimensional data acquisition module collects IT load data, meteorological environment data, cooling system full-link data, and data center environment data at preset measurement points, with a collection frequency of no less than once per minute, and key control parameters are encrypted to once per minute; outlier filtering and data cross-validation are performed on the collected raw data, and a data anomaly alarm is triggered when the data error exceeds ±5%.

[0058] Load and Weather Forecasting Module: Based on multi-source feature datasets, this module performs steady-state load forecasting and sudden load forecasting using a dual-channel forecasting architecture, outputting load forecast curves with forecast confidence levels.

[0059] Specifically, the load and weather forecasting module includes a steady-state load forecasting channel and a sudden load forecasting channel. The steady-state load forecasting channel uses a long short-term memory network to output the load trend baseline and steady-state forecast confidence interval for the next 1-24 hours. The sudden load forecasting channel uses a temporal convolutional network and outputs the load mutation probability, estimated mutation increment, and sudden forecast confidence level for the next 15-60 minutes based on server hardware performance counter indicators. The outputs of the two channels are dynamically weighted by a Bayesian fusion method to generate the final IT load forecasting curve and load mutation probability heatmap.

[0060] Multi-objective optimization decision module: Based on the load forecast curve and safety constraints in the demand database, it uses an adaptive model predictive control framework with an embedded thermodynamic model to solve for the optimal temperature regulation command sequence in a rolling manner.

[0061] Specifically, the multi-objective optimization decision module establishes a simplified lumped parameter thermodynamic model of the cooling system, treating the core components as independent lumped heat capacity nodes. It estimates the heat capacity parameters of each node and the thermal resistance parameters between adjacent nodes through system identification methods. Based on the prediction confidence, it dynamically adjusts the prediction time domain length and temperature change rate constraints of the model predictive control. In each control cycle, based on the current state, load prediction curve, and thermodynamic model, and with the chip temperature not exceeding the safety constraint in the future time domain as a rigid condition, it uses a mixed integer nonlinear programming algorithm to solve for the globally optimal control sequence. The instruction sequence of the first control cycle is then issued and executed as the optimal temperature regulation instruction sequence.

[0062] Closed-loop execution control module: used to adjust the cooling system parameters according to the optimal temperature control command sequence, and perform closed-loop correction based on real-time feedback to generate closed-loop correction deviation.

[0063] Specifically, the closed-loop execution control module performs adjustments according to the model predictive control command cycle for slowly changing parameters such as water supply temperature setpoint and number of operating equipment, and performs PID closed-loop control on a minute-level cycle for rapidly changing parameters such as pump frequency, fan frequency, and valve opening. When the operating deviation exceeds the set threshold, PID fine-tuning is triggered to generate a closed-loop correction deviation amount and add it to the execution command. When the actual IT load suddenly increases by more than 20% of the predicted value or the chip temperature approaches the safety threshold, model predictive control re-optimization is immediately triggered.

[0064] Safety interlock and emergency control module: used to activate the hierarchical safety interlock mechanism and output the interlock protection status when the monitored parameters trigger the safety threshold.

[0065] Specifically, the safety interlock and emergency control module establishes a three-level interlock system: when the monitored parameter reaches 90% of the safety threshold, a level one warning is triggered, suspending the optimization and adjustment of the temperature parameters; when the monitored parameter exceeds the safety warning value, a level two over-limit protection is triggered, switching to the safety priority mode and forcibly reducing the cooling temperature; when there are extreme safety risks such as a major leakage in the cooling system, the temperature of IT equipment soaring to the shutdown threshold, or the fire alarm signal being triggered, a level three emergency interlock is triggered, locking the temperature control system and switching to emergency manual mode.

[0066] The effect verification and iterative optimization module is used to retrain the prediction model based on historical running data and update the thermodynamic model parameters, and output iterative optimization parameters.

[0067] Specifically, the effect verification and iterative optimization module uses the same-period comparison method to quantitatively verify energy-saving indicators such as PUE and CLF; the load forecasting model and energy efficiency optimization model are retrained and their parameters are optimized every month; when IT equipment changes or cooling equipment is upgraded in the data center, the equipment energy efficiency MAP library and thermodynamic model parameters are updated immediately to form a closed-loop evolution capability.

[0068] Example 1: Implementation of variable temperature cooling energy consumption control for water-cooled chiller unit cooling system.

[0069] This embodiment is applied to a data center in North China. The data center has a total building area of ​​12,000 square meters and a total installed capacity of 8 MW of IT equipment. The cooling system consists of 3 water-cooled magnetic levitation centrifugal chiller units (each with a cooling capacity of 2800 kW, N+1 redundancy configuration) + 4 open cooling towers (each with a water flow rate of 400 m³ / h) + a terminal precision air conditioning system (room-level, totaling 48 units). The design PUE target is ≤1.25, and the actual operating load rate is approximately 65%. The specific implementation process is as follows:

[0070] Pre-configuration constraints and baselines: The ambient temperature of the main equipment room is set to 22℃±2℃, the relative humidity to 50%±10%, and the supply air temperature to 19℃, which is more than 2℃ above the dew point temperature; the chilled water supply temperature is set within the range of 7℃-15℃, and the cooling water supply and return temperatures are set to 32℃ / 37℃; the cooling system has a 25% cooling margin. Complete the configuration of all safety constraint baselines according to step S1.

[0071] Data Acquisition and Prediction: A data acquisition gateway is deployed to collect data from the power monitoring system (total IT power consumption), smart meters (chimney, water pump, and fan power consumption), temperature sensors (supply and return water temperatures, outdoor temperature and humidity), flow meters, etc., via the Modbus TCP protocol. Data is collected once per minute, generating a multi-source feature dataset. A dual-channel prediction architecture is adopted: The steady-state channel uses LSTM based on three years of historical power consumption data to predict the daily load curve for the next 24 hours, with the model retrained every 24 hours; the burst channel collects hardware performance counters such as CPU instruction queue depth and memory bandwidth utilization through the server's out-of-band management interface, and uses a TCN model to predict the probability and confidence level of load surges in the next 15-60 minutes. The outputs of the two channels are fused using Bayesian methods to generate a load prediction curve with confidence intervals and a heatmap of burst probability.

[0072] Temperature regulation decision-making and execution under different operating conditions: An adaptive MPC framework is deployed on an industrial control server, with a control cycle of 15 minutes and a basic prediction time domain of 4 steps (1 hour). Dynamic adjustment is based on prediction confidence: when confidence > 0.8, the prediction time domain expands to 6 steps, and the temperature change rate is relaxed to ≤1℃ / 5 minutes; when confidence < 0.5, the prediction time domain shrinks to 2 steps, and the change rate is tightened to ≤1℃ / 15 minutes. A lumped-parameter thermodynamic model of the cooling system (N=6 nodes) is established, and the results are obtained through system identification. and There are 11 parameters in total, which are embedded in the MPC optimizer. The optimizer uses the cabinet inlet air temperature (which is dynamically determined by the liquid supply temperature and load) as a constraint to directly plan the future trajectory of the chilled water supply temperature.

[0073] High-temperature conditions (≥30℃): On a typical summer day, the IT load peak is predicted to occur between 2-4 pm (predicted peak 7.2MW, confidence level 0.85). Under the constraint of temperature change rate, MPC gradually lowers the supply water temperature from 12℃ to 8℃ 1.5 hours in advance, utilizing the thermal inertia of the water system to store cooling capacity. At the same time, it optimizes the cooling water outlet temperature to approach the wet-bulb temperature, and the chiller COP increases from 5.8 to 6.4.

[0074] Medium and high temperature conditions (20-30℃): During the transitional season, the nighttime load is predicted to decrease, so the water supply temperature is gradually increased from 10℃ to 13℃ in advance, and the supply and return water temperature difference is optimized from 5℃ to 7℃ to reduce the power consumption of the chilled water pump.

[0075] Low temperature conditions (<20℃): In winter, if the outdoor wet-bulb temperature is predicted to be more than 2℃ lower than the return water temperature for the next 8 hours, the plate heat exchanger will be turned on in advance for natural cooling. As the temperature drops, the supply water temperature will be gradually increased to 16℃-18℃. The chiller unit will be completely shut down, and cooling will be provided only by the plate heat exchanger and cooling tower.

[0076] Safety Interlock and Effect Verification: When the actual IT load suddenly increases by more than 22% of the predicted value on a certain day, MPC re-optimization is immediately triggered. Within 3 minutes, the solution is recalculated and the water supply temperature is lowered by 1.5℃, and the maximum deviation of the rack intake air temperature is controlled within 1.2℃. After adopting variable temperature control, the annual rack intake air temperature fluctuation is reduced from ±3℃ to within ±1℃, the annual energy consumption of the cooling system is reduced by 22%, the utilization time of natural cold source is increased from 35% to 62%, and the annual average PUE measured value is 1.13, meeting the advanced value requirements of local standards.

[0077] Example 2: Implementation of variable temperature cooling energy consumption control for immersion single-phase liquid cooling system.

[0078] This embodiment is applied to the intelligent computing data center of a large internet company in East China. The center employs single-phase immersion liquid cooling technology, deploying a total of 2000 AI training servers, with a single rack power density of 30kW and a total IT capacity of 60MW. The cooling system includes 12 immersion chambers (each accommodating 16 servers), 12 primary-side dry coolers (each with a heat dissipation capacity of 500kW), 6 secondary-side liquid cooling circulation pumps (4 in operation and 2 on standby), 4 plate heat exchangers, and 2 backup chillers (for extreme high-temperature conditions). The design PUE target is ≤1.1. The specific implementation process is as follows:

[0079] Pre-configured constraints: liquid phase temperature difference ≤ 5℃, liquid supply temperature setting range is 25℃-45℃, chip junction temperature upper limit is 85℃ (5℃ safety redundancy is reserved, actual control target ≤ 80℃), cooling system reserves 20% margin, all constraint parameters are configured according to step S1 to form a safety constraint baseline.

[0080] Data Acquisition and Prediction: CPU / GPU chip temperature and power consumption are collected every 30 seconds via the server's BMC interface. Inlet and outlet liquid temperatures, flow rates, dry cooler fan frequencies, and water pump frequencies are collected via PLC, at a frequency of once per minute. Addressing the bursty characteristics of AI training loads, the burst channel focuses on collecting 12 hardware performance counters, including GPU utilization, memory bandwidth, gradient synchronization frequency, and NVLink throughput. A load surge prediction model is constructed using TCN to output IT power consumption and chip temperature predictions for the next 15-60 minutes, along with prediction confidence levels. The steady-state channel uses LSTM based on 30 days of historical power consumption data to predict the daily load baseline. The dual-channel fusion output includes prediction results with confidence levels, which are updated every 5 minutes.

[0081] Temperature regulation decision-making and execution under different operating conditions: Adaptive MPC is adopted with a control cycle of 10 minutes and a basic prediction time domain of 4 steps (40 minutes). Dynamic adjustment based on confidence level: When the confidence level is >0.8, the prediction time domain is expanded to 6 steps (60 minutes), and the temperature change rate is relaxed to ≤1℃ / 5 minutes; when the confidence level is <0.5, the prediction time domain is reduced to 2 steps (20 minutes), and the change rate is tightened to ≤1℃ / 15 minutes. A lumped parameter thermodynamic model of the liquid cooling system is established (N=5 nodes: chip node, coolant node, cavity wall node, heat exchanger node, and outdoor air node). Five heat capacity parameters and seven thermal resistance parameters (a total of 12 physical parameters) are identified through system identification. The future trajectory of the supply liquid temperature is planned directly with the chip junction temperature as a constraint.

[0082] High-temperature operation (≥30℃): On a summer day, a training task was predicted to start in 15 minutes, causing a sudden increase in load (predicted power consumption increased from 45kW / cabinet to 58kW / cabinet, confidence level 0.9). The MPC preemptively lowered the liquid supply temperature from 32℃ to 25℃ (in 3 steps, decreasing by 2.3℃ every 5 minutes), utilizing the high heat capacity of the liquid cooling system to store cold energy, and increasing the dry cooler fan frequency from 40Hz to 50Hz. When the load peak arrived, the chip temperature reached a maximum of 78.5℃, but did not reach the 80℃ threshold.

[0083] Medium temperature conditions (10-30℃): In spring and autumn, based on the prediction of the stable load period, the liquid supply temperature will be increased to 32℃-35℃, relying entirely on the natural cooling of the dry cooler, and the standby chiller will be kept off.

[0084] Low temperature conditions (<10℃): In winter, while ensuring the chip temperature is ≤78℃, the liquid supply temperature is increased to 42℃-45℃, the dry cooler fan is reduced to 20Hz, and the circulation pump frequency is optimized from 45Hz to 35Hz (flow rate is reduced by 22%, pump power consumption is reduced by 35%).

[0085] Safety Interlock and Performance Verification: When the deviation between the predicted and measured chip temperature reaches 3.5℃ (exceeding the 3℃ threshold), MPC re-optimization is triggered. After resolving the problem, the liquid supply temperature setpoint is fine-tuned, and subsequent deviations are controlled within 1.5℃. After adopting variable temperature control, the chip temperature fluctuation amplitude is reduced by 40%, the annual cooling system CLF (cooling load factor) remains stable within 0.08, and the annual average PUE is 1.08, achieving the design energy efficiency target.

[0086] The technical scope of this invention is not limited to the content described above. Those skilled in the art can make various modifications and variations to the above embodiments without departing from the technical concept of this invention, and all such modifications and variations should fall within the protection scope of this invention.

Claims

1. A method for controlling the energy consumption of variable-temperature cooling in a data center cooling system, characterized in that, This includes: analyzing the cooling demand distribution of the data center and constructing a demand database containing temperature safety thresholds for IT equipment, safety thresholds for cooling system equipment, and energy efficiency targets; acquiring multi-dimensional real-time operational data of the data center and constructing a multi-source feature dataset that has undergone time-series alignment and validity verification; based on the multi-source feature dataset, performing steady-state load prediction and burst load prediction respectively through a dual-channel prediction architecture, and outputting a load prediction curve with prediction confidence; based on the load prediction curve and the safety constraints in the demand database, using an adaptive model predictive control framework with an embedded thermodynamic model to solve for the optimal variable temperature control command sequence; executing cooling system parameter adjustments according to the optimal variable temperature control command sequence, and performing closed-loop correction based on real-time feedback to generate a closed-loop correction deviation; and when the monitored parameters trigger safety thresholds, activating a graded safety interlocking mechanism and outputting the interlocking protection status.

2. The method according to claim 1, characterized in that, The method for analyzing the cooling demand distribution of the data center includes: spatially partitioning the data processing equipment according to the temperature distribution gradient of the data center space to obtain M temperature space intervals; performing load tracking on the data processing equipment within the M temperature space intervals to obtain the load of each temperature space interval; establishing a relationship between load and spatial heat accumulation calculation based on the temperature characteristics of the M temperature space intervals and the load, and obtaining the heat accumulation value of each temperature space interval through heat accumulation calculation; determining the cooling demand of each temperature space interval based on the heat accumulation value, and analyzing the correspondence between the M temperature space intervals, load, and cooling demand to construct the demand database.

3. The method according to claim 2, characterized in that, The construction of the demand database also includes: fitting the cooling penalty probability of data processing devices in M ​​temperature space intervals based on historical tracking samples, and converting it into cooling penalty coefficients for each temperature space region; and setting demand weights for the demand database using the cooling penalty coefficients for each temperature space region.

4. The method according to claim 1, characterized in that, The dual-channel prediction architecture performs steady-state load prediction and burst load prediction separately, including: constructing a steady-state load prediction channel based on historical IT load operation data, using a long short-term memory network to output the load trend baseline and steady-state prediction confidence interval for the next 1-24 hours; constructing a burst load prediction channel based on server hardware performance counter indicators, using a temporal convolutional network to output the probability of load mutation, estimated burst increment, and burst prediction confidence for the next 15-60 minutes; and using a Bayesian fusion method to dynamically weight the prediction confidence of the two channels to generate the final IT load prediction curve and load mutation probability heatmap, while outputting the prediction confidence at each prediction time.

5. The method according to claim 4, characterized in that, The server hardware performance counter metrics include one or more of the following: CPU instruction reordering buffer utilization, cache miss rate, memory bandwidth utilization, PCIe link throughput, GPU utilization, video memory bandwidth, gradient synchronization frequency, and NVLink throughput.

6. The method according to claim 1, characterized in that, The adaptive model predictive control framework, which employs an embedded thermodynamic model, uses a rolling solution to obtain the optimal temperature regulation command sequence. This includes: establishing a simplified lumped-parameter thermodynamic model of the cooling system, treating the core components of the chiller, plate heat exchanger, cooling tower, circulation pipeline, and terminal heat exchange equipment as independent lumped heat capacity nodes; estimating the heat capacity parameters of each node and the thermal resistance parameters between adjacent nodes using a system identification method to form a thermodynamic model parameter library; dynamically adjusting the prediction time domain length and temperature change rate constraint of the model predictive control based on the prediction confidence level; extending the prediction time domain and allowing a faster temperature change rate when the prediction confidence is high, and shortening the prediction time domain and tightening the temperature change rate when the prediction confidence is low; and in each control cycle, based on the current state, load prediction curve, and thermodynamic model, using a mixed-integer nonlinear programming algorithm to solve for the globally optimal control sequence, with the chip temperature not exceeding the safety constraint in the future time domain as a rigid condition, and issuing the command sequence of the first control cycle as the optimal temperature regulation command sequence for execution.

7. The method according to claim 6, characterized in that, The simplified lumped parameter thermodynamic model of the cooling system includes N lumped heat capacity nodes. The temperature recursion relationship of each node is determined by the node heat capacity, the thermal resistance between adjacent nodes, the node heat source power, and the discrete time step. The heat capacity parameters and thermal resistance parameters are estimated and dynamically updated online by recursive least squares method or extended Kalman filter.

8. The method according to claim 1, characterized in that, The process of adjusting cooling system parameters according to the optimal variable temperature control command sequence and performing closed-loop correction based on real-time feedback includes: adjusting slowly varying parameters such as water supply temperature setpoint and number of operating devices according to the model predictive control command cycle; performing PID closed-loop control on a minute-level cycle for rapidly varying parameters such as pump frequency, fan frequency, and valve opening; triggering PID fine-tuning when the operating deviation exceeds the set threshold, generating a closed-loop correction deviation amount and adding it to the execution command; and immediately triggering model predictive control re-optimization and updating the control parameters when the actual IT load suddenly increases by more than 20% of the predicted value or the chip temperature approaches the safety threshold.

9. The method according to claim 1, characterized in that, The activation of the graded safety interlock mechanism includes: triggering a level one warning when the monitored parameter reaches 90% of the safety threshold, suspending the optimization and adjustment of the temperature parameters; triggering a level two over-limit protection when the monitored parameter exceeds the safety warning value, switching to the safety priority mode and forcibly reducing the cooling temperature; and triggering a level three emergency interlock when there are extreme safety risks such as a major leakage in the cooling system, the temperature of IT equipment soaring to the shutdown threshold, or the fire alarm signal being triggered, locking the temperature control system and switching to emergency manual mode.

10. A variable-temperature cooling energy consumption control system for a data center cooling system, characterized in that, The system includes: a demand database construction module for analyzing the cooling demand distribution of the data center and constructing a demand database containing temperature safety thresholds for IT equipment, safety thresholds for cooling system equipment, and energy efficiency targets; a multi-dimensional data acquisition module for acquiring multi-dimensional real-time operational data of the data center and constructing a multi-source feature dataset that has undergone time-series alignment and validity verification; a load and weather forecasting module for performing steady-state load forecasting and burst load forecasting respectively based on the multi-source feature dataset using a dual-channel forecasting architecture, and outputting a load forecasting curve with prediction confidence; a multi-objective optimization decision-making module for using an adaptive model predictive control framework with an embedded thermodynamic model to solve for the optimal variable temperature control command sequence based on the load forecasting curve and the safety constraints in the demand database; a closed-loop execution control module for adjusting cooling system parameters according to the optimal variable temperature control command sequence and performing closed-loop correction based on real-time feedback, generating a closed-loop correction deviation; a safety interlocking and emergency control module for activating a graded safety interlocking mechanism when the monitored parameters trigger safety thresholds and outputting the interlocking protection status; and an effect verification and iterative optimization module for retraining the prediction model based on historical operational data and updating the thermodynamic model parameters, and outputting iterative optimization parameters.