Intelligent control method and system for immersion liquid cooling system based on AI
By using an AI-based intelligent control method for immersion liquid cooling systems, the coordinated optimization of the cooling system and the IT computing system is achieved. This solves the problem of poor overall energy efficiency caused by the independent operation of the cooling system and the IT computing system in existing technologies, reduces the total energy consumption of the system, and improves operational safety and adaptability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANGZHOU SUANXIANG TECHNOLOGY CO LTD
- Filing Date
- 2026-03-20
- Publication Date
- 2026-04-17
AI Technical Summary
In existing immersion liquid cooling systems, the control of the cooling system and the IT computing system are independent of each other, each pursuing local optima, which makes it impossible to achieve the overall energy efficiency optimization of the data center.
An AI-based intelligent control method for immersion liquid cooling systems is adopted. By collecting multi-dimensional operating status data in real time, performing feature fusion processing, generating a unified status feature vector, and using an AI collaborative optimization model to output joint control commands, the system achieves collaborative optimization of IT equipment performance status and cooling subsystem parameters, and constructs a global optimization target.
It achieves a globally optimal match between IT computing and cooling, reduces total system energy consumption, improves system security and adaptability, and has online learning capabilities to continuously optimize control strategies.
Smart Images

Figure CN121879541A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data center heat dissipation technology, and in particular to an AI-based intelligent control method and system for immersion liquid cooling systems. Background Technology
[0002] Immersion liquid cooling technology has become a key heat dissipation solution for high-performance computing, artificial intelligence training and edge data centers due to its advantages such as high heat dissipation density, low energy consumption and low noise.
[0003] In existing immersion liquid cooling systems, control strategies often employ PID control based on fixed thresholds or simple rule-based control. For example, when the coolant outlet temperature exceeds a set value, the circulation pump speed is increased or auxiliary cooling equipment is activated; when the chip temperature is too high, the CPU / GPU frequency is reduced via the server management interface.
[0004] The main problem with existing control methods is that the control of the cooling system and the IT computing system are independent of each other, each pursuing only local optimization. For example, the cooling system aims to reduce the coolant temperature to the target value with the minimum power consumption, while the IT system aims to maximize computing energy efficiency while ensuring business performance. The two goals are disconnected, making it impossible to optimize the overall energy efficiency of the data center. Summary of the Invention
[0005] The purpose of this application is to provide an AI-based intelligent control method and system for immersion liquid cooling systems, which enables collaborative optimization of the cooling system and computing load, minimizing the total energy consumption of the system while ensuring the safe temperature of the equipment.
[0006] Firstly, this application provides an AI-based intelligent control method for an immersion liquid cooling system, comprising: Real-time acquisition of multi-dimensional operating status data of the immersion liquid cooling system, including load data and temperature data of IT equipment in the immersion tank, and operating parameters of the cooling subsystem; Multi-dimensional operational status data is fused to generate a unified status feature vector. Using the state feature vector as input, a set of joint control instructions is output through a preset AI collaborative optimization model. The joint control instructions include a first control instruction for adjusting the performance state of IT equipment and a second control instruction for adjusting the operating parameters of the cooling subsystem. The AI collaborative optimization model is configured to perform optimization search in a joint action space with the total energy efficiency of the immersion liquid cooling system as the target. The joint action space consists of IT equipment performance adjustment actions and cooling subsystem parameter adjustment actions. The immersion liquid cooling system is controlled and regulated based on joint control commands.
[0007] By using the above technical solutions, the control barriers between the IT system and the cooling system are broken down, and a unified AI collaborative optimization model with the overall system energy efficiency as the global optimization goal is constructed. This optimization model performs collaborative search and decision-making in a joint action space composed of IT equipment adjustment actions and cooling adjustment actions, realizing the global optimal matching of IT computing and cooling heat dissipation, which helps to reduce the overall total energy consumption.
[0008] Optionally, the step of performing feature fusion processing on multi-dimensional operational state data to generate a unified state feature vector includes: Perform clock synchronization and outlier handling on multi-dimensional operational status data to generate time-aligned data sequences; By using a pre-defined method, advanced physical characteristics reflecting the thermodynamics, fluid dynamics, and energy efficiency of a system can be extracted or constructed from the data sequence. The basic data and advanced physical features in the data sequence are selected and concatenated to form an initial feature vector; The initial feature vector is standardized to generate a unified state feature vector.
[0009] Optionally, before outputting a set of joint control commands by taking the state feature vector as input and using a preset AI collaborative optimization model, the following steps are included: Based on the state feature vector, a system state prediction sequence is obtained through a pre-set digital twin model; Based on the system state prediction sequence, extract the trend features of the system state prediction sequence relative to the state feature vector; The state feature vector is fused with the trend feature to generate an enhanced state feature vector.
[0010] Optionally, fusing the state feature vector with the trend feature to generate an enhanced state feature vector includes: Calculate the attention weights of the trend features on each element in the state feature vector; Based on the attention weights, the state feature vectors are reweighted to generate weighted state feature vectors; The weighted state feature vector is concatenated with the trend feature to generate an enhanced state feature vector.
[0011] Optionally, the step of taking the state feature vector as input and outputting a set of joint control commands through a preset AI collaborative optimization model includes: Map the state feature vector to an action selection strategy in the joint action space; Based on state feature vectors and action selection strategies, a pre-set value assessment model is used to obtain corresponding value assessment results. The value assessment model is used to predict the long-term cumulative reward after performing a given action. Based on the value assessment results, the action selection strategy is optimized with the goal of maximizing long-term cumulative rewards, and a set of joint control action parameters is generated based on the optimized action selection strategy. The combined control action parameters are converted into executable first and second control commands.
[0012] Optionally, the step of taking the state feature vector as input and outputting a set of joint control commands through a preset AI collaborative optimization model includes: Time-series analysis is performed on multi-dimensional operational status data to identify the current load mode of the system, which includes at least steady-state calculation mode, intermittent burst mode, and training iteration fluctuation mode; Select an AI collaborative optimization sub-model that matches the load pattern from the preset strategy model library, and denot it as the exclusive optimization sub-model; Using the state feature vector as input, a set of joint control commands is output through a dedicated optimization sub-model.
[0013] The AI collaborative optimization model is pre-trained with corresponding sub-models or configured with differentiated reward function weights for each load mode; Based on the identified load patterns, the specific strategies of the AI collaborative optimization model are dynamically switched or adjusted to adapt to the optimization preferences under different load patterns.
[0014] Optionally, after controlling and adjusting the immersion liquid cooling system based on joint control commands, the process includes: Obtain the actual response data of the system after the joint control command is executed, and calculate the actual reward value based on the actual response data; The state feature vector, joint control action parameters, and actual reward value are used as new samples and stored in the training experience pool of the AI collaborative optimization model. Based on the updated experience pool, the AI collaborative optimization model is incrementally trained to generate an optimized AI collaborative optimization model.
[0015] Optionally, after controlling and adjusting the immersion liquid cooling system based on joint control commands, the method further includes: Acquire the actual system response data within a preset time window after the instruction is executed. The actual response data includes at least the total system power consumption, the temperature sequence of the IT equipment, and the action sequence of the actuator. Based on the actual response data, a multi-dimensional strategy effect quantification index is calculated, which includes at least an energy efficiency improvement index, a thermal safety index, and a control stability index. The associated data tuples, consisting of state feature vectors and multi-dimensional policy effect quantification indicators, are stored in a preset policy knowledge base.
[0016] Secondly, this application provides an AI-based intelligent control system for an immersion liquid cooling system, comprising: The data acquisition module 101 is used to collect multi-dimensional operating status data of the immersion liquid cooling system in real time. The multi-dimensional operating status data includes load data and temperature data of IT equipment in the immersion tank, and operating parameters of the cooling subsystem. The feature fusion processing module 102 is used to perform feature fusion processing on multi-dimensional running status data to generate a unified state feature vector. The control command generation module 103 is used to take the state feature vector as input and output a set of joint control commands through a preset AI collaborative optimization model. The joint control commands include a first control command for adjusting the performance state of IT equipment and a second control command for adjusting the operating parameters of the cooling subsystem. The AI collaborative optimization model is configured to perform optimization search in a joint action space with the total energy efficiency of the immersion liquid cooling system as the target. The joint action space consists of IT equipment performance adjustment actions and cooling subsystem parameter adjustment actions. The instruction control and adjustment module 104 is used to control and adjust the immersion liquid cooling system based on joint control instructions.
[0017] Thirdly, this application provides a computer-readable storage medium storing a computer program that can be loaded by a processor and executed as described above regarding an AI-based intelligent control method for an immersion liquid cooling system.
[0018] In summary, this application firstly achieves a globally optimal match between IT computing and cooling by using an AI model for collaborative optimization in the joint action space, effectively reducing the system's total energy consumption. Secondly, by combining digital twin prediction and trend analysis, it can identify and intervene in potential hotspots in advance, transforming passive response into proactive prevention and greatly improving the security of system operation. Furthermore, it possesses online learning and knowledge accumulation capabilities, enabling it to automatically adapt to dynamic changes in the system, continuously optimize control strategies, and reduce long-term operation and maintenance costs. Attached Figure Description
[0019] Figure 1 This is a flowchart of an AI-based intelligent control method for an immersion liquid cooling system provided in an embodiment of this application; Figure 2 This is a flowchart provided in this application embodiment for performing feature fusion processing on multi-dimensional operating state data to generate a unified state feature vector; Figure 3 This is a flowchart provided in the embodiments of this application that takes a state feature vector as input, outputs a set of joint control instructions through a preset AI collaborative optimization model; Figure 4This is a flowchart illustrating the online learning process of the collaborative optimization model after the execution of the joint control instructions provided in this application embodiment; Figure 5 This is a schematic diagram of an AI-based intelligent control system for an immersion liquid cooling system provided in an embodiment of this application. Detailed Implementation
[0020] The following is in conjunction with the appendix Figure 1 -Appendix Figure 5 This application will be described in further detail below.
[0021] This application provides an AI-based intelligent control method for an immersion liquid cooling system, see [link to relevant documentation]. Figure 1 This includes the following steps: S100: Real-time acquisition of multi-dimensional operating status data of immersion liquid cooling system.
[0022] S200: Perform feature fusion processing on multi-dimensional operating status data to generate a unified state feature vector.
[0023] S300 takes the state feature vector as input and outputs a set of joint control commands through a preset AI collaborative optimization model.
[0024] S400 controls and regulates the immersion liquid cooling system based on joint control commands.
[0025] In this embodiment of the application, multi-dimensional operating status data of the immersion liquid cooling system will be collected in real time.
[0026] The multi-dimensional operational status data includes load and temperature data of IT equipment in the immersion tank, and operating parameters of the cooling subsystem. The load data of the IT equipment includes: CPU / GPU utilization, computing task type, and power consumption. The temperature data includes: chip junction temperature, equipment surface temperature, and temperature field distribution of coolant in the tank. The operating parameters of the cooling subsystem include: coolant flow rate, pressure, temperature, viscosity, and level, pump speed, valve opening, operating power of heat exchange module, and inlet and outlet temperature difference.
[0027] After acquiring multi-dimensional operational status data, it is necessary to perform fusion feature processing on the data. The purpose is to transform the raw data from different physical domains (electricity, heat, fluid), different sampling rates, and different dimensions into a standardized feature representation with high information density.
[0028] Specifically, see Figure 2 The process involves feature fusion of multi-dimensional operational status data to generate a unified status feature vector, including the following steps: S210. Perform clock synchronization and outlier processing on multi-dimensional operating status data to generate time-aligned data sequences.
[0029] S220. Using a preset method, extract or construct advanced physical characteristics from the data sequence that reflect the thermodynamic, fluid dynamic, and energy efficiency status of the system.
[0030] S230. Select and concatenate the basic data and advanced physical features in the data sequence to form an initial feature vector.
[0031] S240. Standardize the initial feature vector to generate a unified state feature vector.
[0032] First, clock synchronization and outlier handling are performed on multi-dimensional operational status data. For example, all sensor data is timestamped using a precision clock protocol (such as PTP), and abnormal data points are removed or repaired using statistical methods (such as the 3σ principle) or known physical parameters within a reasonable range. This generates a time-aligned data sequence to ensure data consistency in the time dimension and avoid data misalignment caused by sensor response delays or communication jitter.
[0033] Then, advanced physical features will be extracted. Advanced physical features refer to parameters generated by calculating or logically processing the basic data based on the heat transfer and fluid dynamics principles of the immersion liquid cooling system. These parameters are used to quantitatively characterize the thermodynamic state, cooling efficiency, or cross-domain coupling relationship of the system, and are features that better reflect the essential state of the system.
[0034] Advanced physical characteristics may include real-time thermal resistance, pumping efficiency index, and local boiling state indicators. The pre-defined method here mainly refers to the sum of a series of mathematical equations, empirical correlations, and calculation rules used to describe the system's heat transfer, fluid flow, and phase change processes.
[0035] Among them, the real-time thermal resistance is used to dynamically characterize the overall heat dissipation capacity from the chip junction to the coolant. The smaller the value, the smoother the heat dissipation path. It can be obtained by calculating the ratio of the maximum chip junction temperature to the temperature difference of the coolant inlet to the total power consumption of the IT equipment.
[0036] The pumping efficiency index is used to quantify the efficiency of a coolant circulation pump in converting electrical energy into effective fluid work, in order to assess whether the pump's operating point is optimal. It can be obtained by calculating the ratio of the product of coolant density, gravitational acceleration, head, and flow rate to the pump's input electrical power.
[0037] Local boiling status indicator is used to accurately identify whether the coolant has undergone phase change boiling on the chip surface. Because boiling heat transfer is orders of magnitude different from single-phase flow, it has a great impact on the control strategy. It can be obtained by comparing the chip surface temperature and the coolant saturation temperature corresponding to the pressure at that point, calculating the superheat, and determining whether the superheat exceeds the initial boiling threshold.
[0038] Next, the basic data (such as CPU temperature, power consumption, pump speed, and inlet water temperature) in the data sequence are filtered with the calculated advanced physical features to remove highly redundant features, and the selected features are concatenated into a one-dimensional vector in a fixed order, which is denoted as the initial feature vector.
[0039] Finally, the initial feature vector is standardized (e.g., Z-score standardization or Min-Max normalization) to map the values of each dimension to a similar scale (e.g., mean 0, variance 1), generating a unified state feature vector to eliminate the adverse effects of differences in the scale and numerical range of different features on the subsequent optimization model.
[0040] After obtaining the state feature vector, the state feature vector can be used as input to output a set of joint control commands through a preset AI collaborative optimization model.
[0041] The AI collaborative optimization model is configured to perform optimization search in a joint action space with the total energy efficiency of the immersion liquid cooling system as the target. The joint action space consists of IT equipment performance adjustment actions and cooling subsystem parameter adjustment actions.
[0042] Specifically, see Figure 3 Using the state feature vector as input, and through a pre-set AI collaborative optimization model, a set of joint control commands is output, including the following steps: S310. Map the state feature vector to an action selection strategy in the joint action space.
[0043] S320: Based on state feature vectors and action selection strategies, obtain the corresponding value assessment results through a preset value assessment model.
[0044] S330. Based on the value assessment results, optimize the action selection strategy with the goal of maximizing long-term cumulative rewards, and generate a set of joint control action parameters based on the optimized action selection strategy.
[0045] S340, convert the combined control action parameters into executable first control instructions and second control instructions.
[0046] The AI collaborative optimization model is generated based on a reinforcement learning framework. Essentially, this model is a pre-trained deep reinforcement learning agent whose goal is to learn an optimal policy that enables the selected joint control actions to minimize the total energy consumption of the system under any system state while satisfying hard constraints such as temperature safety.
[0047] The AI collaborative optimization model comprises two core modules: a policy generation module, which generates all possible control actions and their probability distributions or directly outputs recommended actions based on the current system state; and a value evaluation module, which evaluates the long-term value of executing the suggested actions given by the policy generation module under the current system state through a set reward function.
[0048] The current system state is represented by the state feature vector. Therefore, the first step is to use the policy generation module to map the state feature vector into an action selection policy in the joint action space, which outputs a definite action vector. Here, t represents the start time of the current control cycle. The system operates with a fixed control cycle, and after each control cycle... To execute one complete control loop, therefore, here and In reality, it represents the state feature vector and the corresponding action vector at the start of the current control cycle (the same applies to subsequent ones).
[0049] Then, based on the state feature vector and action selection strategy, the corresponding value assessment results are obtained through a pre-set value assessment model.
[0050] The value assessment model is used to predict the long-term cumulative reward after performing a given action. This model is trained with the goal of setting a reward function. During the training phase, the objective is to maximize the long-term cumulative expected value of the set reward function, learned through extensive trial and error with the system environment. After training, for any input state feature vector... and suggested actions The valuation model can output a scalar. This scalar value represents the value of the action performed. Predictive assessment of the long-term value it can bring.
[0051] The reward function is constructed as a weighted sum of the negative correlation function of the total system power consumption and the temperature exceedance penalty function. It can be represented as: in, This is an energy consumption penalty term, representing the total power consumption of the system. The temperature penalty function is... , , Indicates the first The core temperature of each chip (or temperature monitoring point), For temperature safety threshold, As a temperature safety penalty measure, it is essential to ensure that the temperature of all monitoring chips remains within a safe range. If the temperature of any chip exceeds the limit, a penalty function will be applied. Negative rewards are generated, with α and β being weighting coefficients used to balance the importance of energy saving and safety.
[0052] It is worth noting that, and It is not an independent variable, but rather an action performed in period t. After that, the system enters a new state. The core of value network assessment is the result value that is measured only at that time. Perform an action in a state The expected long-term cumulative R, which includes the future... , ...prediction.
[0053] Based on the value assessment results, the action selection strategy is optimized with the goal of maximizing long-term cumulative rewards, and a set of joint control action parameters is generated based on the optimized action selection strategy.
[0054] This involves using the predictive information provided by the value assessment model to fine-tune the initial action selection strategy so that the resulting actions can obtain higher predictive long-term value. In other words, the abstract goal of "maximizing long-term cumulative reward" is transformed into a specific and computable optimization problem to find the optimal policy network parameters that maximize its expected value.
[0055] For example, using policy gradient as an optimization algorithm, the policy network parameters are updated through the gradient direction. The gradient direction indicates the minute change in the action vector in the current state. How will the long-term cumulative reward Q change for each component of the network? Through iteration, the optimal policy network parameters can be obtained.
[0056] Based on the optimal strategy network parameters, a set of joint control action parameters can be generated. The joint control action parameters are specific, digital instruction vectors that directly specify the target adjustment amount of each controlled object in the IT equipment and cooling subsystem within the current control cycle.
[0057] Finally, the joint control action parameters are converted into executable first and second control instructions, that is, the instruction vector is decoded into specific commands that can drive the actuators in the physical world. The joint control instructions include a first control instruction for adjusting the performance state of IT equipment and a second control instruction for adjusting the operating parameters of the cooling subsystem.
[0058] In this way, by using the AI collaborative optimization model, the control problem is formalized into a policy optimization problem with value prediction as its core. By leveraging the powerful function approximation capability of deep neural networks, an intelligent policy that can directly output globally approximate optimal collaborative control commands is learned.
[0059] In order to improve the foresight of decision-making and transform passive response into proactive prevention, this application embodiment will also use a preset digital twin model to simulate the system state within a preset time period in the future and obtain a system state prediction sequence.
[0060] Specifically, before outputting a set of joint control commands by using the state feature vector as input and through a preset AI collaborative optimization model, the following steps are also included: S301. Based on the state feature vector, obtain the system state prediction sequence through a preset digital twin model.
[0061] S302. Based on the system state prediction sequence, extract the trend features of the system state prediction sequence relative to the state feature vector.
[0062] S303. Fuse the state feature vector with the change trend feature to generate an enhanced state feature vector.
[0063] Among them, the preset digital twin model is a virtual mapping of the real system. It is usually simplified from computational fluid dynamics, heat transfer and circuit thermal models. By receiving the current state as the initial condition, it can quickly simulate the dynamic response of the system under different control strategies.
[0064] First, by inputting the state feature vector and using a pre-defined digital twin model, a system state prediction sequence can be obtained.
[0065] Then, from the system state prediction sequence, the trend features of the system state prediction sequence relative to the state feature vector (current state) can be extracted. The trend features are characterized as key trend indicators that have a significant impact on control optimization decisions, such as "the rise of the highest predicted junction temperature in the next 5 minutes", "the average slope of the total power consumption change", and "the cumulative time when the predicted temperature exceeds the safety threshold".
[0066] Finally, by fusing the state feature vector with the trend feature, an enhanced state feature vector can be generated.
[0067] Specifically, the state feature vector is fused with the trend feature to generate an enhanced state feature vector, including the following steps: S3031. Calculate the attention weight of the trend feature on each element in the state feature vector.
[0068] S3032. Based on the attention weights, reweight the state feature vector to generate a weighted state feature vector.
[0069] S3033. The weighted state feature vector is concatenated with the trend feature to generate an enhanced state feature vector.
[0070] The state feature vector and the trend feature are fused, mainly by using an attention mechanism. First, the attention weight of the trend feature on each element of the state feature vector is calculated.
[0071] Then, based on the attention weights, the state feature vectors are reweighted to generate weighted state feature vectors. That is, the original features that are highly correlated with future trends are preserved or even amplified, while irrelevant features are suppressed.
[0072] For example, if a trend feature indicates that "CPU2 will overheat in 5 minutes", then the attention mechanism will automatically increase the weight of features related to CPU2 heat dissipation in the current state (such as the current power consumption of CPU2 and the fluid flow rate in its vicinity).
[0073] Finally, the weighted state feature vector is concatenated with the trend feature to generate an enhanced state feature vector.
[0074] In addition, in order to adapt to the diverse operating scenarios of the system, the load modes are classified and identified in this embodiment of the application, and the control strategy is optimized for the characteristics of different load modes, taking into account both energy efficiency and stability.
[0075] Specifically, using the state feature vector as input, and through a pre-set AI collaborative optimization model, a set of joint control commands is output, which also includes the following steps: S350 performs time-series analysis on multi-dimensional operating status data to identify the current load mode of the system.
[0076] S360. Select an AI collaborative optimization sub-model that matches the load mode from the preset strategy model library, and denot it as the exclusive optimization sub-model.
[0077] S370 takes the state feature vector as input and outputs a set of joint control commands through a dedicated optimization sub-model.
[0078] The load modes include at least steady-state computing mode, intermittent burst mode, and training iteration fluctuation mode. Steady-state computing mode: the load is stable and changes slowly; intermittent burst mode: the load fluctuates drastically periodically; training iteration fluctuation mode: the load fluctuates in a regular sawtooth pattern.
[0079] First, by performing real-time time-series analysis on multi-dimensional operational status data, such as using a sliding window to calculate the mean, variance, and spectral characteristics of the load, the current load mode of the system can be identified.
[0080] Then, from the preset strategy model library, select the AI collaborative optimization sub-model that matches the load mode and record it as the exclusive optimization sub-model.
[0081] The preset strategy model library stores multiple AI collaborative optimization sub-models associated with load modes. Alternatively, it can be viewed as AI collaborative optimization models pre-training specialized sub-models for each typical load mode. Different sub-models are configured with differentiated reward function weights. For example, in intermittent burst mode, more emphasis is placed on temperature safety weights; in stable computing mode, more emphasis is placed on energy consumption weights, pursuing ultimate energy efficiency.
[0082] Finally, based on the state feature vector, joint control commands are generated according to a dedicated optimization sub-model adapted to the current load mode.
[0083] In this way, the system can dynamically switch or adjust the model strategy according to the identified load pattern to adapt to the optimization preferences under different load patterns, which can avoid the performance bottleneck of a single model in complex scenarios and improve the system's adaptability.
[0084] After obtaining the joint control command, the immersion liquid cooling system can be controlled and regulated based on the joint control command.
[0085] In addition, to achieve continuous self-optimization of the system, a closed loop of online learning and knowledge accumulation will be added.
[0086] Online learning primarily focuses on verifying the effectiveness of control commands through actual results after execution. This allows for the determination of the validity of the joint control commands generated by the AI collaborative optimization model, enabling targeted updates and optimizations of the AI collaborative optimization model.
[0087] Specifically, see Figure 4 After controlling and regulating the immersion liquid cooling system based on the joint control commands, the following steps are also included: S510: Obtain the actual response data of the system after the joint control command is executed, and calculate the actual reward value based on the actual response data.
[0088] S520. The state feature vector, joint control action parameters, and actual reward value are used as new samples and stored in the training experience pool of the AI collaborative optimization model.
[0089] S530: Based on the updated experience pool, incremental training is performed on the AI collaborative optimization model to generate an optimized AI collaborative optimization model.
[0090] First, obtain the actual response data of the system after the joint control command is executed, such as the actual temperature and power consumption changes, and calculate the actual reward value through the reward function based on the actual response data.
[0091] That is, after the command is executed, a stable observation window is set, for example, from 30 seconds to 60 seconds after the command is executed, to obtain the average system power consumption (IT equipment power consumption + cooling system power consumption), which can be denoted as... The average temperature of each monitoring chip can be denoted as: With other parameters remaining unchanged, by substituting them into the reward function mentioned above, the actual reward value can be obtained.
[0092] The actual reward value represents the actual benefit brought by this action. The smaller the negative value, the greater the penalty and the worse the effect.
[0093] A reward threshold is set, such as 80% of the average reward of recent experience, in order to avoid learning strategies that have obviously failed. If the actual reward value is not less than the set reward threshold, it means that the execution strategy can be used as a subsequent learning sample. Therefore, the state feature vector, joint control action parameters and actual reward value are used as new samples and stored in the training experience pool of the AI collaborative optimization model.
[0094] Finally, periodically or when the experience pool reaches a certain capacity, batch data is sampled from the training experience pool to incrementally train the AI collaborative optimization model and update its network parameters (mainly the policy network and value assessment model). In this way, the model can adapt to the slow changes in the system itself and the environment, such as coolant performance degradation, equipment aging, and load mode changes.
[0095] Although AI collaborative optimization models can be continuously updated and optimized through online learning, the essence of AI collaborative optimization models is to explore the unknown. When faced with a newly deployed system or a sudden load pattern that has never been seen before, the AI model needs time to re-explore, and its initial performance may be poor, requiring retraining or lengthy fine-tuning.
[0096] Therefore, by constructing a strategy knowledge base as a guiding mechanism, and by actively using historical experience (knowledge) to constrain safety boundaries and guide efficiency, the exploration time can be significantly shortened and the quality of decision-making can be improved.
[0097] Specifically, after obtaining the actual response data of the system after the joint control command is executed, the following steps are also included: S511. Obtain the actual system response data within a preset time window after the instruction is executed.
[0098] S512. Calculate multi-dimensional strategy effectiveness quantification indicators based on actual response data.
[0099] S513. Store the associated data tuple consisting of state feature vectors and multi-dimensional strategy effect quantification indicators into the preset strategy knowledge base.
[0100] First, it will obtain the actual system response data within a preset time window after the command is executed.
[0101] The actual response data includes at least the total system power consumption, the temperature sequence of IT equipment, and the action sequence of actuators; the preset time window here is also regarded as the effect evaluation time window, ensuring that the system reaches a relatively stable new state before evaluation.
[0102] Then, based on the actual response data, multi-dimensional strategy effectiveness quantification indicators are calculated, including at least energy efficiency improvement indicators, thermal safety indicators, and control stability indicators.
[0103] Among them, the energy efficiency improvement index is calculated based on the rate of change of the total power consumption of the system before and after execution and the rate of change of the IT task throughput. It is used to represent energy saving. The larger the value, the more significant the energy saving effect. The thermal safety index is calculated based on the highest proximity of the IT equipment temperature to the safety threshold and the duration of overheating after execution. It is used to represent the minimum margin of the hottest point from the safety threshold. The larger the value, the higher the safety margin. The control stability index is calculated based on the variance or frequency of change of the actuator action sequence. It is a value between 0 and 1. The closer the value is to 1, the smoother the actuator action is, the more stable the system operation is, which is beneficial to the equipment life.
[0104] Finally, the associated data tuples consisting of state feature vectors and multi-dimensional strategy effect quantification indicators are stored in the preset strategy knowledge base. The strategy knowledge base construction involves associating and storing state features, executed actions, and effect indicators to form a searchable knowledge unit.
[0105] In subsequent control decisions, the current state feature vector is matched with the historical state feature vectors stored in the policy knowledge base for similarity. If the matching degree exceeds a preset threshold, the historical joint control action parameters and effect quantification indicators associated with it are retrieved from the policy knowledge base and output as reference information for the current AI collaborative optimization model to optimize decisions.
[0106] This application also provides an AI-based intelligent control system for an immersion liquid cooling system, see [link to relevant documentation]. Figure 5 The system includes: a data acquisition module 101, a feature fusion processing module 102, a control command generation module 103, and a command control adjustment module 104.
[0107] Among them, the data acquisition module 101 is used to collect multi-dimensional operating status data of the immersion liquid cooling system in real time.
[0108] The feature fusion processing module 102 is used to perform feature fusion processing on multi-dimensional operating status data to generate a unified state feature vector.
[0109] The control command generation module 103 is used to take the state feature vector as input and output a set of joint control commands through a preset AI collaborative optimization model.
[0110] The instruction control and adjustment module 104 is used to control and adjust the immersion liquid cooling system based on joint control instructions.
[0111] In this embodiment of the application, the data acquisition module 101 is specifically used to collect multi-dimensional operating status data of the immersion liquid cooling system in real time. The multi-dimensional operating status data includes load data and temperature data of IT equipment in the immersion tank, and operating parameters of the cooling subsystem.
[0112] The feature fusion processing module 102 is specifically used to perform feature fusion processing on the multi-dimensional running status data acquired by the data acquisition module 101 to generate a unified state feature vector.
[0113] The control command generation module 103 is specifically used to take the state feature vector generated by the feature fusion processing module 102 as input, and output a set of joint control commands through a preset AI collaborative optimization model.
[0114] The instruction control and adjustment module 104 is specifically used to control and adjust the immersion liquid cooling system based on the joint control instructions generated by the control instruction generation module 103.
[0115] This application also provides a computer-readable storage medium storing a computer program that can be loaded by a processor and executed by any of the above-described AI-based intelligent control methods for immersion liquid cooling systems.
[0116] The embodiments described in this application are preferred embodiments of this application and are not intended to limit the scope of protection of this application. Therefore, all equivalent changes made in accordance with the principles of this application should be included within the scope of protection of this application.
Claims
1. An AI-based intelligent control method for an immersion liquid cooling system, characterized in that, include: Real-time acquisition of multi-dimensional operating status data of the immersion liquid cooling system, including load data and temperature data of IT equipment in the immersion tank, and operating parameters of the cooling subsystem; Multi-dimensional operational status data is fused to generate a unified status feature vector. Using the state feature vector as input, a set of joint control instructions is output through a preset AI collaborative optimization model. The joint control instructions include a first control instruction for adjusting the performance state of IT equipment and a second control instruction for adjusting the operating parameters of the cooling subsystem. The AI collaborative optimization model is configured to perform optimization search in a joint action space with the total energy efficiency of the immersion liquid cooling system as the target. The joint action space consists of IT equipment performance adjustment actions and cooling subsystem parameter adjustment actions. The immersion liquid cooling system is controlled and regulated based on joint control commands.
2. The intelligent control method for an AI-based immersion liquid cooling system according to claim 1, characterized in that, The process of performing feature fusion processing on multi-dimensional operational status data to generate a unified status feature vector includes: Perform clock synchronization and outlier handling on multi-dimensional operational status data to generate time-aligned data sequences; By using a pre-defined method, advanced physical characteristics reflecting the thermodynamics, fluid dynamics, and energy efficiency of a system can be extracted or constructed from the data sequence. The basic data and advanced physical features in the data sequence are selected and concatenated to form an initial feature vector; The initial feature vector is standardized to generate a unified state feature vector.
3. The intelligent control method for an AI-based immersion liquid cooling system according to claim 1, characterized in that, Before the step of taking the state feature vector as input and outputting a set of joint control commands through a preset AI collaborative optimization model, the following steps are included: Based on the state feature vector, a system state prediction sequence is obtained through a pre-set digital twin model; Based on the system state prediction sequence, extract the trend features of the system state prediction sequence relative to the state feature vector; The state feature vector is fused with the trend feature to generate an enhanced state feature vector.
4. The intelligent control method for an AI-based immersion liquid cooling system according to claim 3, characterized in that, The process of fusing the state feature vector with the trend feature to generate an enhanced state feature vector includes: Calculate the attention weights of the trend features on each element in the state feature vector; Based on the attention weights, the state feature vectors are reweighted to generate weighted state feature vectors; The weighted state feature vector is concatenated with the trend feature to generate an enhanced state feature vector.
5. The intelligent control method for an AI-based immersion liquid cooling system according to claim 1, characterized in that, The process involves taking a state feature vector as input, using a pre-defined AI collaborative optimization model, and outputting a set of joint control commands, including: Map the state feature vector to an action selection strategy in the joint action space; Based on state feature vectors and action selection strategies, a pre-set value assessment model is used to obtain corresponding value assessment results. The value assessment model is used to predict the long-term cumulative reward after performing a given action. Based on the value assessment results, the action selection strategy is optimized with the goal of maximizing long-term cumulative rewards, and a set of joint control action parameters is generated based on the optimized action selection strategy. The combined control action parameters are converted into executable first and second control commands.
6. The intelligent control method for an AI-based immersion liquid cooling system according to claim 1, characterized in that, The process involves taking a state feature vector as input, using a pre-defined AI collaborative optimization model, and outputting a set of joint control commands, including: Time-series analysis is performed on multi-dimensional operational status data to identify the current load mode of the system, which includes at least steady-state calculation mode, intermittent burst mode, and training iteration fluctuation mode; Select an AI collaborative optimization sub-model that matches the load pattern from the preset strategy model library, and denot it as the exclusive optimization sub-model; Using the state feature vector as input, a set of joint control commands is output through a dedicated optimization sub-model; The AI collaborative optimization model is pre-trained with corresponding sub-models or configured with differentiated reward function weights for each load mode; Based on the identified load patterns, the specific strategies of the AI collaborative optimization model are dynamically switched or adjusted to adapt to the optimization preferences under different load patterns.
7. The intelligent control method for an AI-based immersion liquid cooling system according to claim 1, characterized in that, After controlling and adjusting the immersion liquid cooling system based on joint control commands, the following steps are included: Obtain the actual response data of the system after the joint control command is executed, and calculate the actual reward value based on the actual response data; The state feature vector, joint control action parameters, and actual reward value are used as new samples and stored in the training experience pool of the AI collaborative optimization model. Based on the updated experience pool, the AI collaborative optimization model is incrementally trained to generate an optimized AI collaborative optimization model.
8. The intelligent control method for an AI-based immersion liquid cooling system according to claim 1, characterized in that, After controlling and regulating the immersion liquid cooling system based on joint control commands, the system further includes: Acquire the actual system response data within a preset time window after the instruction is executed. The actual response data includes at least the total system power consumption, the temperature sequence of the IT equipment, and the action sequence of the actuator. Based on the actual response data, a multi-dimensional strategy effect quantification index is calculated, which includes at least an energy efficiency improvement index, a thermal safety index, and a control stability index. The associated data tuples, consisting of state feature vectors and multi-dimensional policy effect quantification indicators, are stored in a preset policy knowledge base.
9. An AI-based intelligent control system for an immersion liquid cooling system, characterized in that, include: The data acquisition module (101) is used to collect multi-dimensional operating status data of the immersion liquid cooling system in real time. The multi-dimensional operating status data includes load data and temperature data of IT equipment in the immersion tank, and operating parameters of the cooling subsystem. The feature fusion processing module (102) is used to perform feature fusion processing on multi-dimensional running status data to generate a unified state feature vector; The control instruction generation module (103) is used to take the state feature vector as input and output a set of joint control instructions through a preset AI collaborative optimization model. The joint control instructions include a first control instruction for adjusting the performance state of IT equipment and a second control instruction for adjusting the operating parameters of the cooling subsystem. The AI collaborative optimization model is configured to perform optimization search in a joint action space with the total energy efficiency of the immersion liquid cooling system as the target. The joint action space consists of IT equipment performance adjustment actions and cooling subsystem parameter adjustment actions. The instruction control and adjustment module (104) is used to control and adjust the immersion liquid cooling system based on joint control instructions.
10. A computer-readable storage medium storing a computer program capable of being loaded by a processor and executing an AI-based intelligent control method for an immersion liquid cooling system as described in any one of claims 1 to 8.
Citation Information
Patent Citations
CPU junction temperature control method and system
CN119493436A
Data center IT load and cooling system cooperative control method based on TD3 algorithm
CN120743544A
Data center liquid cooling accurate temperature control method and system
CN121116028A
Compressor energy-saving operation control method and system based on reinforcement learning
CN121165467A
Data center global temperature control optimization method and system based on hybrid reinforcement learning
CN121503286A