Liquid cooling resource intelligent distribution regulation and control method and system
By collecting and processing operation data in the liquid cooling system, establishing a model and using reinforcement learning and adaptive algorithms to dynamically adjust the cooling parameters, the problem of uneven allocation of cooling resources in the liquid cooling system is solved, cooling efficiency is improved, energy consumption is reduced, and the stability of the system is ensured.
Patent Information
- Application Number
- CN202510176478.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-07-04
AI Technical Summary
Traditional liquid cooling systems cannot dynamically adjust the flow rate, flow rate and pressure of the coolant according to actual operating conditions, resulting in insufficient cooling in high-load areas or excessive cooling in low-load areas, uneven allocation of cooling resources, affecting heat dissipation efficiency and increasing energy consumption.
By collecting the operating data of the target computing power cluster, preprocessing it and establishing a model, combining reinforcement learning and adaptive algorithms, dynamically adjusting the flow rate, flow rate and pressure of the coolant, optimizing the heat exchanger temperature and fan speed, and achieving intelligent regulation.
It realizes intelligent and dynamic management of computing power cluster cooling resources, improves cooling efficiency, reduces energy consumption, and ensures the stable operation of the system.
Smart Images

Figure CN120264673A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent regulation and control, and particularly to an intelligent allocation and regulation method and system for liquid cooling resources. Background Art
[0002] With the rapid development of big data, artificial intelligence, and cloud computing technologies, the demand for computing power has increased sharply. Especially in high-density computing environments, the thermal management of computing power clusters has become a key challenge. Traditional air cooling systems are no longer able to meet the heat dissipation requirements of high-density computing nodes, and liquid cooling systems have gradually become the preferred cooling solution for large data centers and supercomputing centers due to their high thermal conductivity performance. Liquid cooling systems can effectively reduce the temperature of computing nodes and improve system stability and energy efficiency. However, the design and management of liquid cooling systems still face some problems, including uneven distribution of cooling resources, low cooling efficiency, and energy consumption waste.
[0003] In current liquid cooling systems, parameters such as the flow rate, flow volume, and pressure of the coolant often cannot be dynamically adjusted according to the actual operating conditions, resulting in insufficient cooling in high-load areas or excessive cooling in low-load areas. Traditional cooling solutions also do not achieve intelligent dynamic allocation and regulation, and cannot optimize resource allocation based on real-time load, temperature, and other data of the computing power cluster. Therefore, how to achieve efficient and intelligent allocation and regulation of liquid cooling resources in large-scale computing power clusters has become a key technical issue for improving the operating efficiency of computing power clusters, reducing energy consumption, and ensuring system stability. Summary of the Invention
[0004] The purpose of this part is to outline some aspects of the embodiments of the present invention and briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this part, as well as in the abstract and title of the present application, to avoid obscuring the purpose of this part, the abstract, and the title. However, such simplifications or omissions shall not be used to limit the scope of the present invention.
[0005] In view of the above existing problems, the present invention is proposed.
[0006] Therefore, the present invention provides an intelligent allocation and regulation method and system for liquid cooling resources, which can solve the problems mentioned in the background art.
[0007] To solve the above technical problems, the present invention provides the following technical solutions:
[0008] In the first aspect, the present invention provides an intelligent allocation and regulation method for liquid cooling resources, including:
[0009] Collect the first operation data of the target computing power cluster, and perform a first preprocessing on the first operation data to obtain second operation data;
[0010] Build a first model, and evaluate the operating status of the target computing power cluster based on the second operating data in combination with the first model;
[0011] Perform regulation according to the operating status in combination with a first regulation strategy, where the first regulation strategy includes a first objective function and constraint conditions;
[0012] Solve the first objective function to obtain an optimal regulation plan.
[0013] As a preferred solution of the intelligent allocation and regulation method for liquid cooling resources of the present invention, wherein: the first model includes:
[0014] The first model is any model used to simulate the heat generation and conduction process of computing power nodes in the target computing power cluster;
[0015] The operating status of the target computing power cluster at least includes cooling efficiency, heat generation amount of nodes, and utilization rate of liquid cooling resources.
[0016] As a preferred solution of the intelligent allocation and regulation method for liquid cooling resources of the present invention, wherein: the first regulation strategy includes:
[0017] The first regulation strategy is any regulation strategy that minimizes the total energy consumption of the liquid cooling system on the premise of meeting the cooling requirements of all nodes;
[0018] The minimization of the total energy consumption of the liquid cooling system at least includes the energy consumption of the liquid cooling pump and the penalty cost for node overheating.
[0019] As a preferred solution of the intelligent allocation and regulation method for liquid cooling resources of the present invention, wherein: the collection of the first operating data of the target computing power cluster includes:
[0020] Deploy temperature sensors at the computing power nodes of the target computing power set, collect the surface temperature of the nodes, and obtain heat generation data;
[0021] Install pressure sensors, flow velocity sensors, and flow meters in the liquid cooling pipeline to monitor the flow state of the coolant and obtain coolant state data;
[0022] Install energy consumption sensors on the liquid cooling pump and heat exchanger to obtain equipment operating status data;
[0023] The first operating data at least includes heat generation data, coolant state data, and equipment operating status data.
[0024] As a preferred solution of the intelligent allocation and regulation method for liquid cooling resources of the present invention, wherein: the first preprocessing includes:
[0025] Establish a first reinforcement learning model, which is used to dynamically adjust the arrangement and working frequency of sensors according to the first operation data;
[0026] Use denoising technology, i.e., moving average filtering, to remove interference signals from the first operation data processed by the first reinforcement learning model, and filter out abnormal data by means of the method based on mean and standard deviation.
[0027] As a preferred solution of the intelligent allocation and regulation method for liquid cooling resources of the present invention, wherein: the first regulation strategy further includes assigning cooling priorities to each node.
[0028] As a preferred solution of the intelligent allocation and regulation method for liquid cooling resources of the present invention, wherein: the obtaining of the optimal regulation plan by solving the first objective function includes:
[0029] Dynamically adjust the coolant outlet temperature of the heat exchanger according to the solution result, and optimize the heat dissipation efficiency by adjusting the fan speed;
[0030] Adopt an adaptive algorithm to continuously adjust the coolant flow rate and pipeline flow direction according to the feedback data for dynamic adjustment results, and adjust the first regulation strategy by means of the adaptive algorithm.
[0031] In a second aspect, the present invention provides an intelligent allocation and regulation system for liquid cooling resources, including:
[0032] A data acquisition and processing module, which is used to acquire the first operation data of the target computing power cluster and perform first preprocessing on the first operation data to obtain second operation data;
[0033] A model establishment module, which is used to establish a first model and evaluate the operation state of the target computing power cluster based on the second operation data in combination with the first model;
[0034] A regulation module, which is used to perform regulation according to the operation state in combination with the first regulation strategy, and the first regulation strategy includes a first objective function and constraint conditions;
[0035] A solving module, which is used to solve the first objective function to obtain an optimal regulation plan.
[0036] In a third aspect, the present invention provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of the method described above are implemented.
[0037] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method described above are implemented.
[0038] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention proposes a method and system for intelligent allocation and regulation of liquid cooling resources, which collect the first operation data of a target computing power cluster, perform a first preprocessing on the first operation data to obtain second operation data; establish a first model, and evaluate the operation state of the target computing power cluster based on the second operation data in combination with the first model; perform regulation according to the operation state in combination with a first regulation strategy, where the first regulation strategy includes a first objective function and constraint conditions; solve the first objective function to obtain an optimal regulation plan. It can realize the intelligent and dynamic management of the cooling resources of the computing power cluster, ensuring efficient and safe operation.
[0039] By collecting and processing the operation data of the target computing power cluster in real time, the system can accurately evaluate the operation state of the cluster and perform intelligent regulation according to the preset regulation strategy. Among them, the data collection and processing module, as the core part of the system, is responsible for collecting and processing various sensor data from the computing power cluster, providing reliable data support for subsequent analysis and regulation. By establishing an accurate mathematical model, the system can simulate the heat generation and conduction process of the computing power nodes, thereby accurately judging key indicators such as the cooling efficiency of the cluster, the heat generation amount of the nodes, and the utilization rate of liquid cooling resources. In terms of the regulation strategy, the system aims to minimize the total energy consumption of the liquid cooling system on the premise of meeting the cooling requirements of all nodes, which includes the energy consumption of the liquid cooling pump and the penalty cost caused by overheating of the nodes. To achieve this goal, the system adopts advanced algorithms and technical means, such as reinforcement learning models, denoising techniques, adaptive algorithms, etc., to ensure the accuracy and effectiveness of the regulation plan. By continuously optimizing parameters such as the flow rate, flow volume, and pressure of the coolant, and dynamically adjusting measures such as the coolant outlet temperature of the heat exchanger and the fan speed, the system can significantly improve the cooling efficiency of the computing power cluster, reduce energy consumption, and ensure the stable operation of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. Among them:
[0041] Figure 1 is a method flow chart of a method and system for intelligent allocation and regulation of liquid cooling resources provided by an embodiment of the present invention;
[0042] Figure 2 is an internal structure diagram of a computer device of a method and system for intelligent allocation and regulation of liquid cooling resources provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] To make the above objects, features, and advantages of the present invention more apparent and understandable, the following will describe in detail the specific embodiments of the present invention with reference to the accompanying drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0044] Embodiment 1
[0045] Referring to Figure 1 - Figure 2 , which is the first embodiment of the present invention. This embodiment provides a method and system for intelligent allocation and regulation of liquid cooling resources, including:
[0046] In the existing related technologies, there are some problems. For example, the cooling resource allocation is unreasonable, resulting in excessive cooling in some areas and insufficient cooling in some areas. This not only affects the heat dissipation efficiency of the system but also increases unnecessary energy consumption.
[0047] This application provides a method that can effectively solve the above-mentioned problems. Next, multiple embodiments will be combined to elaborate in detail how to implement the method for intelligent allocation and regulation of liquid cooling resources;
[0048] Figure 1 shows a method flow chart of a method for intelligent allocation and regulation of liquid cooling resources, including:
[0049] S101, collect the first operation data of the target computing power cluster, and perform a first preprocessing on the first operation data to obtain the second operation data;
[0050] In an optional embodiment, the target computing power cluster is several groups of computing power nodes configured in a liquid cooling system, and these nodes cooperate with each other to execute complex computing tasks.
[0051] In an optional embodiment, the first operation data is collected in real time by various sensors deployed on the computing power nodes, liquid cooling pipelines, liquid cooling pumps, and heat exchangers. These data include but are not limited to the surface temperature of the nodes, the flow state of the coolant, and the energy consumption of the equipment, etc. They jointly reflect the real-time operation state of the target computing power cluster.
[0052] In the embodiments of this application, collecting the first operation data of the target computing power cluster includes:
[0053] Deploy temperature sensors on the computing power nodes of the target computing power set to collect the surface temperature of the nodes and obtain heat generation data;
[0054] Install pressure sensors, flow velocity sensors, and flow meters in the liquid cooling pipeline to monitor the flow state of the coolant and obtain coolant state data;
[0055] Install energy consumption sensors on the liquid cooling pump and heat exchanger to obtain the operation status data of the equipment;
[0056] The first operation data at least includes heat generation data, coolant status data and equipment operation status data.
[0057] It should be noted that the purpose of the first preprocessing is to improve the accuracy and reliability of the data and provide a solid foundation for subsequent analysis and regulation. In the preprocessing stage, the system may perform a series of operations, such as data cleaning, denoising, outlier detection and processing, etc., to ensure the accuracy and integrity of the data.
[0058] In an optional embodiment, the first preprocessing may include data standardization processing, which converts data with different dimensions into the same scale to improve the efficiency and accuracy of model training.
[0059] In an optional embodiment, data augmentation can also be performed to improve the generalization ability of the model by generating more training samples and ensure that the model can perform well when dealing with various complex situations.
[0060] In the embodiment of the present application, the first preprocessing includes:
[0061] Establish a first reinforcement learning model, which is used to dynamically adjust the arrangement and working frequency of the sensors according to the first operation data;
[0062] Use denoising technology, moving average filtering, to remove interference signals from the first operation data processed by the first reinforcement learning model, and filter out abnormal data by the method based on mean and standard deviation.
[0063] In an optional embodiment, the reinforcement learning model can learn and optimize the configuration and working mode of the sensors to ensure that the most critical data is collected at critical moments. By dynamically adjusting the arrangement and working frequency of the sensors, the system can use resources more efficiently while reducing unnecessary energy consumption. In addition, denoising technology and abnormal data filtering further improve the accuracy of the data, providing a reliable basis for subsequent analysis and regulation.
[0064] In the embodiment of the present application, the first reinforcement learning model is to dynamically adjust the configuration and working mode of the sensors to adapt to different operating environments and load conditions. Through continuous learning and optimization, the model can identify which data is most critical for evaluating the operation status of the computing power cluster, so that under limited resources, these data are preferentially collected and processed. This intelligent data collection and processing strategy not only improves the data collection efficiency, but also ensures the accuracy and representativeness of the data, providing strong support for subsequent analysis and regulation.
[0065] Exemplarily, LoRa is used to connect each sensor, and the first operation data of the computing power cluster is transmitted to a position close to the data source. After receiving the first operation data of the computing power cluster, the edge device uses the denoising technology of moving average filtering to remove interference signals, and filters abnormal data by means of the method based on the mean and standard deviation;
[0066] Furthermore, the first reinforcement learning model of the present application defines the state space through the operation data, real-time temperature, pressure and flow state of the computing power node; defines the action space through the start and stop of the sensor, the adjustment of the sampling frequency, and the adjustment of the sensor position. The data acquisition module designs an appropriate reward function by balancing the two objectives of cooling effect and energy efficiency optimization, guiding the model to learn towards the optimal objective, and giving rewards according to the temperature and load conditions of the node. The specific formula of the reward function is based on:
[0067] R t = w1×Cooling Efficiency - w2×Energy Consumption
[0068] where, Cooling Efficiency represents the cooling efficiency, Energy Consumption represents the energy efficiency penalty, and w1 and w2 are weight coefficients;
[0069] It should be noted that the present application also adopts the ε-greedy strategy, selects a random action with a fixed probability ε, where ε represents the probability of "exploration". After accumulating a certain amount of experience, the model will select the best action by evaluating the Q value of the state-action pair. Among them, the Q value represents the expected return, and the higher the Q value, the better the model considers this action;
[0070] In an optional embodiment, the operating state can be continuously monitored, and the feedback information is transmitted to the reinforcement learning model for adjustment. At the same time, a feedback period is set. After each period, by comparing the actual cooling effect and energy consumption, the first reinforcement learning model is adjusted and the data acquisition strategy is adjusted.
[0071] It should be noted that the second operation data is a more accurate and reliable data set after preprocessing, and these data will be used in the subsequent model establishment and state evaluation stages.
[0072] It should be noted that collecting the first running data of the target computing power cluster and performing the first preprocessing on the first running data to obtain the second running data can ensure the accuracy and effectiveness of subsequent analysis and regulation. Through operations such as data cleaning, denoising, and outlier processing in the preprocessing stage, the system can eliminate noise and outliers in the original data, improving the accuracy and reliability of the data. This provides a solid foundation for subsequent model establishment and status evaluation, enabling the system to more accurately evaluate the running status of the computing power cluster and formulate more effective regulation strategies. At the same time, by dynamically adjusting the configuration and working mode of the sensors through the reinforcement learning model, the system can use resources more efficiently, reduce unnecessary energy consumption, and further improve the energy efficiency of the entire liquid cooling system.
[0073] S102, establish a first model, and evaluate the running status of the target computing power cluster based on the second running data in combination with the first model;
[0074] In the embodiment of the present application, the first model includes:
[0075] The first model is any model used to simulate the heat generation and conduction process of computing power nodes in the target computing power cluster;
[0076] The running status of the target computing power cluster at least includes cooling efficiency, heat generation amount of nodes, and liquid cooling resource utilization rate.
[0077] In an optional embodiment, the first model can be a machine learning-based prediction model. It uses the second running data as input and outputs an evaluation result of the running status of the target computing power cluster through complex algorithms and calculation processes. This evaluation result is a multi-dimensional index, including cooling efficiency, heat generation amount of nodes, and liquid cooling resource utilization rate, etc., which together reflect the current heat dissipation performance, energy consumption status, and liquid cooling resource utilization efficiency of the cluster.
[0078] It should be noted that in order to establish this first model, the system may need to collect a large amount of historical data and conduct in-depth analysis and processing. These data should cover various running scenarios and load conditions to ensure that the model can accurately simulate and predict the actual running status of the cluster. In the model training stage, the system will use these data to optimize the parameters and structure of the model to make it better adapt to the actual application scenario. Once the model is established, the system can input the second running data into the model in real time for running status evaluation. This evaluation result will be an important basis for formulating subsequent regulation strategies, helping the system achieve intelligent and dynamic management of cooling resources. Through continuous iteration and optimization, this first model will gradually become more accurate and efficient, providing strong support for the stable operation of the system.
[0079] In an alternative embodiment, the first model can also simulate the heat generation and conduction of computing power nodes during actual operation by converting current power into heat based on physical principles, and calibrate and optimize the model in combination with actual operation data to improve the accuracy and applicability of the model. This method combining physical principles and data-driven approach enables the first model to more accurately reflect the actual operation status of the target computing power cluster and provide a more reliable basis for subsequent regulation strategies.
[0080] In an alternative embodiment, the first model can also be built through historical data regression analysis to explore the potential relationships between the operation status of the computing power cluster, cooling efficiency, heat generation amount, and liquid cooling resource utilization rate. This method uses the principles of statistics to find the correlations between variables from historical data, thereby establishing a prediction model.
[0081] It should be noted that through continuous training and optimization, this model can gradually learn the complex characteristics of the cluster operation status and provide strong support for subsequent intelligent regulation. Regardless of which modeling method is adopted, the goal of the system is to establish an accurate and efficient first model to achieve a precise assessment of the operation status of the target computing power cluster.
[0082] In the embodiment of the present application, modeling is performed based on the conversion of current power into heat according to physical principles and historical data regression analysis. According to the workload and power consumption data of the computing power nodes, the heat generation of each node is estimated, specifically based on the formula: heat generation amount of the node = power consumption × conversion coefficient, where the specific conversion coefficient is obtained through experiments or data analysis;
[0083] Furthermore, the computational domain is divided into multiple small units, and the heat conduction equation is used to solve each unit to simulate how heat is transferred from high-temperature regions to low-temperature regions. The heat conduction equation: where κ is the thermal conductivity, is the temperature, is the source term. At the same time, as the status of the computing power cluster changes, the input data in the model is adjusted in real time, and the finite element method is used to dynamically solve the heat distribution and liquid cooling efficiency;
[0084] Furthermore, the cooling efficiency is defined by calculating the ratio of the heat actually carried away by the coolant to the heat that should be carried away theoretically, specifically based on the formula: cooling efficiency = (heat actually carried away / heat that should be carried away theoretically) × 100%. The liquid cooling resource utilization rate is defined by the relationship between the actual effect of the coolant during the cooling process and the preset target, specifically based on the formula: liquid cooling resource utilization rate = sum of actual effective cooling amounts / sum of available cooling resources;
[0085] Furthermore, based on the historical data of multiple variables, a relationship model between temperature and factors such as coolant flow rate, pressure, and flow velocity is established using regression analysis (linear regression). The cooling effect in different regions is judged, and according to the regression model, the cooling effect, temperature distribution, and liquid cooling resource allocation in each region are statistically analyzed to generate a detailed hotspot report.
[0086] It should be noted that establishing the first model and evaluating the operating state of the target computing power cluster based on the second operating data in combination with the first model can more accurately understand the real-time operating state of the target computing power cluster, including key indicators such as cooling efficiency, heat generation amount of nodes, and liquid cooling resource utilization rate. These indicators are important bases for formulating subsequent control strategies and can help the system achieve intelligent and refined allocation of cooling resources. Through the evaluation of the first model, the system can timely discover problems such as uneven heat dissipation and excessive energy consumption in the cluster and take corresponding control measures for optimization. This not only improves the heat dissipation efficiency and energy efficiency of the system but also extends the service life of the equipment and reduces the operation and maintenance costs. In addition, the establishment of the first model also provides the possibility for the continuous optimization and upgrade of the system, enabling the system to continuously adapt to new application scenarios and load conditions and maintain an efficient and stable operating state.
[0087] S103, perform regulation according to the operating state in combination with the first regulation strategy, and the first regulation strategy includes a first objective function and constraint conditions;
[0088] In an optional embodiment, the first regulation strategy is to optimize the allocation of cooling resources according to the operating state of the target computing power cluster obtained from the above evaluation. The first objective function aims to maximize the cooling efficiency while minimizing the energy consumption. This function may consider multiple variables, such as coolant flow rate, temperature distribution, node load, etc., to ensure that while meeting the heat dissipation requirements, unnecessary energy consumption is minimized as much as possible.
[0089] In an optional embodiment, the constraint conditions may include physical property limitations of the coolant, operating parameter limitations of the equipment, and safety specifications, etc. These constraint conditions ensure the feasibility and safety of the regulation strategy in practical applications.
[0090] In an optional embodiment, the first regulation strategy may include dynamically adjusting the flow rate and pressure of the coolant to optimize the cooling effect of different computing power nodes. For example, when the load of a certain node is high and the temperature rises, the system may increase the flow rate of the coolant flowing to that node to more effectively remove heat. On the contrary, for nodes with low load and relatively stable temperature, the system may appropriately reduce the flow rate of the coolant to avoid unnecessary energy consumption. This dynamic adjustment strategy not only improves the cooling efficiency but also ensures the reasonable allocation of energy consumption, making the entire liquid cooling system more efficient and energy-saving.
[0091] In the embodiments of the present application, the first regulation strategy includes:
[0092] The first regulation strategy is any regulation strategy that minimizes the total energy consumption of the liquid cooling system on the premise of meeting the cooling requirements of all nodes;
[0093] Minimizing the total energy consumption of the liquid cooling system includes at least the energy consumption of the liquid cooling pump and the penalty cost for node overheating.
[0094] In an alternative embodiment, the first regulation strategy may also involve dynamically adjusting the operating state of the computing power nodes according to the real-time load conditions. For example, when the load of a certain node is low, the system may reduce the heat generation by reducing its operating frequency or reducing its operating time, thereby reducing the cooling demand. This strategy not only helps to reduce energy consumption, but also extends the service life of the nodes to a certain extent. At the same time, in order to ensure the stability and reliability of the system, the first regulation strategy also needs to consider the overheat protection mechanism of the nodes. Once the temperature of a certain node exceeds the preset safety threshold, the system should immediately take measures, such as increasing the coolant flow rate or starting the standby cooling equipment, to prevent performance degradation or failure caused by node overheating. Through these refined regulation strategies, the intelligent allocation and regulation system of liquid cooling resources of the present invention can achieve precise control and optimization of the operating state of the computing power cluster, ensuring the efficient and stable operation of the system.
[0095] In the embodiments of the present application, according to the real-time operating state of the computing power cluster and the operating parameters of the cooling system, the allocation strategy and execution control of the liquid cooling resources are dynamically adjusted. The goal is to achieve efficient cooling while optimizing energy consumption and resource utilization.
[0096] First, using the real-time operating data of the computing power cluster, including: the operating load, surface temperature, coolant flow rate, pressure and flow velocity of the computing power nodes, the heat demand matrix and cooling priority of the nodes are generated through a thermodynamic model.
[0097] Specifically, according to the data such as the load and power consumption of the computing power nodes, the heat generation of each node is estimated using thermodynamic formulas. Among them, the heat generation formula of the node is: heat generation = power consumption × conversion coefficient. Combining the surface temperature of each node with the flow parameters (flow velocity, flow rate, pressure) of the coolant, according to the heat conduction theory and the working principle of the liquid cooling system, a heat conduction model is used to calculate the cooling demand of the node. By calculating the cooling demand of each node, a cooling priority is generated based on the temperature and load of the node, specifically according to the formula:
[0098] Priority = α × (T node - T avg ) + β × L node
[0099] Where, T nodeRepresents the node temperature, T avg Represents the average temperature of all nodes, L node Represents the node load, and α and β are adjustment weights.
[0100] Furthermore, a mathematical optimization model (such as linear programming) is used to formulate a resource allocation strategy to minimize the total energy consumption of the liquid cooling system and ensure that the cooling requirements are met.
[0101] Specifically, the objective function is to minimize the total energy consumption of the liquid cooling system while meeting the cooling requirements of all nodes. Specifically, according to the formula:
[0102] min∑P pumpi+∑ C overheati
[0103] Among them, P pumpi Represents the energy consumption of the liquid cooling pump, and C overheati Represents the penalty cost for node overheating; at the same time, the allocation strategy will give priority to meeting the needs of high-load and high-temperature nodes, and determine the priority according to the weight calculation formula. Specifically, according to the formula W i =a×(T i -T avg )+b×L i , where, T i Represents the node temperature, T avg Represents the average temperature, and L i Represents the real-time load of the node, and a and b are adjustment weights.
[0104] Furthermore, the liquid cooling pump is controlled by frequency conversion technology to dynamically adjust the coolant flow rate to meet the distribution requirements of each node. The pump speed control formula is:
[0105] P pump =k·Q 3
[0106] Among them, Q represents the coolant flow rate, and k represents the equipment constant, which determines the efficiency curve of the pump. At the same time, in the liquid cooling loop, the electric control valve adjusts the opening according to the cooling requirements of each node to ensure the precise and controllable distribution of liquid cooling resources. Specifically, according to the formula:
[0107]
[0108] Among them, θ represents the valve opening percentage, F desired Represents the current node demand flow rate, F max Represents the maximum coolant flow rate. Finally, according to the coolant flow rate, temperature and cooling requirements, the coolant outlet temperature of the heat exchanger is dynamically adjusted to ensure the maximization of cooling efficiency. And by adjusting the fan speed, the heat dissipation efficiency is optimized to reduce unnecessary energy loss.
[0109] It should be noted that the liquid cooling pipeline is fixed and cannot be dynamically optimized according to the change of heat distribution. When significant changes occur in the heat distribution (such as the transfer of high-load areas), the fixed pipeline design will lead to a decrease in cooling efficiency. Especially in the case of large load changes, hot spots are likely to occur and energy waste will be caused. Therefore, in this embodiment, through the pipeline dynamic reconstruction technology, the flow direction and distribution method of liquid cooling resources are adjusted according to the actual load and heat changes, optimizing the cooling efficiency, reducing energy consumption, and improving adaptability and stability.
[0110] Specifically, sensors are used to collect the temperature, coolant flow rate and pressure data of the computing power nodes in real time, and these real-time data are used to construct a thermodynamic data model to accurately monitor the heat distribution and flow state, and identify the change trends of temperature, flow rate and pressure.
[0111] Furthermore, according to the real-time temperature, pressure and flow rate data, the heat conduction process of the coolant is calculated through CFD and FEA models to evaluate the temperature distribution of each node. And the temperature change of the node is detected in real time. When the temperature of a certain node exceeds the preset threshold, it is marked as a hot spot area. If the heat distribution changes (such as load transfer or a sharp increase in temperature in some areas), the module will update the thermodynamic model in time and recalculate the cooling demand.
[0112] Furthermore, based on the data such as the temperature, flow rate and pressure of the nodes, the cooling demand of each node is calculated through the thermodynamic model. And according to the information such as the temperature, load and cooling demand of the nodes, the weighted average method or other sorting algorithms are used to assign cooling priorities to each node. At the same time, the cooling demand and priority sorting matrix of the nodes provide a basis for adjusting the coolant flow direction and resource allocation.
[0113] Furthermore, based on the real-time thermodynamic data and load changes, an intelligent control algorithm is used to dynamically adjust the flow direction of the liquid cooling pipeline.
[0114] In the embodiment of the present application, the first regulation strategy also includes assigning cooling priorities to each node.
[0115] Specifically, according to the priority of the cooling demand, the coolant flow rate of each node is calculated through an optimization algorithm, and the coolant flow direction is determined. A priority-based scheduling algorithm (such as a greedy algorithm, reinforcement learning, etc.) is used to dynamically adjust the pipeline flow direction according to the cooling demand. And according to the calculated cooling demand and flow rate, the liquid cooling flow is controlled to the specified node by adjusting the pump speed and valve opening to ensure that high-temperature or high-load nodes are cooled first. At the same time, according to the fuzzy control rules, the control range of the coolant flow rate is determined according to the information such as the temperature and load of the nodes. Improve the cooling efficiency and ensure that the hot spot area is cooled first.
[0116] It should be noted that the regulation is carried out according to the operating state in combination with the first regulation strategy. The first regulation strategy includes the first objective function and constraint conditions, which can ensure the efficient and safe operation of the liquid cooling resource intelligent allocation and regulation system in practical applications. First of all, by comprehensively considering the cooling efficiency and energy consumption, the first objective function ensures that while meeting the heat dissipation requirements of the computing power cluster, the energy consumption of the system is reduced as much as possible, which is of great significance for improving the overall energy efficiency and reducing the operation cost. Secondly, the setting of constraint conditions, such as the physical property limitations of the coolant, the operating parameter limitations of the equipment, and safety specifications, etc., ensures the feasibility and safety of the regulation strategy during the actual implementation process, and avoids equipment damage or safety accidents caused by over-cooling or improper operation. In addition, by dynamically adjusting the flow rate and pressure of the coolant, and dynamically adjusting the operating state of the computing power nodes according to the real-time load conditions, the system can achieve precise control and optimization of the operating state of the computing power cluster, ensuring the efficient and stable operation of the system. These refined regulation strategies not only improve the heat dissipation efficiency and energy efficiency of the system, but also extend the service life of the equipment, reduce the operation and maintenance cost, and provide a reliable liquid cooling solution for computing power-intensive application scenarios such as data centers.
[0117] S104, solve the first objective function to obtain the optimal regulation plan.
[0118] In the embodiment of the present application, solving the first objective function to obtain the optimal regulation plan includes:
[0119] Dynamically adjust the coolant outlet temperature of the heat exchanger according to the solution result, and optimize the heat dissipation efficiency by adjusting the fan speed;
[0120] Adopt an adaptive algorithm to continuously adjust the coolant flow rate and pipeline flow direction according to the feedback data, and adjust the first regulation strategy through the adaptive algorithm.
[0121] In an optional embodiment, through the preset first regulation strategy, intelligently determine the cooling requirements of each area or node, dynamically reconstruct the pipeline flow direction, avoid over-cooling or under-cooling in the high-heat area and low-efficiency area, optimize the cooling resource allocation, and combine the dynamic reconstruction and the overall design of the liquid cooling system. Through the adjustable pipeline and intelligent valve control design, it can be adaptively adjusted according to the real-time change of the heat distribution.
[0122] Specifically, formulate the coolant distribution plan for each node through mathematical optimization (such as linear programming, integer programming, etc.). Perform dynamic scheduling according to the demand to ensure the on-demand allocation of cooling resources. And according to the scheduling result, adopt an adjustable design, combined with intelligent regulating valves and variable frequency pumps, to dynamically adjust the pipeline flow direction and flow rate to ensure that the coolant can flow to the high-load and high-temperature areas in real time and accurately.
[0123] By continuously monitoring the cooling effect and coolant flow rate, an adaptive algorithm is used to dynamically adjust the coolant flow rate and pipeline flow direction according to the feedback data. Based on the feedback data, the control strategy is adjusted through the adaptive algorithm to ensure that the cooling effect reaches the optimal level.
[0124] In summary, the present invention proposes an intelligent allocation and regulation method for liquid cooling resources, which collects the first operation data of the target computing power cluster and performs a first preprocessing on the first operation data to obtain the second operation data; establishes a first model, and evaluates the operation state of the target computing power cluster based on the second operation data combined with the first model; performs regulation according to the operation state combined with the first regulation strategy, and the first regulation strategy includes a first objective function and constraint conditions; solves the first objective function to obtain an optimal regulation plan. It can realize the intelligent and dynamic management of the cooling resources of the computing power cluster and ensure efficient and safe operation.
[0125] By collecting and processing the operation data of the target computing power cluster in real time, the system can accurately evaluate the operation state of the cluster and perform intelligent regulation according to the preset regulation strategy. Among them, the data collection and processing module, as the core part of the system, is responsible for collecting and processing various sensor data from the computing power cluster, providing reliable data support for subsequent analysis and regulation. By establishing an accurate mathematical model, the system can simulate the heat generation and conduction process of the computing power nodes, thereby accurately judging key indicators such as the cooling efficiency of the cluster, the heat generation amount of the nodes, and the utilization rate of liquid cooling resources. In terms of the regulation strategy, the system aims to minimize the total energy consumption of the liquid cooling system on the premise of meeting the cooling requirements of all nodes, which includes the energy consumption of the liquid cooling pump and the penalty cost caused by overheating of the nodes. To achieve this goal, the system adopts advanced algorithms and technical means, such as reinforcement learning models, denoising techniques, adaptive algorithms, etc., to ensure the accuracy and effectiveness of the regulation plan. By continuously optimizing parameters such as the flow rate, flow volume, and pressure of the coolant, and dynamically adjusting measures such as the coolant outlet temperature of the heat exchanger and the fan speed, the system can significantly improve the cooling efficiency of the computing power cluster, reduce energy consumption, and ensure the stable operation of the system.
[0126] Embodiment 2
[0127] In a preferred embodiment, the present application can also be designed to provide functions of real-time status monitoring, abnormal alarm, and emergency handling of the computing power cluster to ensure the safety of the computing power cluster.
[0128] First, use a chart library (such as D3.js, Highcharts, Grafana, etc.) to implement the display of data through a Web or desktop visualization interface (based on the chart library and front-end frameworks such as React, Vue). Use forms such as heatmaps, line charts, and bar charts to display key indicators such as temperature distribution, flow status, and pressure changes, and use statistical analysis and machine learning techniques (such as regression analysis, time series analysis) to perform trend analysis on the collected historical data. Based on the historical data, predict possible future temperature changes, flow fluctuations, etc., to help identify long-term trends and potential abnormal patterns.
[0129] Furthermore, monitor the temperature, flow, and pressure data of each node in real time, and detect abnormal changes in the data through threshold judgment (such as when indicators such as temperature, pressure, and flow exceed the set thresholds) or machine learning-based anomaly detection algorithms (such as Isolation Forest, SVM, etc.). Set multiple alarm thresholds (such as issuing a warning when the temperature is 50°C and starting an emergency cooling strategy when it is 60°C), and each threshold corresponds to a different alarm level (such as warning, alert, emergency alarm).
[0130] Furthermore, after confirming the anomaly, automatically start the cooling system or regulating valve for emergency cooling, and the control system executes the cooling measures through PLC or SCADA. Integrate with the mail server and SMS gateway (such as Twilio) through the API, or transmit the alarm information to the administrator and operation and maintenance personnel through the voice broadcast system to ensure that they can handle other potential problems in a timely manner. At the same time, record the occurrence time, type, and processing process of all abnormal events to provide a basis for later analysis and optimization.
[0131] Embodiment 3
[0132] This embodiment also provides an intelligent allocation and regulation system for liquid cooling resources, which is characterized by including:
[0133] A data collection and processing module for collecting the first operation data of the target computing power cluster and performing a first preprocessing on the first operation data to obtain second operation data;
[0134] A model establishment module for establishing a first model and evaluating the operation status of the target computing power cluster based on the second operation data in combination with the first model;
[0135] A regulation module for performing regulation according to the operation status in combination with a first regulation strategy, and the first regulation strategy includes a first objective function and constraint conditions;
[0136] A solution module for solving the first objective function to obtain an optimal regulation plan.
[0137] Each of the above unit modules can be embedded in or independent of a processor in a computer device in hardware form, or stored in a memory in a computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.
[0138] This embodiment also provides a computer device, which can be a terminal, and its internal structure diagram can be as Figure 2 shown. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it realizes a method for intelligent allocation and regulation of liquid cooling resources. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0139] This embodiment also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0140] Collect the first operation data of the target computing power cluster, and perform a first preprocessing on the first operation data to obtain second operation data;
[0141] Establish a first model, and evaluate the operation state of the target computing power cluster based on the second operation data in combination with the first model;
[0142] Perform regulation according to the operation state in combination with a first regulation strategy, and the first regulation strategy includes a first objective function and constraint conditions;
[0143] Solve the first objective function to obtain an optimal regulation plan.
[0144] Embodiment 4
[0145] In an alternative embodiment, a detailed system structure based on the above method and system can be designed, specifically including a data acquisition module, a liquid cooling state evaluation module, a liquid cooling resource allocation and control module, a monitoring and alarm module, and a data storage module, and the modules are connected by signals.
[0146] The data storage module is used to store all data during the processing process.
[0147] The data acquisition module is responsible for real-time acquisition of the operation data of the computing power cluster, including heat generation data, coolant status data, and equipment operation status data, providing basic data support for subsequent modules.
[0148] First, the data acquisition module deploys temperature sensors at computing power nodes to collect the surface temperature of the nodes and obtain heat generation data; installs pressure sensors, flow velocity sensors, and flow meters in the liquid cooling pipeline to monitor the flow status of the coolant and obtain coolant status data; installs energy consumption sensors on the liquid cooling pump and heat exchanger to obtain equipment operation status data.
[0149] Furthermore, the data acquisition module uses Internet of Things technologies (such as LoRa or ZigBee) with the characteristics of low power consumption and long transmission distance, which are suitable for large-scale distributed computing power cluster environments, to connect each sensor, transmit the operation data of the computing power cluster to the edge device located close to the data source, which has strong computing and storage capabilities and can realize preliminary data processing and analysis.
[0150] Furthermore, after receiving the operation data of the computing power cluster, the edge device preprocesses the operation data of the computing power cluster, uses denoising techniques (such as moving average filtering) to remove interference signals, and filters out abnormal data through methods based on mean and standard deviation to improve data quality.
[0151] It should be noted that traditional sensor configurations are usually fixed and cannot be dynamically adjusted, which may lead to monitoring blind spots or resource waste in high-temperature areas. Therefore, in this embodiment, a sensor configuration optimization mechanism based on a self-learning algorithm is designed to dynamically adjust the layout and working frequency of sensors according to historical acquisition data and environmental conditions, thereby optimizing data acquisition efficiency and reducing energy consumption.
[0152] First, the data acquisition module uses a reinforcement learning algorithm (such as DQN) to construct a dynamic configuration model.
[0153] Specifically, the data acquisition module defines the state space through the operation data, real-time temperature, pressure, and flow status of the computing power nodes; for example: Node 1: Temperature 70°C, load 85%, coolant flow rate 6 L / min, pressure 1.2 MPa, Node 2: Temperature 45°C, load 30%, coolant flow rate 4 L / min, pressure 0.9 MPa. The action space is defined by sensor start / stop, sampling frequency adjustment, and sensor position adjustment, for example: enable more sensors in high-temperature areas and increase the sampling frequency, and turn off some sensors in low-load areas and reduce the sampling frequency.
[0154] Furthermore, the data acquisition module designs an appropriate reward function by balancing two objectives: cooling effect and energy efficiency optimization, guiding the model to learn towards the optimal objective. Specifically, rewards are given based on the temperature and load conditions of the nodes. For example, a positive reward is given when the node temperature remains within a reasonable range, and a negative reward is given when the temperature is too high. When the sensor activation frequency in a certain area is too high or too low, or the sampling frequency is not optimized, the data acquisition module will receive a negative penalty. The specific formula of the reward function is based on:
[0155] R t = w1 × Cooling Efficiency - w2 × Energy Consumption
[0156] where Cooling Efficiency represents the cooling efficiency, Energy Consumption represents the energy efficiency penalty, and w1 and w2 are weight coefficients, reflecting the relative importance of cooling efficiency and energy consumption respectively. For example: If Node 1 (temperature 70°C) reduces its temperature to 60°C by increasing sensors and sampling frequency, a certain positive reward is given, and the cooling effect is improved. If Node 2 (temperature 45°C) does not require high-frequency sampling but the sensors still maintain high-frequency sampling, it will be penalized.
[0157] Furthermore, in the initial stage, since the model has insufficient understanding of the environment, randomly selecting actions helps to better understand the relationship between the model state and actions. Relying solely on exploitation too early may lead to local optimal solutions. As the training progresses, the model's understanding of the environment gradually increases. Using existing experience to select optimal actions can improve decision-making efficiency and avoid wasting resources. Therefore, the data acquisition module adopts the ε-greedy strategy, randomly selecting actions with a fixed probability ε, where ε represents the probability of "exploration". After the model has accumulated a certain amount of experience, it will select the best action by evaluating the Q-value (expected return) of the state-action pair. The higher the Q-value, the better the model considers this action. For example: Every time a decision needs to be made, the model first generates a random number r between [0, 1]. If r < ε, then explore: select a random action. If r ≥ ε, then exploit: select the action with the maximum Q-value. At the same time, the data acquisition module uses Q-learning technology to continuously update the Q-value, evaluate the expected return of taking a certain action in each state, and gradually converge to the optimal strategy.
[0158] Furthermore, the data acquisition module continuously monitors the operating status (such as node temperature, coolant flow, etc.), and transmits this feedback information to the reinforcement learning model for adjustment. At the same time, after each cycle, the data acquisition module adjusts the strategy and optimizes the model by comparing the actual cooling effect and energy consumption.
[0159] The liquid cooling status evaluation module evaluates the operating status of the computing power cluster in real time based on the operation data of the computing power cluster obtained by the data acquisition module, including cooling efficiency, heat distribution, and liquid cooling resource utilization.
[0160] First, the liquid cooling status evaluation module establishes a dynamic thermal model to simulate the heat generation of computing nodes. Specifically, the liquid cooling status evaluation module performs modeling based on physical principles (such as the conversion of electrical power into heat) or historical data regression analysis. According to data such as the workload and power consumption of computing nodes, it estimates the heat generation of each node. Specifically, it is based on the formula: heat generation of the node = power consumption * conversion coefficient, where the specific conversion coefficient can be obtained through experiments or data analysis.
[0161] Furthermore, the liquid cooling status evaluation module uses the finite element analysis (FEA) method to establish the heat conduction relationship between the coolant and the computing nodes, considering factors such as the flow rate, flow volume, and temperature of the coolant.
[0162] Specifically, the liquid cooling status evaluation module divides the computational domain (including computing nodes and the liquid cooling system) into multiple small units (finite element meshes). Each unit will simulate certain temperature, pressure, and flow rate conditions. And it uses the heat conduction equation to solve each unit to simulate how heat is transferred from high-temperature regions to low-temperature regions. Heat conduction equation: where κ is the thermal conductivity, is the temperature, and is the source term (heat generation). At the same time, as the status of the computing power cluster changes (such as load changes), the input data (temperature, flow rate, etc.) in the model is adjusted in real time, and the finite element method is used to dynamically solve the heat distribution and liquid cooling efficiency.
[0163] Furthermore, the liquid cooling status evaluation module calculates the cooling efficiency and the utilization of liquid cooling resources through the solution results of the dynamic thermal model. Among them, the cooling efficiency is defined by calculating the ratio of the heat actually carried away by the coolant to the heat that should be carried away theoretically. Specifically, it is based on the formula: cooling efficiency = (heat actually carried away / heat that should be carried away theoretically) × 100%, where the heat actually carried away is the change in the heat of the coolant obtained through simulation. The heat that should be carried away theoretically is calculated based on the power consumption and heat generation of the computing nodes. The liquid cooling resource utilization rate is defined by the relationship between the actual effect of the coolant during the cooling process and the preset target. Specifically, it is based on the formula: liquid cooling resource utilization rate = sum of actual effective cooling amounts / sum of available cooling resources, where the actual effective cooling amount is calculated through the cooling efficiency. The available cooling resources are estimated through parameters such as the coolant flow rate and flow velocity.
[0164] Further, the liquid cooling status evaluation module uses historical data of multiple variables (such as temperature, pressure, flow rate, etc.) and employs regression analysis (such as linear regression, polynomial regression, etc.) to establish a relationship model between temperature and factors such as coolant flow rate, pressure, and flow velocity, and judges the cooling effects in different regions. And based on the regression model and real-time data, it statistics the cooling effects, temperature distribution, and liquid cooling resource allocation in each region, and generates a detailed hot spot report.
[0165] The liquid cooling resource allocation and control module dynamically adjusts the allocation strategy and execution control of liquid cooling resources according to the real-time operating status of the computing power cluster and the operating parameters of the cooling system. Its goal is to achieve efficient cooling while optimizing energy consumption and resource utilization.
[0166] First, the liquid cooling resource allocation and control module utilizes the real-time operating data of the computing power cluster, including: the operating load, surface temperature of the computing power nodes, the flow rate, pressure, and flow velocity of the coolant, and generates a heat demand matrix and cooling priority of the nodes through a thermodynamic model.
[0167] Specifically, the liquid cooling resource allocation and control module estimates the heat generation of each node according to data such as the load and power consumption of the computing power nodes. Among them, the heat generation formula of the node is: heat generation = power consumption × conversion coefficient. The liquid cooling resource allocation and control module combines the surface temperature of each node with the flow parameters (flow rate, flow, pressure) of the coolant, and according to the heat conduction theory and the working principle of the liquid cooling system, uses a heat conduction model to calculate the cooling demand of the node. By calculating the cooling demand of each node, a cooling priority is generated based on the temperature and load of the node, specifically according to the formula:
[0168] Priority = α × (T node - T avg ) + β × L node Where, T node represents the node temperature, T avg represents the average temperature of all nodes, L node represents the node load, and α and β are adjustment weights.
[0169] Further, the liquid cooling resource allocation and control module uses a mathematical optimization model (such as linear programming) to formulate a resource allocation strategy to minimize the total energy consumption of the liquid cooling system and ensure that the cooling demand is met.
[0170] Specifically, the objective function is to minimize the total energy consumption of the liquid cooling system on the premise of meeting the cooling demand of all nodes, specifically according to the formula:
[0171] min∑P pumpi + ∑C overheati
[0172] Where, P pumpiIndicates the energy consumption of the liquid cooling pump, C overheati Indicates the penalty cost for node overheating; at the same time, the allocation strategy will prioritize meeting the needs of high-load and high-temperature nodes, and determine the priority according to the weight calculation formula, specifically based on the formula W i = a×(T i -T avg ) + b×L i , where T i Indicates the node temperature, T avg Indicates the average temperature, L i Indicates the real-time load of the node, and a and b are adjustment weights.
[0173] Furthermore, the liquid cooling resource allocation and control module controls the liquid cooling pump through frequency conversion technology, dynamically adjusts the coolant flow rate, and meets the allocation requirements of each node. The pump speed control formula is:
[0174] P pump = k·Q 3
[0175] where Q represents the coolant flow rate, and k represents the equipment constant, which determines the efficiency curve of the pump. At the same time, in the liquid cooling loop, the electric control valve adjusts the opening degree according to the cooling requirements of each node to ensure the accurate and controllable allocation of liquid cooling resources, specifically based on the formula:
[0176]
[0177] where θ represents the valve opening percentage, F desired represents the current node demand flow rate, F max represents the maximum coolant flow rate. Finally, the liquid cooling resource allocation and control module dynamically adjusts the coolant outlet temperature of the heat exchanger according to the coolant flow rate, temperature, and cooling requirements to ensure the maximization of cooling efficiency. And by adjusting the fan speed, it optimizes the heat dissipation efficiency and reduces unnecessary energy losses.
[0178] It should be noted that the liquid cooling pipeline is fixed and cannot be dynamically optimized for changes in heat distribution. When significant changes occur in heat distribution (such as the transfer of high-load areas), the fixed pipeline design will lead to a decrease in cooling efficiency. Especially in the case of large load changes, hot spots are likely to occur and energy waste will result. Therefore, in this embodiment, through the pipeline dynamic reconstruction technology, the flow direction and allocation method of liquid cooling resources are adjusted according to the actual load and heat changes, optimizing the cooling efficiency, reducing energy consumption, and improving adaptability and stability.
[0179] Specifically, the liquid cooling resource allocation and control module uses sensors to collect real-time data on the temperature, coolant flow rate, and pressure of the computing power nodes, and uses these real-time data to construct a thermodynamic data model to accurately monitor the heat distribution and flow state, and identify the change trends of temperature, flow rate, and pressure.
[0180] Furthermore, based on the real-time temperature, pressure, and flow rate data, the liquid cooling resource allocation and control module calculates the heat conduction process of the coolant through CFD and FEA models, and evaluates the temperature distribution of each node. It also continuously monitors the temperature changes of the nodes. When the temperature of a certain node exceeds the preset threshold, it is marked as a hot spot area. If the heat distribution changes (such as load transfer or a sharp increase in temperature in some areas), the liquid cooling resource allocation and control module will update the thermodynamic model in a timely manner and recalculate the cooling requirements.
[0181] Furthermore, based on data such as the temperature, flow rate, and pressure of the nodes, the liquid cooling resource allocation and control module calculates the cooling requirements of each node through a thermodynamic model. Then, according to information such as the temperature, load, and cooling requirements of the nodes, it uses the weighted average method or other sorting algorithms to assign a cooling priority to each node. Meanwhile, the cooling requirements and priority sorting matrix of the nodes provide a basis for adjusting the coolant flow direction and resource allocation.
[0182] Furthermore, based on real-time thermodynamic data and load changes, the liquid cooling resource allocation and control module dynamically adjusts the flow direction of the liquid cooling pipeline using intelligent control algorithms.
[0183] Specifically, the liquid cooling resource allocation and control module calculates the coolant flow rate of each node through an optimization algorithm according to the priority of the cooling requirements, determines the coolant flow direction, and uses a priority-based scheduling algorithm (such as the greedy algorithm, reinforcement learning, etc.) to dynamically adjust the pipeline flow direction according to the cooling requirements. According to the calculated cooling requirements and flow rates, it controls the liquid cooling flow to the specified nodes by adjusting the pump speed and valve opening, ensuring that high-temperature or high-load nodes are cooled first. At the same time, according to the fuzzy control rules, it determines the control range of the coolant flow rate based on information such as the temperature and load of the nodes, improving the cooling efficiency and ensuring that the hot spot areas are cooled first.
[0184] Furthermore, through a preset optimized scheduling strategy, the liquid cooling resource allocation and control module intelligently determines the cooling requirements of each area or node, dynamically reconstructs the pipeline flow direction, avoids over-cooling or under-cooling of high-heat areas and inefficient areas, optimizes the cooling resource allocation, and combines the dynamic reconstruction with the overall design of the liquid cooling system. Through an adjustable pipeline and intelligent valve control design, it can adaptively adjust according to the real-time changes in the heat distribution.
[0185] Specifically, the liquid cooling resource allocation and control module formulates the coolant allocation plan for each node through mathematical optimization (such as linear programming, integer programming, etc.). It performs dynamic scheduling according to the requirements to ensure that the cooling resources are allocated as needed. According to the scheduling results, it adopts an adjustable design, combines intelligent regulating valves and variable-frequency pumps, and dynamically adjusts the pipeline flow direction and flow rate to ensure that the coolant can accurately flow to high-load and high-temperature areas in real time.
[0186] The liquid cooling resource allocation and control module continuously monitors the cooling effect and coolant flow rate, and uses an adaptive algorithm to dynamically adjust the coolant flow rate and pipeline flow direction according to the feedback data. Based on the feedback data, it adjusts the control strategy through the adaptive algorithm to ensure that the cooling effect reaches the optimal level.
[0187] The monitoring and alarm module provides real-time status monitoring, abnormal alarm, and emergency handling functions for the computing power cluster to ensure the safety of the computing power cluster.
[0188] First of all, the monitoring and alarm module uses chart libraries (such as D3.js, Highcharts, Grafana, etc.) to display data through a Web or desktop visualization interface (based on chart libraries and front-end frameworks such as React, Vue). It uses forms such as heat maps, line charts, and bar charts to display key indicators such as temperature distribution, flow status, and pressure changes, and uses statistical analysis and machine learning techniques (such as regression analysis, time series analysis) to perform trend analysis on the collected historical data. Based on the historical data, it predicts possible future temperature changes, flow fluctuations, etc., to help identify long-term trends and potential abnormal patterns.
[0189] Furthermore, the monitoring and alarm module monitors the temperature, flow rate, and pressure data of each node in real time, and detects abnormal changes in the data through threshold judgment (such as when indicators such as temperature, pressure, and flow rate exceed the set thresholds) or machine learning-based anomaly detection algorithms (such as Isolation Forest, SVM, etc.). Multiple alarm thresholds are set (such as issuing a warning when the temperature is 50°C and starting an emergency cooling strategy when it is 60°C), and each threshold corresponds to a different alarm level (such as warning, alarm, emergency alarm).
[0190] Furthermore, after confirming an abnormality, the monitoring and alarm module automatically starts the cooling system or regulating valve for emergency cooling, and the control system executes the cooling measures through PLC or SCADA. It is integrated with the mail server and SMS gateway (such as Twilio) through the API, or transmits the alarm information to the administrator and operation and maintenance personnel through the voice broadcast system to ensure that they can handle other potential problems in a timely manner. At the same time, the monitoring and alarm module records the occurrence time, type, and handling process of all abnormal events, providing a basis for later analysis and optimization.
[0191] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not restrictive. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
[0192] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code. The solutions in the embodiments of the present application can be implemented in various computer languages. For example, object-oriented programming languages such as Java and interpreted scripting languages such as JavaScript can be used.
[0193] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0194] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0195] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0196] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the present application.
[0197] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these modifications and variations.
Claims
1. An intelligent allocation and regulation method for liquid cooling resources, characterized in that, Including: Collecting first operation data of a target computing power cluster, and performing first preprocessing on the first operation data to obtain second operation data; Establishing a first model, and evaluating the operation state of the target computing power cluster based on the second operation data in combination with the first model; Performing regulation according to the operation state in combination with a first regulation strategy, where the first regulation strategy includes a first objective function and constraint conditions; Solving the first objective function to obtain an optimal regulation plan.
2. The intelligent allocation and regulation method of liquid cooling resources according to claim 1, wherein The first model includes: The first model is any model used to simulate the heat generation and conduction process of computing power nodes in a target computing power cluster; The operation state of the target computing power cluster at least includes cooling efficiency, heat generation amount of nodes, and utilization rate of liquid cooling resources.
3. The liquid cooling resource intelligent allocation and regulation method according to claim 2, characterized in that, The first regulation strategy includes: The first regulation strategy is any regulation strategy that minimizes the total energy consumption of the liquid cooling system on the premise of meeting the cooling requirements of all nodes; The minimization of the total energy consumption of the liquid cooling system at least includes the energy consumption of the liquid cooling pump and the penalty cost for node overheating.
4. The liquid cooling resource intelligent allocation and regulation method according to claim 3, characterized in that, The collecting of the first operation data of the target computing power cluster includes: Deploying temperature sensors at the computing power nodes of the target computing power set, collecting the surface temperature of the nodes, and obtaining heat generation data; Installing pressure sensors, flow velocity sensors, and flow meters in the liquid cooling pipeline to monitor the flow state of the coolant and obtain coolant state data; Installing energy consumption sensors on the liquid cooling pump and the heat exchanger to obtain equipment operation state data; The first operation data at least includes heat generation data, coolant state data, and equipment operation state data.
5. The intelligent allocation and regulation method of liquid cooling resources according to claim 4, characterized in that, The first preprocessing includes: Establishing a first reinforcement learning model, which is used to dynamically adjust the arrangement and working frequency of sensors according to the first operation data; Using denoising technology, sliding average filtering to remove interference signals from the first operation data processed by the first reinforcement learning model, and filtering abnormal data by a method based on mean and standard deviation.
6. The intelligent allocation and regulation method of liquid cooling resources according to claim 5, characterized in that The first regulation strategy further includes assigning cooling priorities to each node.
7. The intelligent allocation and regulation method of liquid cooling resources according to claim 6, wherein The solving of the first objective function to obtain an optimal regulation plan includes: Dynamically adjusting the coolant outlet temperature of the heat exchanger according to the solution result, and optimizing the heat dissipation efficiency by adjusting the fan speed; Adopting an adaptive algorithm to continuously adjust the coolant flow rate and pipeline flow direction according to the feedback data for dynamic adjustment results, and adjusting the first regulation strategy by the adaptive algorithm.
8. An intelligent allocation and regulation system for liquid cooling resources, characterized in that, Including: A data collection and processing module, which is used to collect the first operation data of the target computing power cluster, and perform first preprocessing on the first operation data to obtain second operation data; A model establishment module, which is used to establish a first model, and evaluate the operation state of the target computing power cluster based on the second operation data in combination with the first model; A regulation module, which is used to perform regulation according to the operation state in combination with a first regulation strategy, where the first regulation strategy includes a first objective function and constraint conditions; A solving module, which is used to solve the first objective function to obtain an optimal regulation plan.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.
Citation Information
Cited By
Immersed liquid cooling edge AI computing power server
CN120428837A
Dynamic flow control method for two-phase cold plate liquid cooling system based on multi-mode perception
CN120909404A
Dynamic flow control method for two-phase cold plate liquid cooling system based on multi-modal perception
CN120909404B
Anti-blocking water circulation radiator
CN120935987A
Self-adaptive operation and maintenance regulation and control method and device for liquid cooling system of intelligent calculation data center
CN121057181A