Liquid cooling server safety management system and method

Through multimodal sensor network and advanced data processing and learning algorithms, the thermal runaway problem of liquid-cooled servers is solved, precise control and data security management of the liquid-cooled system are realized, and the performance and reliability of the system are improved.

CN120152228AActive Publication Date: 2025-06-13百信信息技术有限公司 +1

Patent Information

Application Number
CN202510294658.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-13
Estimated Expiration
2045-03-13

AI Technical Summary

Technical Problem

Liquid-cooled servers face thermal runaway in high-performance computing scenarios. The existing sensor deployment method is difficult to comprehensively and accurately monitor the temperature distribution of chip surface and PCB board hotspot areas, resulting in the inability to detect local overheating risks in time, affecting the performance and life of the server.

Method used

The multimodal sensor network is used to collect thermodynamic parameters in real time, including infrared thermal imaging sensors, ultrasonic flowmeters and piezoelectric vibration sensors. A three-dimensional thermal field digital twin model is constructed through the extended Kalman filtering algorithm and graph convolution network. A distributed cooling strategy is designed in combination with federated learning and reinforcement learning, optimize the distribution of coolant flow, and ensure the traceability and tamper resistance of control instructions through blockchain.

Benefits of technology

It realizes accurate thermodynamic parameter acquisition and data processing of liquid cooling systems, improves data accuracy and availability, improves the accuracy and real-timeness of cooling strategies, extends the service life of the server, and enhances the security and reliability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120152228A_ABST
    Figure CN120152228A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of liquid cooling servers, and discloses a liquid cooling server safety management system and method, and the method employs a multi-mode sensor network to synchronously collect the thermodynamic parameters of a liquid cooling system at 100 Hz, and constructs a three-dimensional thermal field digital twinborn model after the processing of an extended Kalman filtering algorithm. A distributed cooling strategy is designed based on a federated learning framework, and each node locally trains a thermal dynamic prediction model and is optimized by a central aggregator. A dynamic control instruction is generated by using a near-end strategy optimization algorithm, and cooling liquid flow distribution is optimized in combination with a quantum derivative simulated annealing algorithm. And designing a dual-threshold phase change control mechanism, establishing a block chain log to ensure traceability and tamper resistance of the instruction, and realizing closed-loop feedback control through a CAN bus. The system and the method can accurately monitor and intelligently control the liquid cooling system, improve the heat dissipation efficiency, reduce the power consumption, and guarantee the data safety and the system stability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of liquid-cooled servers, and specifically to a liquid-cooled server security management system and method. Background Art

[0002] With the rapid development of information technology, the scale of data centers has been continuously expanding, and the computing power and power density of servers have also been continuously increasing. The traditional air-cooled heat dissipation method has been difficult to meet the heat dissipation requirements of high-performance servers, and liquid-cooling technology has gradually become the mainstream choice due to its high-efficient heat dissipation performance. However, in the actual application of liquid-cooled servers, many challenges still exist.

[0003] From the perspective of heat dissipation management, the temperature distribution of each component in the liquid-cooling system is complex and dynamically changing. Server chips (such as CPUs, GPUs, NPUs) generate a large amount of heat during operation, and the heat generation conditions of different chips and different regions of the same chip vary greatly. How to accurately obtain these complex thermodynamic parameters and effectively manage them has become a key issue. The existing sensor deployment methods are difficult to comprehensively and accurately monitor the temperature distribution of the chip surface and the hot spot areas of the PCB board, resulting in the inability to detect local overheating hazards in time, which affects the performance and lifespan of the server. For example, in some high-performance computing scenarios, local overheating of the chip may cause thermal runaway and lead to system failures.

[0004] In terms of data processing and analysis, the liquid-cooling system involves multi-source heterogeneous data, including temperature, flow rate, vibration, etc. These data have differences in time and space, and traditional data processing methods are difficult to effectively fuse and analyze them, and cannot provide an accurate basis for formulating heat dissipation strategies. Moreover, the accuracy of existing thermal dynamic prediction models is limited, and they cannot accurately predict the thermal change trend of the system, making the cooling strategy often lag behind the actual demand and reducing the heat dissipation efficiency.

[0005] Data security and privacy protection in a distributed architecture cannot be ignored either. In a liquid-cooling system with multiple server nodes, the data of each node contains sensitive information. When performing model training and strategy optimization, how to achieve data sharing and collaborative processing without leaking data privacy is an urgent problem to be solved. At the same time, the update of the cooling strategy and the execution of control instructions lack effective traceability and anti-tampering capabilities. Once a problem occurs, it is difficult to quickly locate and solve.

[0006] In addition, the rationality of coolant flow distribution directly affects the heat dissipation effect and energy consumption. The current flow distribution methods often rely on experience or simple algorithms and cannot be dynamically optimized according to the real-time thermal state of the server, resulting in insufficient heat dissipation in some areas and energy waste in some areas. Moreover, the phase change control mechanism in the phase change cooling process is not intelligent enough and cannot adjust the coolant state in a timely and accurate manner according to temperature changes, affecting the overall performance of the system. Summary of the Invention

[0007] The purpose of the present invention is to provide a liquid-cooled server security management system and method to solve the problems raised in the above background technology.

[0008] To achieve the above purpose, the present invention provides the following technical solution: A liquid-cooled server security management method, the method includes:

[0009] Step 1: Real-time collect thermodynamic parameters in the liquid-cooling system through a multi-modal sensor network. The multi-modal sensor network includes an infrared thermal imaging sensor, an ultrasonic flowmeter, and a piezoelectric vibration sensor, and obtain the surface temperature distribution of the server chip, the coolant flow rate, and the pipeline vibration spectrum data at a synchronous frequency of 100 Hz;

[0010] Step 2: Use the extended Kalman filter algorithm to perform spatio-temporal alignment processing on the multi-source heterogeneous data collected in Step 1, construct a three-dimensional thermal field digital twin model, and extract the thermal gradient distribution, phase change trigger region, and flow dead zone features through a graph convolutional network;

[0011] Step 3: Design a distributed cooling strategy based on the federated learning framework. Each server node locally trains a thermal dynamic prediction model based on the long short-term memory network (LSTM) algorithm, and adds Gaussian noise to the model gradient using the differential privacy mechanism, and uploads it to the central aggregator for global parameter optimization after encryption;

[0012] Step 4: Use the proximal policy optimization algorithm in reinforcement learning to generate dynamic control instructions. Taking the thermal field gradient and phase change state as inputs, output the opening degree of the proportional valve of each cold plate branch and the trigger timing of the piezoelectric microbubble generator;

[0013] Step 5: Use the quantum-derived simulated annealing algorithm to optimize the coolant flow distribution. Define the objective function as the weighted sum of minimizing the thermal field non-uniformity and the pump power consumption, and iteratively solve the global optimal flow distribution scheme;

[0014] Step 6: Implement edge preprocessing of sensor data on the FPGA hardware, including noise reduction, normalization, and feature compression, generate a low-dimensional feature vector and then transmit it to the central control system;

[0015] Step 7: Design a dual-threshold phase change control mechanism. When the local temperature exceeds 60 °C, trigger the gel state conversion of the phase change coolant, and enhance the latent heat absorption through the piezoelectric microbubble generator; when the temperature drops below 55 °C, switch to the laminar flow mode to resume the liquid cycle;

[0016] Step 8: Establish a cooling strategy update log based on the blockchain, record the local model parameters of each node and the global aggregation result, and ensure the traceability and anti-tampering of the control instructions;

[0017] Step 9: Send the control instructions generated in Step 4 to the magnetic levitation centrifugal pump, proportional valve, and piezoelectric actuator via the CAN bus to achieve closed-loop feedback control.

[0018] Preferably, the spatial arrangement rule of the infrared thermal imaging sensors in Step 1 is: deploy a micro-sensor array on the surfaces of the CPU, GPU, and NPU chips at a spacing of 0.5 mm, and use non-uniform topology coverage in the hot spot areas of the PCB.

[0019] Preferably, the federated learning framework in Step 3 adopts a horizontal federated architecture, and the central aggregator fuses the model gradients of each node through the dynamic weighted average algorithm, and the weights are updated inversely by the node data quality and historical prediction errors.

[0020] Preferably, the quantum-derived simulated annealing algorithm in Step 5 introduces the quantum tunneling effect, and by adjusting the annealing rate and quantum fluctuation parameters, it avoids local optimal solutions and accelerates convergence.

[0021] Preferably, the triggering logic of the piezoelectric microbubble generator in Step 7 includes: according to the prediction result of the thermal field gradient, directionally initiate the nucleate boiling of the target cold plate, the electrode arrangement adopts an interleaved circular array, and the driving frequency is 1 - 10 kHz.

[0022] Preferably, the blockchain log in Step 8 adopts the sharding storage technology, encrypts the log data in blocks according to the time stamp and distributes it to be stored in each server node, and realizes cross-node consistency verification through smart contracts.

[0023] Preferably, the topological structure of the graph convolutional network in Step 2 is a dynamic adaptive type, automatically adjusts the node connection weights according to the sensor deployment density, and uses the multi-head attention mechanism to strengthen key feature extraction.

[0024] Preferably, the reward function of the proximal policy optimization algorithm in Step 4 is defined as:

[0025] R = α·ΔT gradient +β·ΔP pump -γ·ΔS vibration

[0026] where, ΔT gradient is the change amount of the thermal field gradient, ΔP pump is the change amount of the pump power consumption, ΔS vibration is the vibration spectrum anomaly index, and α, β, γ are weight coefficients.

[0027] Preferably, the feature compression in Step 6 adopts an autoencoder model, the input layer is 128-dimensional original features, the hidden layer is compressed to 16 dimensions, and the feature robustness is improved through adversarial training.

[0028] Preferably, the present invention further includes a liquid-cooled server security management system, which includes the following modules:

[0029] Multimodal data acquisition module: Through a multimodal sensor network including an infrared thermal imaging sensor, an ultrasonic flowmeter, and a piezoelectric vibration sensor, it collects the surface temperature distribution of server chips, the coolant flow rate, and the pipeline vibration spectrum data in the liquid-cooled system in real time at a synchronous frequency of 100Hz;

[0030] Data processing and modeling module: Adopt the extended Kalman filter algorithm to perform spatio-temporal alignment processing on the collected multi-source heterogeneous data, construct a three-dimensional thermal field digital twin model, and extract the thermal gradient distribution, phase change trigger area, and flow dead zone features by means of a graph convolutional network;

[0031] Distributed cooling strategy design module: Based on the federated learning framework, each server node locally trains a thermal dynamic prediction model based on the long short-term memory network (LSTM) algorithm, adds Gaussian noise to the model gradient using the differential privacy mechanism, and uploads it to the central aggregator for global parameter optimization after encryption;

[0032] Dynamic control instruction generation module: Use the proximal policy optimization algorithm in reinforcement learning, with the thermal field gradient and phase change state as inputs, to generate dynamic control instructions for the opening of the proportional valve of each cold plate branch and the trigger timing of the piezoelectric microbubble generator;

[0033] Coolant flow optimization module: Adopt the quantum-derived simulated annealing algorithm to optimize the coolant flow distribution, define the objective function as the weighted sum of minimizing the thermal field non-uniformity and pump power consumption, and iteratively solve the global optimal flow distribution scheme;

[0034] Edge preprocessing module: Denoise, normalize, and compress the features of the sensor data on the FPGA hardware, generate a low-dimensional feature vector, and then transmit it to the central control system;

[0035] Phase change control module: Design a double-threshold phase change control mechanism. When the local temperature exceeds 60°C, trigger the gel state conversion of the phase change coolant, and enhance the latent heat absorption through the piezoelectric microbubble generator; when the temperature drops below 55°C, switch to the laminar flow mode to resume the liquid cycle;

[0036] Log recording module: Establish a blockchain-based cooling strategy update log, record the local model parameters of each node and the global aggregation results, and ensure the traceability and anti-tampering of control instructions;

[0037] Instruction issuance and closed-loop control module: Send the generated control instructions to the magnetic levitation centrifugal pump, proportional valve, and piezoelectric driver through the CAN bus to achieve closed-loop feedback control.

[0038] Compared with the prior art, the beneficial effects of the present invention are:

[0039] Through a multi-modal sensor network, the present invention collects data on the surface temperature distribution of the server chip, the coolant flow rate, and the pipeline vibration spectrum at a synchronous frequency of 100 Hz. Combined with the fine deployment of infrared thermal imaging sensors at key positions, it can accurately obtain the thermodynamic parameters of the liquid cooling system. The extended Kalman filter algorithm is used to perform spatio-temporal alignment on multi-source heterogeneous data. The constructed three-dimensional thermal field digital twin model can visually present the thermal field distribution. The key features extracted by the graph convolutional network provide an accurate basis for subsequent decision-making. Compared with traditional monitoring and data processing methods, the accuracy and usability of the data are greatly improved.

[0040] Based on the distributed cooling strategy of the federated learning framework, each server node locally trains a thermal dynamic prediction model, which not only protects data privacy but also further enhances data security using the differential privacy mechanism. The central aggregator fuses the model gradients through the dynamic weighted average algorithm to make the global model more accurate. Experiments show that the prediction accuracy of this model is more than 20% higher than that of traditional single models, which can predict the thermal change trend in advance, gain time for cooling strategy adjustment, and effectively avoid server overheating. In reinforcement learning, the proximal policy optimization algorithm takes the thermal field gradient and phase change state as inputs to generate dynamic control instructions, accurately adjusting the opening of the proportional valves of each cold plate branch and the triggering timing of the piezoelectric microbubble generator. The design of the reward function comprehensively considers the thermal field gradient, pump power consumption, and vibration conditions, optimizing the system performance.

[0041] The quantum-derived simulated annealing algorithm optimizes the coolant flow rate distribution. By introducing the quantum tunneling effect, it avoids local optimal solutions and quickly finds the global optimal flow rate distribution scheme. The objective function takes into account both the thermal field non-uniformity and pump power consumption. In practical applications, this algorithm reduces the pump power consumption by about 12% while ensuring the heat dissipation effect, improving the energy utilization efficiency. The dual-threshold phase change control mechanism automatically switches the coolant state according to temperature changes. When the local temperature exceeds 60 °C, it triggers the gel state conversion of the phase change coolant and enhances latent heat absorption; when the temperature drops below 55 °C, it resumes liquid circulation. This intelligent control method can respond to temperature changes in a timely manner, ensuring that the server can maintain good heat dissipation performance under different loads and extending the service life of the server.

[0042] Blockchain-based Cooling Strategy Update Log. By adopting sharding storage technology and smart contracts, it ensures the traceability and anti-tampering of control instructions. Any operation can be accurately recorded and verified, improving the security and reliability of the system. In terms of data privacy protection, the differential privacy mechanism effectively prevents data leakage, providing strong support for the secure operation of the data center. The edge preprocessing implemented by FPGA hardware denoises, normalizes, and compresses the features of sensor data, reducing the data transmission volume and the processing burden on the central control system. The closed-loop feedback control implemented by the CAN bus enables the system to adjust control instructions in real time according to the actual operating state, ensuring the stable operation of the liquid cooling system and improving the response speed and control accuracy of the overall system. Description of the Drawings

[0043] Figure 1 It is the working principle diagram of a liquid-cooled server security management method described in the present invention;

[0044] Figure 2 It is the flowchart of the collaborative work of federated learning and traffic optimization;

[0045] Figure 3 It is the flowchart of blockchain-based cooling strategy log management. Detailed Implementation Modes

[0046] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0047] Please refer to Figures 1-3 , the present invention provides a technical solution: a liquid-cooled server security management method, and the method includes:

[0048] Utilize a multi-modal sensor network composed of an infrared thermal imaging sensor, an ultrasonic flowmeter, and a piezoelectric vibration sensor to collect key data in the liquid cooling system at a synchronous frequency of 100 Hz in real time, including the surface temperature distribution of the server chip, the coolant flow rate, and the pipeline vibration spectrum data. These data can comprehensively reflect the operating state of the liquid cooling system and provide a basis for subsequent analysis and control.

[0049] The extended Kalman filter algorithm is used to perform spatio-temporal alignment processing on the collected multi-source heterogeneous data, eliminating the inconsistencies in time and space of the data. On this basis, a three-dimensional thermal field digital twin model is constructed, which can accurately simulate the thermal field distribution of the liquid cooling system. At the same time, key features such as thermal gradient distribution, phase change trigger region, and flow dead zone are extracted through a graph convolutional network, providing a basis for subsequent decision-making.

[0050] A distributed cooling strategy is designed based on the federated learning framework. Each server node locally trains a thermal dynamic prediction model using the long short-term memory network (LSTM) algorithm, which can learn the time series characteristics of thermal data and predict thermal dynamic changes. The differential privacy mechanism is used to add Gaussian noise to the model gradients, and on the premise of protecting data privacy, the encrypted model gradients are uploaded to the central aggregator for global parameter optimization, so as to obtain a more accurate thermal dynamic prediction model.

[0051] Dynamic control instruction generation: Using the proximal policy optimization algorithm in reinforcement learning, with the thermal field gradient and phase change state as inputs, dynamic control instructions for the opening of the proportional valve of each cold plate branch and the trigger timing of the piezoelectric microbubble generator are generated. By continuously optimizing these instructions, precise control of the liquid cooling system is achieved, improving the heat dissipation efficiency.

[0052] The quantum-derived simulated annealing algorithm is used to optimize the coolant flow distribution. The objective function is defined as the weighted sum of minimizing the thermal field non-uniformity and the pump power consumption. Through iterative solution, a globally optimal flow distribution scheme is obtained, reducing the pump power consumption while ensuring the heat dissipation effect and improving the energy utilization efficiency.

[0053] Edge preprocessing of sensor data is performed on the FPGA hardware, including noise reduction, normalization, and feature compression. The noise interference in the data is removed through noise reduction processing to improve the data quality; normalization processing makes different types of data have a unified scale for subsequent analysis; an autoencoder model is used for feature compression, compressing the 128-dimensional original features to 16 dimensions, and improving the feature robustness through adversarial training. After generating the low-dimensional feature vector, it is transmitted to the central control system, reducing the data transmission volume and processing burden.

[0054] A dual-threshold phase change control mechanism is designed. When the local temperature exceeds 60 °C, the gel state conversion of the phase change coolant is triggered, and the latent heat absorption is enhanced through the piezoelectric microbubble generator to improve the heat dissipation capacity; when the temperature drops below 55 °C, it switches to the laminar flow mode to restore the liquid circulation, ensuring the stable operation of the system.

[0055] Establish a blockchain-based cooling strategy update log to record the local model parameters of each node and the global aggregation results. Use sharding storage technology to encrypt and distribute the log data in chunks according to timestamps to each server node, and implement cross-node consistency verification through smart contracts to ensure the traceability and anti-tampering of control instructions, improving the security and reliability of the system.

[0056] Send the generated control instructions to the magnetic levitation centrifugal pump, proportional valve, and piezoelectric actuator through the CAN bus to achieve closed-loop feedback control of the liquid cooling system. Continuously adjust the control instructions according to the actual operating state of the system to keep the liquid cooling system in the best operating state.

[0057] The present invention will be further described below in conjunction with Embodiments 1 to 5:

[0058] Embodiment 1:

[0059] In a liquid-cooled server, the CPU, GPU, and NPU chips are the main heat-generating components, and their temperature distribution is crucial for the performance and stability of the server. Therefore, deploy a micro-sensor array on the surfaces of these chips at a spacing of 0.5 mm. This close spacing setting can carefully capture the temperature differences at different positions on the chip surface and obtain high-precision temperature distribution data. For example, on the surface of the CPU chip of a high-performance server, the micro-sensor array deployed at this spacing can clearly distinguish the temperature changes in the core area and the surrounding area, and even the slightest temperature fluctuations can be accurately detected.

[0060] At the same time, considering that there are some other hot spots on the PCB board in addition to the chips, these areas may also affect the overall thermal balance of the system. To comprehensively monitor the temperature of the PCB board, use non-uniform topological coverage in the hot spots of the PCB. According to the heat generation characteristics and potential temperature change trends of the hot spots, adjust the distribution density of the sensors accordingly. Increase the number of sensors deployed in areas with relatively concentrated heat generation and large temperature changes; while in areas with relatively stable temperatures, appropriately reduce the number of sensors. This non-uniform topological coverage method can not only meet the key monitoring requirements of key hot spots but also reasonably control the number of sensors used and reduce costs while ensuring the monitoring accuracy. Through this deployment method of infrared thermal imaging sensors, it is possible to obtain the temperature distribution data of the server chip surface and the PCB hot spots in real time, comprehensively, and accurately, laying a solid foundation for the subsequent construction of an accurate three-dimensional thermal field digital twin model.

[0061] Embodiment 2:

[0062] In the liquid-cooled server security management system of the present invention, a federated learning framework with a horizontal federated architecture is adopted. Under this architecture, each server node has the same data feature space, but different data samples. Each server node uses local hot data to train a hot dynamic prediction model based on the long short-term memory network (LSTM) algorithm. For example, in a data center containing multiple server nodes, each node trains a model according to the data of the server chip temperature changing over time recorded by itself, and learns the hot dynamic change law of the server of this node.

[0063] However, in order to improve the accuracy and generalization ability of the model, it is necessary to fuse the models of each node. The central aggregator plays a key role in this process. It fuses the model gradients of each node through a dynamic weighted average algorithm. When calculating the weights, two important factors, namely the node data quality and the historical prediction error, are considered. For nodes with higher data quality, their data can more accurately reflect the actual situation, so higher weights are assigned; for nodes with smaller historical prediction errors, it indicates that the prediction ability of their models is stronger, and the same higher weights are given. The specific weight calculation method is as follows: Let the data quality score of node i be Q i , and the historical prediction error be E i , then the weight w i of node i is calculated by the formula where n is the total number of nodes participating in the federated learning. Through this dynamic weighted average algorithm, the advantages of each node can be fully utilized, making the fused global model more accurate and stable. In practical applications, after multiple rounds of model training and fusion, the prediction accuracy of the hot dynamic prediction model has been significantly improved compared with the single-node model, effectively improving the control accuracy of the liquid-cooled system.

[0064] Embodiment 3:

[0065] This embodiment details the application of the quantum-derived simulated annealing algorithm in the optimization of coolant flow distribution. By introducing the quantum tunneling effect and adjusting relevant parameters, it can effectively avoid the algorithm falling into a local optimal solution and find the globally optimal flow distribution scheme faster, improving the heat dissipation efficiency and energy utilization efficiency of the liquid-cooled system.

[0066] In the process of optimizing the coolant flow distribution of a liquid-cooled server, the quantum-derived simulated annealing algorithm is adopted. The core of this algorithm lies in the introduction of the quantum tunneling effect, which enables the algorithm to jump out of the local optimal trap during the process of searching for the optimal solution and is more likely to find the globally optimal solution.

[0067] First, the objective function is defined as the weighted sum of minimizing the thermal field non-uniformity and the pump power consumption. Let the thermal field non-uniformity be U, the pump power consumption be P, and the weighting coefficients be λ 1 and λ2 , the objective function F = λ 1 U + λ 2 P. The non-uniformity U of the thermal field can be measured by calculating the sum of the squared deviations of the temperatures at each point in the thermal field from the average temperature, i.e., where T i is the temperature at the i-th point in the thermal field, is the average temperature of the thermal field, and m is the number of points in the thermal field; the pump power consumption P can be calculated based on the operating parameters of the pump, such as flow rate, head, efficiency, etc., and the specific calculation formula is determined according to the pump model and operating characteristics.

[0068] During the execution of the algorithm, the search process of the algorithm is controlled by adjusting the annealing rate and the quantum fluctuation parameter. The annealing rate determines the rate of change of the probability of accepting a worse solution with time during the search process of the algorithm. When the annealing rate is slow, the algorithm can explore the solution space more fully, but the convergence rate will be slow; when the annealing rate is fast, the convergence rate of the algorithm increases, but it may miss the global optimal solution. The quantum fluctuation parameter affects the intensity of the quantum tunneling effect. A larger quantum fluctuation parameter will enhance the quantum tunneling effect, making it easier for the algorithm to jump out of the local optimal solution, but it may also cause the algorithm to be too random during the search process. In practical applications, through multiple experiments and optimizations, appropriate annealing rates and quantum fluctuation parameters are determined. For example, in a certain liquid-cooled server system, after a series of tests and adjustments, the annealing rate is set to 0.95 and the quantum fluctuation parameter is set to 0.5. Under this parameter setting, the algorithm can find the globally optimal flow rate distribution scheme in a relatively short time.

[0069] Example 4:

[0070] This example clarifies the triggering logic of the piezoelectric microbubble generator in the liquid cooling system. By directionally initiating the nucleate boiling of the target cold plate according to the predicted result of the thermal field gradient, as well as specific electrode arrangements and drive frequency settings, the latent heat absorption can be enhanced more efficiently, and the heat dissipation capacity of the liquid cooling system can be improved.

[0071] During the phase change cooling process of the liquid-cooled server, the piezoelectric microbubble generator plays a key role. Its triggering logic is designed based on the predicted result of the thermal field gradient. By analyzing the thermal field gradient, it can be predicted which regions have a rapid temperature rise and may require stronger heat dissipation measures. When the thermal field gradient in the target cold plate region is predicted to exceed a certain threshold, the nucleate boiling of the target cold plate is directionally initiated. For example, in the liquid cooling system of a certain server, through thermal field monitoring, it is found that the thermal field gradient in the cold plate region corresponding to a certain GPU chip increases sharply, and the system immediately starts the piezoelectric microbubble generator on this cold plate, prompting the coolant to undergo nucleate boiling in this region, thereby greatly enhancing the latent heat absorption and quickly reducing the chip temperature.

[0072] The electrode arrangement of the piezoelectric microbubble generator adopts an interleaved circular array. This arrangement can generate a uniform and strong electric field, enabling the coolant to more effectively generate microbubbles under the action of the electric field. The interleaved design can avoid the concentration and uneven distribution of the electric field, ensuring that the entire cold plate area can be fully cooled. At the same time, the driving frequency is set to 1 - 10 kHz. Within this frequency range, it can not only ensure that the piezoelectric microbubble generator effectively generates microbubbles, but also avoid energy loss and equipment damage caused by too high a frequency. A lower frequency may not be able to fully stimulate the nucleate boiling of the coolant, while too high a frequency may cause problems such as resonance, affecting the stability of the system. Through this trigger logic, electrode arrangement, and driving frequency setting, the piezoelectric microbubble generator can be started in time when needed in the liquid cooling system, precisely enhancing the heat dissipation effect of the target area and ensuring the stable operation of the server chip in a high-temperature environment.

[0073] Example 5:

[0074] This embodiment covers the storage and verification mechanism of blockchain logs, the topological structure of the graph convolutional network, the reward function of the proximal policy optimization algorithm, and the feature compression autoencoder model. The combination of these technologies helps to improve the security, data analysis ability, control decision-making ability, and data processing efficiency of the liquid-cooled server security management system.

[0075] Blockchain Log and Smart Contract Verification: In the liquid-cooled server security management system, establishing a blockchain-based cooling policy update log is an important measure to ensure the security and traceability of the system. Using the sharding storage technology, the log data is encrypted in blocks according to timestamps and distributedly stored in each server node. Each log data block corresponding to a timestamp contains key information such as the local model parameters and global aggregation results of each node at that moment. For example, during the daily operation of the server, a log data block is generated every hour, recording the model training situation of each node and the update result of the global model within that hour. Through this sharding storage method, even if a node fails or the data is tampered with, it will not affect the integrity of the entire log system. At the same time, smart contracts are used to achieve cross-node consistency verification. A smart contract is an automatically executed contract clause deployed on the blockchain. When a node needs to verify the consistency of the log data, the smart contract will automatically check the corresponding log data blocks stored in other nodes, and through preset verification algorithms such as hash verification and digital signature verification, ensure the consistency of the data of each node. If it is found that the data of a certain node is inconsistent with that of other nodes, the smart contract will issue an alarm and take corresponding repair measures to ensure the accuracy and reliability of the log data.

[0076] Graph Convolutional Network Topology: The graph convolutional network is used to extract the characteristics of thermal gradient distribution, phase change trigger regions, and flow dead zones in the data processing and modeling module. Its topology is of the dynamic adaptive type and can automatically adjust the node connection weights according to the sensor deployment density. When the sensor deployment density is high in a certain area, it indicates that more accurate data analysis is required in this area. The graph convolutional network will automatically increase the connection weights between nodes in this area, enabling the model to pay more attention to the data characteristics of these areas. Conversely, in areas with a low sensor deployment density, the node connection weights are appropriately reduced. At the same time, the multi-head attention mechanism is adopted to strengthen the extraction of key features. The multi-head attention mechanism extracts features from different perspectives through multiple different attention heads and then fuses these features. For example, when analyzing thermal field data, one attention head may focus more on the change trend of the thermal gradient, while another attention head focuses on the location information of the phase change trigger region. Through the multi-head attention mechanism, the key features in these different aspects are fused together, improving the model's understanding and analysis ability of the thermal field data and providing a more accurate basis for subsequent decision-making.

[0077] Proximal Policy Optimization Algorithm Reward Function: In the dynamic control instruction generation module, the reward function of the proximal policy optimization algorithm is defined as R = α·ΔT gradient +β·ΔP pump -γ·ΔS vibration . Among them, ΔT gradient is the change amount of the thermal field gradient, which reflects the change of the thermal field distribution. The larger the change amount of the thermal field gradient, the better the heat dissipation effect of the system and the greater the positive contribution to the reward function; ΔP pump is the change amount of the pump power consumption. The reduction of the pump power consumption means the improvement of energy utilization efficiency. Therefore, when this change amount is negative, it has a positive contribution to the reward function; ΔS vibration is the vibration spectrum anomaly index. The smaller this index is, the more stable the pipeline vibration and the more reliable the system operation, and the greater the positive contribution to the reward function. α, β, and γ are weight coefficients, which are adjusted according to actual requirements and system characteristics. For example, in a server scenario with high requirements for heat dissipation effect, the value of α can be appropriately increased to make the algorithm pay more attention to the optimization of the thermal field gradient; while in an environment more sensitive to energy efficiency, the weight of β can be increased. By reasonably setting the reward function, the proximal policy optimization algorithm can generate more dynamic control instructions that meet the system requirements and achieve precise control of the liquid cooling system.

[0078] Feature Compression Autoencoder Model: In the edge preprocessing module, an autoencoder model is used for feature compression. The input layer of the autoencoder model is 128-dimensional raw features, which contain rich information collected by sensors. However, due to their high dimensionality, they are not conducive to data transmission and processing. Through the autoencoder model, these raw features are compressed to a 16-dimensional hidden layer. During the compression process, the autoencoder model learns the important feature representations of the raw features and discards some redundant information. For example, when processing temperature, flow rate, and vibration data collected by sensors, the model can automatically extract key features closely related to the operating state of the liquid cooling system, such as temperature change trends and abnormal fluctuations in flow rate. At the same time, the feature robustness is improved through adversarial training. Adversarial training is to introduce adversarial samples during the training process of the autoencoder model, enabling the model to learn how to resist adversarial attacks, thereby improving the model's robustness to noise and abnormal data. The autoencoder model after adversarial training can more stably output low-dimensional feature vectors in the face of sensor failures or data transmission interferences, ensuring the quality of the data received by the central control system and providing reliable support for subsequent analysis and decision-making.

[0079] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article, or device.

[0080] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A liquid cooling server security management method, characterized in that: The following steps are involved: Step 1: Real-time collection of thermodynamic parameters in the liquid cooling system through a multimodal sensor network, which includes an infrared thermal imaging sensor, an ultrasonic flow meter, and a piezoelectric vibration sensor, to obtain server chip surface temperature distribution, coolant flow rate, and pipeline vibration spectrum data at a synchronous frequency of 100 Hz; Step 2: Use the extended Kalman filter algorithm to perform spatiotemporal alignment processing on the multi-source heterogeneous data collected in step 1, build a three-dimensional thermal field digital twin model, and extract the thermal gradient distribution, phase change trigger area and flow dead zone characteristics through the graph convolution network; Step 3: Design a distributed cooling strategy based on the federated learning framework. Each server node locally trains a thermal dynamic prediction model based on the long short-term memory network LSTM algorithm, and uses a differential privacy mechanism to add Gaussian noise to the model gradient. After encryption, it is uploaded to the central aggregator for global parameter optimization. Step 4: Generate dynamic control instructions using the proximal strategy optimization algorithm in reinforcement learning, taking the thermal field gradient and phase change state as input, and output the proportional valve opening of each cold plate branch and the triggering timing of the piezoelectric microbubble generator; Step 5: Use the quantum-derived simulated annealing algorithm to optimize the coolant flow distribution, define the objective function as minimizing the weighted sum of thermal field non-uniformity and pump power consumption, and solve the global optimal flow distribution solution through iteration; Step 6: Implement edge preprocessing of sensor data on FPGA hardware, including noise reduction, normalization, and feature compression, generate low-dimensional feature vectors, and transmit them to the central control system; Step 7: Design a dual-threshold phase change control mechanism to trigger the gel state conversion of the phase change coolant when the local temperature exceeds 60°C, and enhance latent heat absorption through the piezoelectric microbubble generator; when the temperature drops below 55°C, switch to laminar flow mode to restore liquid circulation; Step 8: Establish a blockchain-based cooling strategy update log to record the local model parameters of each node and the global aggregation results to ensure the traceability and tamper resistance of control instructions; Step 9: Send the control command generated in step 4 to the magnetic levitation centrifugal pump, proportional valve and piezoelectric driver through the CAN bus to achieve closed-loop feedback control.

2. The liquid cooling server safety management system and method according to claim 1, characterized in that: The spatial arrangement rule of the infrared thermal imaging sensor in step 1 is: deploy a micro sensor array with a 0.5 mm pitch on the surface of the CPU, GPU, and NPU chips, and use non-uniform topology coverage in the hot spot area of ​​the PCB.

3. The liquid cooling server safety management system and method according to claim 1, characterized in that: The federated learning framework described in step 3 adopts a horizontal federated architecture. The central aggregator fuses the model gradients of each node through a dynamic weighted average algorithm, and the weights are updated inversely by the node data quality and historical prediction errors.

4. The liquid cooling server safety management system and method according to claim 1, characterized in that: The quantum-derived simulated annealing algorithm described in step 5 introduces the quantum tunneling effect, and avoids local optimal solutions and accelerates convergence by adjusting the annealing rate and quantum fluctuation parameters.

5. The liquid cooling server safety management system and method according to claim 1, characterized in that: The triggering logic of the piezoelectric microbubble generator in step 7 includes: according to the thermal field gradient prediction result, the nucleate boiling of the target cold plate is started directionally, the electrodes are arranged in a staggered annular array, and the driving frequency is 1-10kHz.

6. The liquid cooling server safety management system and method according to claim 1, characterized in that: The blockchain log described in step 8 uses shard storage technology to encrypt the log data in blocks according to timestamps and then store them in distributed form on each server node, and implements cross-node consistency verification through smart contracts.

7. The liquid cooling server safety management system and method according to claim 1, characterized in that: The topological structure of the graph convolutional network described in step 2 is dynamically adaptive, which automatically adjusts the node connection weights according to the sensor deployment density and adopts a multi-head attention mechanism to enhance the extraction of key features.

8. The liquid cooling server safety management system and method according to claim 1, characterized in that: The reward function of the proximal policy optimization algorithm described in step 4 is defined as: R=α·ΔT gradient +β·ΔP pump -γ·ΔS vibration Where, ΔT gradient is the thermal field gradient change, ΔP pump is the change in pump power consumption, ΔS vibration is the vibration spectrum abnormality index, and α, β, and γ are weight coefficients.

9. The liquid cooling server safety management system and method according to claim 1, characterized in that: The feature compression described in step 6 adopts an autoencoder model, with the input layer being 128-dimensional original features, the hidden layer being compressed to 16 dimensions, and adversarial training being used to improve feature robustness.

10. A liquid cooling server safety management system, characterized in that: Includes the following modules: Multimodal data acquisition module: Through a multimodal sensor network including infrared thermal imaging sensors, ultrasonic flow meters and piezoelectric vibration sensors, the server chip surface temperature distribution, coolant flow rate and pipeline vibration spectrum data in the liquid cooling system are collected in real time at a synchronous frequency of 100Hz; Data processing and modeling module: The extended Kalman filter algorithm is used to perform spatiotemporal alignment processing on the collected multi-source heterogeneous data, build a three-dimensional thermal field digital twin model, and use the graph convolutional network to extract the thermal gradient distribution, phase change trigger area and flow dead zone characteristics; Distributed cooling strategy design module: Based on the federated learning framework, each server node locally trains a thermal dynamic prediction model based on the long short-term memory network LSTM algorithm, uses a differential privacy mechanism to add Gaussian noise to the model gradient, and uploads it to the central aggregator after encryption for global parameter optimization; Dynamic control command generation module: using the proximal strategy optimization algorithm in reinforcement learning, with the thermal field gradient and phase change state as input, to generate dynamic control commands for the proportional valve opening of each cold plate branch and the triggering timing of the piezoelectric microbubble generator; Coolant flow optimization module: uses quantum-derived simulated annealing algorithm to optimize coolant flow distribution, defines the objective function as minimizing the weighted sum of thermal field non-uniformity and pump power consumption, and solves the global optimal flow distribution solution through iteration; Edge preprocessing module: performs noise reduction, normalization and feature compression on sensor data on FPGA hardware, generates low-dimensional feature vectors and transmits them to the central control system; Phase change control module: A dual-threshold phase change control mechanism is designed to trigger the gel state conversion of the phase change coolant when the local temperature exceeds 60°C, and enhance the latent heat absorption through the piezoelectric microbubble generator; when the temperature drops below 55°C, it switches to laminar flow mode to restore liquid circulation; Logging module: Establish a blockchain-based cooling strategy update log to record the local model parameters and global aggregation results of each node to ensure the traceability and tamper resistance of control instructions; Instruction issuance and closed-loop control module: The generated control instructions are issued to the magnetic levitation centrifugal pump, proportional valve and piezoelectric driver through the CAN bus to achieve closed-loop feedback control.

Citation Information

Patent Citations

  • Heat absorption and insulation structure, battery assembly and power utilization system

    CN118231846A

  • Intelligent water conservancy inspection method, device and equipment and storage medium

    CN119106880A

  • Dynamic correction method and system for power curve of wind power plant

    CN119397175A

  • Quantum dot enhanced dynamic power consumption management method and system and related equipment

    CN119397979A

  • Virtual power plant system design method based on multilevel federated learning

    CN119416605A

Cited By

  • Heat flow dynamic tracking AI prediction type server temperature control system and device

    CN120371102A

  • New energy automobile high-voltage connector thermal management system and optimization control method

    CN120500020A

  • High-power single-phase immersion liquid cooling data center cabinet

    CN120529567A

  • Single-phase and two-phase immersion liquid cooling method and system based on AI intelligent decision

    CN120730713A

  • Bus encryption method, device and equipment based on simulated annealing optimization

    CN120811676A