Liquid cooling server safety management system and method
Through multimodal sensor network and advanced data processing algorithms, a three-dimensional thermal field digital twin model of liquid-cooled server is built to realize dynamic cooling strategy optimization and security management, solving the complexity and data security problems of liquid-cooled servers in thermal dissipation management, and improving the stability and efficiency of the system.
Patent Information
- Application Number
- CN202510294658.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-03-13
AI Technical Summary
In the thermal management of liquid cooling servers, there is complex temperature distribution and difficult to accurately monitor dynamic changes, difficult to effectively integrate and analyze multi-source heterogeneous data, difficult to achieve data security and privacy protection, and lack of dynamic optimization and intelligent control of cooling strategies, resulting in low heat dissipation efficiency, waste of energy and poor system stability.
A multimodal sensor network is used to collect thermodynamic parameters in real time, combine extended Kalman filtering and graph convolution network to build a three-dimensional thermal field digital twin model, generate dynamic cooling strategies based on federated learning and reinforcement learning, optimize traffic allocation using quantum-derived simulation annealing algorithm, design a dual-threshold phase change control mechanism, and record control instructions through blockchain to achieve closed-loop feedback control.
Accurately obtain the thermodynamic parameters of the liquid-cooled system, improve data processing accuracy and safety, optimize cooling strategies, reduce pump power consumption, ensure system stability and reliability, extend server life, and improve heat dissipation efficiency and energy utilization efficiency.
Smart Images

Figure CN120152228B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of liquid cooling servers, and in particular to a liquid cooling server security management system and method. Background Art
[0002] With the rapid development of information technology, data centers are expanding, and the computing power and power density of servers are also increasing. Traditional air cooling is no longer able to meet the cooling needs of high-performance servers. Liquid cooling technology, due to its efficient heat dissipation performance, is gradually becoming the mainstream choice. However, the actual application of liquid-cooled servers still faces many challenges.
[0003] From the perspective of heat dissipation management, the temperature distribution of various components within the liquid cooling system is complex and changes dynamically. Server chips (such as CPUs, GPUs, and NPUs) generate a large amount of heat during operation. The heat generation conditions of different chips and different areas of the same chip vary greatly. How to accurately obtain these complex thermodynamic parameters and effectively manage them has become a key issue. Existing sensor deployment methods make it difficult to comprehensively and accurately monitor the temperature distribution of chip surfaces and hot spots on PCB boards. As a result, local overheating risks cannot be discovered in a timely manner, affecting the performance and lifespan of the server. For example, in certain high-performance computing scenarios, excessively high local chip temperatures may trigger thermal runaway, leading to system failure.
[0004] Liquid cooling systems involve heterogeneous data from multiple sources, including temperature, flow rate, and vibration. This data varies in time and space, making it difficult for traditional data processing methods to effectively integrate and analyze it, thus failing to provide an accurate basis for developing cooling strategies. Furthermore, existing thermal dynamic prediction models have limited accuracy and are unable to accurately predict the system's thermal trends. This often causes cooling strategies to lag behind actual needs, reducing cooling efficiency.
[0005] Data security and privacy protection within a distributed architecture are equally crucial. In a liquid cooling system with multiple server nodes, each node's data contains sensitive information. Enabling data sharing and collaborative processing without compromising privacy is a pressing issue during model training and policy optimization. Furthermore, cooling policy updates and control command execution lack effective traceability and tamper resistance, making it difficult to quickly locate and resolve any issues that arise.
[0006] Furthermore, the rationality of coolant flow distribution directly impacts cooling effectiveness and energy consumption. Current flow distribution methods are often based on experience or simple algorithms, failing to dynamically optimize based on the server's real-time thermal state. This results in insufficient cooling in some areas and wasted energy in others. Furthermore, the phase change control mechanism during phase change cooling is not intelligent enough to accurately and promptly adjust the coolant state based on temperature fluctuations, impacting overall system performance. Summary of the Invention
[0007] The object of the present invention is to provide a liquid cooling server safety management system and method to solve the problems raised in the above background technology.
[0008] To achieve the above objectives, the present invention provides the following technical solution: a liquid cooling server security management method, the method comprising:
[0009] Step 1: Real-time acquisition of thermodynamic parameters within the liquid cooling system using a multimodal sensor network. This multimodal sensor network includes infrared thermal imaging sensors, ultrasonic flow meters, and piezoelectric vibration sensors. The network acquires server chip surface temperature distribution, coolant flow rate, and pipeline vibration spectrum data at a synchronous frequency of 100 Hz.
[0010] Step 2: Use the extended Kalman filter algorithm to perform spatiotemporal alignment on the multi-source heterogeneous data collected in step 1, build a 3D thermal field digital twin model, and extract the thermal gradient distribution, phase change trigger area, and flow dead zone characteristics through a graph convolutional network;
[0011] Step 3: Design a distributed cooling strategy based on the federated learning framework. Each server node locally trains a thermal dynamic prediction model based on the long short-term memory (LSTM) algorithm. A differential privacy mechanism is used to add Gaussian noise to the model gradients. The data is encrypted and uploaded to the central aggregator for global parameter optimization.
[0012] Step 4: Use the proximal strategy optimization algorithm in reinforcement learning to generate dynamic control instructions. Taking the thermal field gradient and phase change state as input, it outputs the proportional valve opening of each cold plate branch and the triggering timing of the piezoelectric microbubble generator.
[0013] Step 5: Use the quantum-derived simulated annealing algorithm to optimize the coolant flow distribution. Define the objective function as minimizing the weighted sum of thermal field non-uniformity and pump power consumption, and iterate to find the global optimal flow distribution solution.
[0014] Step 6: Implement edge preprocessing of sensor data on FPGA hardware, including noise reduction, normalization, and feature compression, to generate low-dimensional feature vectors that are then transmitted to the central control system.
[0015] Step 7: Design a dual-threshold phase change control mechanism. When the local temperature exceeds 60°C, the phase change coolant is triggered to transform into a gel state and enhance latent heat absorption through a piezoelectric microbubble generator. When the temperature drops below 55°C, the system switches to laminar flow mode to resume liquid circulation.
[0016] Step 8: Establish a blockchain-based cooling strategy update log to record the local model parameters of each node and the global aggregation results to ensure the traceability and tamper resistance of control instructions;
[0017] Step 9: Send the control instructions generated in step 4 to the magnetic levitation centrifugal pump, proportional valve and piezoelectric driver through the CAN bus to achieve closed-loop feedback control.
[0018] Preferably, the spatial arrangement rule of the infrared thermal imaging sensor in step 1 is: deploying a micro sensor array with a 0.5 mm pitch on the surface of the CPU, GPU, and NPU chips, and adopting non-uniform topological coverage in the PCB hot spot area.
[0019] Preferably, the federated learning framework described in step 3 adopts a horizontal federated architecture, and the central aggregator fuses the model gradients of each node through a dynamic weighted average algorithm, and the weights are updated inversely by the node data quality and historical prediction error.
[0020] Preferably, the quantum-derived simulated annealing algorithm in step 5 introduces a quantum tunneling effect, and avoids local optimal solutions and accelerates convergence by adjusting the annealing rate and quantum fluctuation parameters.
[0021] Preferably, the triggering logic of the piezoelectric microbubble generator in step 7 includes: directionally starting nucleate boiling of the target cold plate according to the thermal field gradient prediction result, the electrodes are arranged in a staggered annular array, and the driving frequency is 1-10 kHz.
[0022] Preferably, the blockchain log described in step 8 adopts shard storage technology, and the log data is encrypted in blocks according to timestamps and then distributedly stored in each server node, and cross-node consistency verification is achieved through smart contracts.
[0023] Preferably, the topology of the graph convolutional network in step 2 is dynamically adaptive, automatically adjusting the node connection weights according to the sensor deployment density, and using a multi-head attention mechanism to enhance key feature extraction.
[0024] Preferably, the reward function of the proximal policy optimization algorithm in step 4 is defined as:
[0025] R=α·ΔT gradient +β·ΔP pump -γ·ΔS vibration
[0026] Where, ΔT gradient is the thermal field gradient change, ΔP pump is the change in pump power consumption, ΔS vibration is the vibration spectrum anomaly index, and α, β, and γ are weight coefficients.
[0027] Preferably, the feature compression in step 6 adopts an autoencoder model, the input layer is 128-dimensional original features, the hidden layer is compressed to 16 dimensions, and the feature robustness is improved through adversarial training.
[0028] Preferably, the present invention further includes a liquid cooling server safety management system, comprising the following modules:
[0029] Multimodal data acquisition module: This module uses a multimodal sensor network consisting of infrared thermal imaging sensors, ultrasonic flow meters, and piezoelectric vibration sensors to collect real-time data on server chip surface temperature distribution, coolant flow rate, and pipe vibration spectrum within the liquid cooling system at a synchronous frequency of 100 Hz.
[0030] Data processing and modeling module: This module uses the extended Kalman filter algorithm to perform spatiotemporal alignment on the collected multi-source heterogeneous data, constructs a three-dimensional thermal field digital twin model, and uses a graph convolutional network to extract thermal gradient distribution, phase change trigger areas, and flow dead zone features.
[0031] Distributed Cooling Strategy Design Module: Based on a federated learning framework, each server node locally trains a thermal dynamic prediction model based on the Long Short-Term Memory (LSTM) algorithm. A differential privacy mechanism is used to add Gaussian noise to the model gradients, which are then encrypted and uploaded to a central aggregator for global parameter optimization.
[0032] Dynamic control instruction generation module: This module uses the proximal strategy optimization algorithm in reinforcement learning to generate dynamic control instructions for the proportional valve opening of each cold plate branch and the triggering timing of the piezoelectric microbubble generator, taking the thermal field gradient and phase change state as input;
[0033] Coolant flow optimization module: This module uses a quantum-derived simulated annealing algorithm to optimize coolant flow distribution. The objective function is defined as minimizing the weighted sum of thermal field non-uniformity and pump power consumption. The global optimal flow distribution solution is solved through iteration.
[0034] Edge pre-processing module: performs noise reduction, normalization, and feature compression on sensor data on FPGA hardware, generates low-dimensional feature vectors, and transmits them to the central control system;
[0035] Phase change control module: A dual-threshold phase change control mechanism is designed to trigger the gel state conversion of the phase change coolant when the local temperature exceeds 60°C, and enhance latent heat absorption through a piezoelectric microbubble generator. When the temperature drops below 55°C, it switches to laminar flow mode to resume liquid circulation.
[0036] Logging module: Establishes a blockchain-based cooling strategy update log, records the local model parameters of each node and the global aggregation results, and ensures the traceability and tamper resistance of control instructions;
[0037] Instruction issuance and closed-loop control module: The generated control instructions are issued to the magnetic levitation centrifugal pump, proportional valve and piezoelectric driver through the CAN bus to achieve closed-loop feedback control.
[0038] Compared with the prior art, the present invention has the following beneficial effects:
[0039] This method uses a multimodal sensor network to collect data on server chip surface temperature distribution, coolant flow rate, and pipeline vibration spectrum at a synchronous frequency of 100Hz. Combined with the careful deployment of infrared thermal imaging sensors in key locations, it can accurately obtain the thermodynamic parameters of the liquid cooling system. Using the extended Kalman filter algorithm to align multi-source heterogeneous data in time and space, the constructed three-dimensional thermal field digital twin model can intuitively present the thermal field distribution. The key features extracted by the graph convolutional network provide an accurate basis for subsequent decision-making. Compared with traditional monitoring and data processing methods, this method significantly improves data accuracy and usability.
[0040] Based on a federated learning framework, a distributed cooling strategy uses a local thermal dynamic prediction model trained on each server node, protecting data privacy while further enhancing data security through differential privacy mechanisms. A central aggregator fuses model gradients using a dynamic weighted average algorithm, enhancing global model accuracy. Experiments demonstrate that this model achieves over 20% higher prediction accuracy than traditional single models. This model can predict thermal trends in advance, allowing time for cooling strategy adjustments and effectively preventing server overheating. A proximal policy optimization algorithm, based on reinforcement learning, generates dynamic control commands based on thermal field gradients and phase transition states, precisely adjusting the proportional valve opening of each cold plate branch and the triggering timing of the piezoelectric microbubble generator. The reward function design comprehensively considers thermal field gradients, pump power consumption, and vibration to optimize system performance.
[0041] A quantum-derived simulated annealing algorithm optimizes coolant flow distribution, introducing the quantum tunneling effect to avoid local optimal solutions and quickly find the globally optimal flow distribution solution. The objective function takes into account both thermal field nonuniformity and pump power consumption. In practical applications, this algorithm reduces pump power consumption by approximately 12% while ensuring effective heat dissipation, thereby improving energy efficiency. A dual-threshold phase change control mechanism automatically switches the coolant state based on temperature changes. When the local temperature exceeds 60°C, the phase change coolant undergoes a gel-like transition and enhances latent heat absorption. When the temperature drops below 55°C, liquid circulation resumes. This intelligent control method responds promptly to temperature changes, ensuring that the server maintains good heat dissipation performance under varying loads, thereby extending the server's service life.
[0042] The blockchain-based cooling strategy update log utilizes sharded storage technology and smart contracts to ensure the traceability and tamper resistance of control commands. All operations are accurately recorded and verified, improving system security and reliability. In terms of data privacy protection, the differential privacy mechanism effectively prevents data leakage, providing strong support for the secure operation of the data center. Edge preprocessing implemented in FPGA hardware reduces noise, normalizes, and compresses features on sensor data, reducing data transmission volume and the processing burden on the central control system. Closed-loop feedback control implemented via the CAN bus enables the system to adjust control commands in real time based on actual operating conditions, ensuring stable operation of the liquid cooling system and improving the overall system's response speed and control accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 This is a working principle diagram of a liquid cooling server safety management method according to the present invention;
[0044] Figure 2 Flowchart for collaborative working of federated learning and traffic optimization;
[0045] Figure 3 Flowchart of blockchain-based cooling strategy log management. DETAILED DESCRIPTION
[0046] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0047] See also Figure 1-3 The present invention provides a technical solution: a liquid cooling server security management method, the method comprising:
[0048] A multimodal sensor network consisting of infrared thermal imaging sensors, ultrasonic flow meters, and piezoelectric vibration sensors collects key data from the liquid cooling system in real time at a synchronous frequency of 100Hz, including the surface temperature distribution of server chips, coolant flow rate, and pipeline vibration spectrum. This data comprehensively reflects the operating status of the liquid cooling system and provides a basis for subsequent analysis and control.
[0049] The extended Kalman filter algorithm is used to perform spatiotemporal alignment on the collected multi-source heterogeneous data, eliminating temporal and spatial inconsistencies. Based on this, a three-dimensional thermal field digital twin model is constructed, which accurately simulates the thermal field distribution of the liquid cooling system. Furthermore, a graph convolutional network is used to extract key features such as thermal gradient distribution, phase change triggering areas, and flow dead zones, providing a basis for subsequent decision-making.
[0050] A distributed cooling strategy is designed based on a federated learning framework. Each server node locally trains a thermal dynamics prediction model using a long short-term memory (LSTM) algorithm. This model learns the time series characteristics of thermal data and predicts thermal dynamics. A differential privacy mechanism is used to add Gaussian noise to the model gradients. While protecting data privacy, the encrypted model gradients are uploaded to a central aggregator for global parameter optimization, resulting in a more accurate thermal dynamics prediction model.
[0051] Dynamic Control Command Generation: Utilizing the proximal strategy optimization algorithm from reinforcement learning, the system uses the thermal field gradient and phase change state as inputs to generate dynamic control commands for the proportional valve opening of each cold plate branch and the triggering sequence of the piezoelectric microbubble generator. By continuously optimizing these commands, precise control of the liquid cooling system is achieved, improving heat dissipation efficiency.
[0052] A quantum-derived simulated annealing algorithm is used to optimize coolant flow distribution. The objective function is defined as minimizing the weighted sum of thermal field nonuniformity and pump power consumption. An iterative solution is used to determine the globally optimal flow distribution solution, ensuring effective cooling while reducing pump power consumption and improving energy efficiency.
[0053] Edge preprocessing of sensor data, including noise reduction, normalization, and feature compression, is performed on FPGA hardware. Noise reduction removes noise interference from the data, improving data quality. Normalization uniformly scales different types of data for ease of subsequent analysis. An autoencoder model is used for feature compression, compressing the original 128-dimensional features to 16 dimensions. Adversarial training is used to enhance feature robustness, generating low-dimensional feature vectors that are then transmitted to the central control system, reducing data transmission and processing overhead.
[0054] A dual-threshold phase-change control mechanism was designed. When the local temperature exceeds 60°C, the phase-change coolant transforms into a gel state, enhancing latent heat absorption and heat dissipation through a piezoelectric microbubble generator. When the temperature drops below 55°C, the system switches to laminar flow mode, resuming liquid circulation and ensuring stable operation.
[0055] A blockchain-based cooling strategy update log is established to record each node's local model parameters and global aggregated results. Sharded storage technology is used to encrypt and distribute log data across server nodes by timestamp. Smart contracts are used to verify cross-node consistency, ensuring traceability and tamper resistance of control instructions, improving system security and reliability.
[0056] The generated control instructions are sent to the magnetic levitation centrifugal pump, proportional valve, and piezoelectric actuator via the CAN bus, achieving closed-loop feedback control of the liquid cooling system. The control instructions are continuously adjusted based on the actual operating status of the system, ensuring that the liquid cooling system is always in optimal operating condition.
[0057] The present invention will be further described below in conjunction with Examples 1 to 5:
[0058] Example 1:
[0059] In liquid-cooled servers, the CPU, GPU, and NPU chips are the main heat-generating components, and their temperature distribution is crucial to the server's performance and stability. Therefore, a micro-sensor array is deployed on the surface of these chips with a pitch of 0.5mm. This tight spacing allows for detailed capture of temperature differences at different locations on the chip surface, acquiring highly accurate temperature distribution data. For example, on the surface of a CPU chip in a high-performance server, a micro-sensor array deployed with this spacing can clearly distinguish temperature changes between the core and surrounding areas, accurately detecting even tiny temperature fluctuations.
[0060] Furthermore, considering that PCBs contain other hotspots besides chips, these areas may also affect the overall thermal balance of the system. To comprehensively monitor the temperature of the PCB board, a non-uniform topology coverage is implemented in these PCB hotspots. The sensor density is adjusted specifically based on the heating characteristics and potential temperature trends of these hotspots. The number of sensors deployed is increased in areas with concentrated heat generation and large temperature fluctuations, while the number of sensors is appropriately reduced in areas with relatively stable temperatures. This non-uniform topology coverage approach not only meets the need for focused monitoring of key hotspots, but also ensures accurate monitoring while rationally controlling the number of sensors used and reducing costs. This deployment of infrared thermal imaging sensors enables real-time, comprehensive, and accurate acquisition of temperature distribution data on the server chip surface and PCB hotspots, laying a solid foundation for the subsequent construction of an accurate three-dimensional thermal field digital twin model.
[0061] Example 2:
[0062] In the liquid-cooled server safety management system of the present invention, a federated learning framework with a horizontal federated architecture is adopted. In this architecture, each server node has the same data feature space, but different data samples. Each server node uses local thermal data to train a thermal dynamic prediction model based on the long short-term memory network (LSTM) algorithm. For example, in a data center containing multiple server nodes, each node performs model training based on its own recorded data on the change in server chip temperature over time to learn the thermal dynamic change patterns of the node server.
[0063] However, in order to improve the accuracy and generalization ability of the model, the models of each node need to be fused. The central aggregator plays a key role in this process. It fuses the gradients of the model of each node through a dynamic weighted average algorithm. When calculating the weight, two important factors are considered: the quality of the node data and the historical prediction error. For nodes with higher data quality, their data can more accurately reflect the actual situation, so they are given a higher weight; and for nodes with smaller historical prediction errors, it means that their models have stronger prediction capabilities, and they are also given a higher weight. The specific weight calculation method is: let the data quality score of node i be Q i , the historical forecast error is E i , then the weight w of node i i The calculation formula is Where n is the total number of nodes participating in federated learning. This dynamic weighted averaging algorithm fully leverages the strengths of each node, making the fused global model more accurate and stable. In practical applications, after multiple rounds of model training and fusion, the prediction accuracy of the thermal dynamic prediction model has significantly improved compared to the single-node model, effectively enhancing the control precision of the liquid cooling system.
[0064] Example 3:
[0065] This example introduces in detail the application of the quantum-derived simulated annealing algorithm in the optimization of coolant flow distribution. By introducing the quantum tunneling effect and adjusting relevant parameters, it can effectively prevent the algorithm from falling into the local optimal solution, more quickly find the global optimal flow distribution solution, and improve the heat dissipation efficiency and energy utilization efficiency of the liquid cooling system.
[0066] A quantum-derived simulated annealing algorithm is used to optimize coolant flow distribution in liquid-cooled servers. The core of this algorithm is the introduction of the quantum tunneling effect, which allows the algorithm to escape local optimality during the search for the optimal solution, increasing the likelihood of finding the global optimal solution.
[0067] First, define the objective function as minimizing the weighted sum of thermal field nonuniformity and pump power consumption. Let thermal field nonuniformity be U, pump power consumption be P, and weighted coefficients be λ1 and λ2 respectively, then the objective function F = λ1U + λ2P. Thermal field nonuniformity U can be measured by calculating the sum of squares of the deviations between the temperature of each point in the thermal field and the average temperature, that is, Where T i is the temperature of the i-th point in the thermal field, is the average temperature of the thermal field, and m is the number of points in the thermal field. The pump power consumption P can be calculated based on the pump's operating parameters, such as flow rate, head, efficiency, etc. The specific calculation formula is determined according to the pump model and operating characteristics.
[0068] During algorithm execution, the algorithm's search process is controlled by adjusting the annealing rate and quantum fluctuation parameters. The annealing rate determines how quickly the probability of accepting a poor solution changes over time during the search. A slower annealing rate allows the algorithm to more fully explore the solution space, but converges more slowly. A faster annealing rate accelerates convergence but may miss the global optimal solution. The quantum fluctuation parameter influences the strength of the quantum tunneling effect. A larger quantum fluctuation parameter enhances the quantum tunneling effect, making it easier for the algorithm to escape from local optimal solutions, but it may also cause the algorithm to be too random during the search. In practical applications, the appropriate annealing rate and quantum fluctuation parameters are determined through multiple trials and optimizations. For example, in a liquid-cooled server system, after a series of tests and adjustments, the annealing rate was set to 0.95 and the quantum fluctuation parameter was set to 0.5. With these parameter settings, the algorithm was able to find the global optimal traffic distribution solution in a relatively short time.
[0069] Embodiment 4:
[0070] This embodiment clarifies the triggering logic of the piezoelectric microbubble generator in the liquid cooling system. By initiating nucleate boiling of the target cold plate in a targeted manner based on the thermal field gradient prediction results, as well as setting a specific electrode arrangement and drive frequency, it can more efficiently enhance latent heat absorption and improve the heat dissipation capacity of the liquid cooling system.
[0071] Piezoelectric microbubble generators play a key role in the phase change cooling process of liquid-cooled servers. Their triggering logic is designed based on thermal field gradient prediction results. By analyzing the thermal field gradient, it is possible to predict which areas will experience a faster temperature rise and may require stronger cooling measures. When the thermal field gradient of a target cold plate area is predicted to exceed a certain threshold, nucleate boiling of that target cold plate is initiated in a targeted manner. For example, in a server's liquid cooling system, thermal field monitoring revealed a sharp increase in the thermal field gradient of the cold plate area corresponding to a certain GPU chip. The system immediately activated the piezoelectric microbubble generator on that cold plate, causing the coolant to undergo nucleate boiling in that area, significantly enhancing latent heat absorption and rapidly reducing the chip temperature.
[0072] The electrodes of the piezoelectric microbubble generator are arranged in a staggered ring array. This arrangement generates a uniform and strong electric field, enabling the coolant to more effectively generate microbubbles. The staggered design avoids electric field concentration and uneven distribution, ensuring sufficient heat dissipation across the entire cold plate area. The drive frequency is set to 1-10kHz. This frequency range ensures effective microbubble generation while avoiding energy loss and equipment damage caused by excessively high frequencies. Lower frequencies may not fully stimulate nucleate boiling of the coolant, while higher frequencies may cause resonance and other issues, impacting system stability. By combining this triggering logic, electrode arrangement, and drive frequency, the piezoelectric microbubble generator activates promptly when needed by the liquid cooling system, precisely enhancing heat dissipation in the target area and ensuring stable operation of server chips in high-temperature environments.
[0073] Example 5:
[0074] This embodiment covers the storage and verification mechanism of blockchain logs, the topological structure of graph convolutional networks, the reward function of the proximal policy optimization algorithm, and the autoencoder model for feature compression. The combination of these technologies helps to improve the security, data analysis capabilities, control decision-making capabilities, and data processing efficiency of the liquid-cooled server safety management system.
[0075] Blockchain Logs and Smart Contract Verification: In the liquid-cooled server security management system, establishing a blockchain-based cooling policy update log is a key measure to ensure system security and traceability. Sharded storage technology is used to encrypt and distribute log data across server nodes by timestamp. Each timestamp corresponds to a log data block containing key information such as the local model parameters and global aggregated results for each node at that moment. For example, during daily server operation, a log data block is generated every hour, recording the model training progress of each node and the global model update results for that hour. This sharded storage approach ensures that even if a node fails or data is tampered with, the integrity of the entire log system is not affected. Furthermore, cross-node consistency verification is implemented using smart contracts. Smart contracts are self-executing contractual clauses deployed on the blockchain. When a node needs to verify the consistency of log data, the smart contract automatically checks the corresponding log data blocks stored by other nodes. Using pre-defined verification algorithms such as hash checksums and digital signature verification, it ensures data consistency across all nodes. If a node's data is found to be inconsistent with that of other nodes, the smart contract issues an alert and takes appropriate remedial measures to ensure the accuracy and reliability of the log data.
[0076] Graph Convolutional Network (GCN) Topology: In the data processing and modeling module, the GCN is used to extract features such as thermal gradient distribution, phase change trigger regions, and flow dead zones. Its topology is dynamically adaptive, automatically adjusting node connection weights based on sensor density. When sensor density is high, indicating that more accurate data analysis is needed, the GCN automatically increases the connection weights between nodes in that area, directing the model's attention to the data features in those areas. Conversely, in areas with lower sensor density, the node connection weights are appropriately reduced. Furthermore, a multi-head attention mechanism is employed to enhance key feature extraction. This mechanism uses multiple attention heads to extract features from different perspectives and then fuses these features. For example, when analyzing thermal field data, one attention head might focus on the changing trend of the thermal gradient, while another might focus on the location of the phase change trigger region. By fusing these key features, the multi-head attention mechanism improves the model's understanding and analysis of the thermal field data, providing a more accurate basis for subsequent decision-making.
[0077] Proximal Policy Optimization Algorithm Reward Function: In the dynamic control instruction generation module, the proximal policy optimization algorithm reward function is defined as R = α·ΔT gradient +β·ΔP pump -γ·ΔS vibration . Among them, ΔT gradient is the thermal field gradient change, which reflects the change in thermal field distribution. The larger the thermal field gradient change, the better the heat dissipation effect of the system and the greater the positive contribution to the reward function; ΔP pump is the change in pump power consumption. A reduction in pump power consumption means an improvement in energy efficiency. Therefore, when this change is negative, it contributes positively to the reward function. vibration α is the vibration spectrum anomaly index. The smaller this index, the more stable the pipeline vibration, the more reliable the system operation, and the greater the positive contribution to the reward function. α, β, and γ are weight coefficients, which are adjusted according to actual needs and system characteristics. For example, in a server scenario with high heat dissipation requirements, the value of α can be appropriately increased to make the algorithm focus more on optimizing the thermal field gradient; in an environment that is more sensitive to energy efficiency, the weight of β can be increased. By properly setting the reward function, the proximal policy optimization algorithm can generate dynamic control instructions that better meet system requirements and achieve precise control of the liquid cooling system.
[0078] Feature Compression Autoencoder Model: In the edge preprocessing module, an autoencoder model is used for feature compression. The input layer of the autoencoder model consists of 128-dimensional raw features. These raw features contain rich information collected by the sensors, but their high dimensionality makes them difficult to transmit and process. The autoencoder model compresses these raw features into a 16-dimensional hidden layer. During this compression process, the autoencoder model learns the important feature representations of the original features and discards some redundant information. For example, when processing temperature, flow rate, and vibration data collected by sensors, the model can automatically extract key features closely related to the operating status of the liquid cooling system, such as temperature trends and abnormal flow rate fluctuations. Furthermore, adversarial training is used to improve feature robustness. Adversarial training introduces adversarial examples during the autoencoder model training process, allowing the model to learn how to resist adversarial attacks and thus improve its robustness to noise and abnormal data. The adversarially trained autoencoder model can more stably output low-dimensional feature vectors in the face of sensor failures or data transmission interference, ensuring the quality of data received by the central control system and providing reliable support for subsequent analysis and decision-making.
[0079] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.
[0080] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A liquid cooling server security management method, characterized in that: The following steps are involved: Step 1: Real-time acquisition of thermodynamic parameters within the liquid cooling system using a multimodal sensor network. This multimodal sensor network includes infrared thermal imaging sensors, ultrasonic flow meters, and piezoelectric vibration sensors. The network acquires server chip surface temperature distribution, coolant flow rate, and pipeline vibration spectrum data at a synchronous frequency of 100 Hz. Step 2: Use the extended Kalman filter algorithm to perform spatiotemporal alignment on the multi-source heterogeneous data collected in step 1, build a 3D thermal field digital twin model, and extract the thermal gradient distribution, phase change trigger area, and flow dead zone characteristics through a graph convolutional network; Step 3: Design a distributed cooling strategy based on the federated learning framework. Each server node locally trains a thermal dynamic prediction model based on the long short-term memory (LSTM) algorithm. A differential privacy mechanism is used to add Gaussian noise to the model gradients. The data is encrypted and uploaded to the central aggregator for global parameter optimization. Step 4: Use the proximal strategy optimization algorithm in reinforcement learning to generate dynamic control instructions. Taking the thermal field gradient and phase change state as input, it outputs the proportional valve opening of each cold plate branch and the triggering timing of the piezoelectric microbubble generator. Step 5: Use the quantum-derived simulated annealing algorithm to optimize the coolant flow distribution. Define the objective function as minimizing the weighted sum of thermal field non-uniformity and pump power consumption, and iterate to find the global optimal flow distribution solution. Step 6: Implement edge preprocessing of sensor data on FPGA hardware, including noise reduction, normalization, and feature compression, to generate low-dimensional feature vectors that are then transmitted to the central control system. Step 7: Design a dual-threshold phase change control mechanism. When the local temperature exceeds 60°C, the phase change coolant is triggered to transform into a gel state and enhance latent heat absorption through a piezoelectric microbubble generator. When the temperature drops below 55°C, the system switches to laminar flow mode to resume liquid circulation. Step 8: Establish a blockchain-based cooling strategy update log to record the local model parameters of each node and the global aggregation results to ensure the traceability and tamper resistance of control instructions; Step 9: Send the control instructions generated in step 4 to the magnetic levitation centrifugal pump, proportional valve and piezoelectric driver through the CAN bus to achieve closed-loop feedback control.
2. The liquid cooling server safety management method according to claim 1, characterized in that: The spatial arrangement rule of the infrared thermal imaging sensors in step 1 is: deploy a micro sensor array with a 0.5mm pitch on the surface of the CPU, GPU, and NPU chips, and use non-uniform topology coverage in the PCB hotspot area.
3. The liquid cooling server safety management method according to claim 1, characterized in that: The federated learning framework described in step 3 adopts a horizontal federated architecture. The central aggregator fuses the model gradients of each node through a dynamic weighted average algorithm, and the weights are updated inversely by the node data quality and historical prediction error.
4. The liquid cooling server safety management method according to claim 1, characterized in that: The quantum-derived simulated annealing algorithm described in step 5 introduces the quantum tunneling effect and avoids local optimal solutions and accelerates convergence by adjusting the annealing rate and quantum fluctuation parameters.
5. The liquid cooling server safety management method according to claim 1, characterized in that: The triggering logic of the piezoelectric microbubble generator in step 7 includes: based on the thermal field gradient prediction result, the nucleate boiling of the target cold plate is directionally started, the electrodes are arranged in a staggered ring array, and the driving frequency is 1-10kHz.
6. The liquid cooling server safety management method according to claim 1, characterized in that: The blockchain log described in step 8 uses shard storage technology to encrypt the log data in blocks according to timestamps and then store them in a distributed manner on each server node, and implement cross-node consistency verification through smart contracts.
7. The liquid cooling server safety management method according to claim 1, characterized in that: The topology of the graph convolutional network described in step 2 is dynamically adaptive, automatically adjusting the node connection weights according to the sensor deployment density, and using a multi-head attention mechanism to enhance key feature extraction.
8. The liquid cooling server safety management method according to claim 1, characterized in that: The reward function of the proximal policy optimization algorithm described in step 4 is defined as: R=α·ΔT gradient +β·ΔP pump -γ·ΔS vibration Where ΔT gradient is the thermal field gradient change, ΔP pump is the change in pump power consumption, ΔS vibration is the vibration spectrum anomaly index, and α, β, and γ are weight coefficients.
9. The liquid cooling server safety management method according to claim 1, characterized in that: The feature compression described in step 6 uses an autoencoder model, with the input layer being 128-dimensional original features, the hidden layer compressed to 16 dimensions, and adversarial training to improve feature robustness.
10. A liquid cooling server safety management system, characterized in that: Includes the following modules: Multimodal data acquisition module: This module uses a multimodal sensor network consisting of infrared thermal imaging sensors, ultrasonic flow meters, and piezoelectric vibration sensors to collect real-time data on server chip surface temperature distribution, coolant flow rate, and pipe vibration spectrum within the liquid cooling system at a synchronous frequency of 100 Hz. Data processing and modeling module: This module uses the extended Kalman filter algorithm to perform spatiotemporal alignment on the collected multi-source heterogeneous data, constructs a three-dimensional thermal field digital twin model, and uses a graph convolutional network to extract thermal gradient distribution, phase change trigger areas, and flow dead zone features. Distributed Cooling Strategy Design Module: Based on a federated learning framework, each server node locally trains a thermal dynamic prediction model based on the Long Short-Term Memory (LSTM) algorithm. A differential privacy mechanism is used to add Gaussian noise to the model gradients, which are then encrypted and uploaded to a central aggregator for global parameter optimization. Dynamic control instruction generation module: This module uses the proximal strategy optimization algorithm in reinforcement learning to generate dynamic control instructions for the proportional valve opening of each cold plate branch and the triggering timing of the piezoelectric microbubble generator, taking the thermal field gradient and phase change state as input; Coolant flow optimization module: This module uses a quantum-derived simulated annealing algorithm to optimize coolant flow distribution. The objective function is defined as minimizing the weighted sum of thermal field non-uniformity and pump power consumption. The global optimal flow distribution solution is solved through iteration. Edge pre-processing module: performs noise reduction, normalization, and feature compression on sensor data on FPGA hardware, generates low-dimensional feature vectors, and transmits them to the central control system; Phase change control module: A dual-threshold phase change control mechanism is designed to trigger the gel state conversion of the phase change coolant when the local temperature exceeds 60°C, and enhance latent heat absorption through a piezoelectric microbubble generator. When the temperature drops below 55°C, it switches to laminar flow mode to resume liquid circulation. Logging module: Establishes a blockchain-based cooling strategy update log, records the local model parameters of each node and the global aggregation results, and ensures the traceability and tamper resistance of control instructions; Instruction issuance and closed-loop control module: The generated control instructions are issued to the magnetic levitation centrifugal pump, proportional valve and piezoelectric driver through the CAN bus to achieve closed-loop feedback control.
Citation Information
Patent Citations
Intelligent water conservancy inspection method, device and equipment and storage medium
CN119106880A
Quantum dot enhanced dynamic power consumption management method and system and related equipment
CN119397979A