Dynamic flow control method for two-phase cold plate liquid cooling system based on multi-modal perception

By using multimodal sensing and big data modeling, multiple signals are collected in real time to construct the comprehensive dryness value of the cold plate and regulate the flow rate. This solves the problems of response lag and inaccurate regulation in the existing two-phase cold plate liquid cooling system, and achieves efficient heat dissipation in high heat density data centers.

CN120909404BActive Publication Date: 2025-12-23TIANJIN TIER TECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511439369.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-10
Publication Date
2025-12-23
Estimated Expiration
2045-10-10

AI Technical Summary

Technical Problem

Existing two-phase cold plate liquid cooling systems mainly rely on a single temperature sensor, which has problems such as response lag, strong localization, susceptibility to interference and signal distortion. This results in inaccurate cooling capacity distribution and adjustment, making it difficult to meet the heat dissipation requirements of high heat density data centers.

Method used

A multimodal sensing method is adopted to collect temperature, pressure, flow rate and microwave attenuation signals in real time. The comprehensive dryness value of the cold plate is constructed by fusing multiple types of physical information. Anomaly detection is carried out by combining signal cross-validation and isolated forest algorithm, and the coolant flow rate is dynamically adjusted. The flow control strategy is optimized by big data modeling and simulation feedback platform.

Benefits of technology

It enhances the sensitivity to complex operating conditions and abnormal risks inside the cold plate, ensures the accuracy and response speed of flow control, and improves the system's self-healing ability and intelligent operation and maintenance level.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120909404B_ABST
    Figure CN120909404B_ABST
Patent Text Reader

Abstract

The application discloses a two-phase cold plate liquid cooling system dynamic flow control method based on multi-modal perception and relates to the technical field of equipment cooling. The method comprises the following steps: S1, real-time acquisition of liquid cooling monitoring data and data preprocessing; S2, determination of cold plate state risk and signal cross-validation, and adoption of abnormal control measures; S3, cooling liquid flow control and evaluation of the matching degree of flow control and chip thermal load, thereby flow control correction; S4, construction of a big data modeling and simulation feedback platform, evaluation of flow control self-healing effect, optimization of parameters and strategies, and self-learning evolution. The method solves the problems that the existing two-phase cold plate liquid cooling system generally only relies on a single temperature signal and lacks a multi-source redundant verification mechanism, has response lag, weak local reflection ability, high risk in case of failure and the like, thereby resulting in insufficient cooling capacity distribution regulation precision and difficulty in meeting the actual demand of high heat density data centers for heat dissipation.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of device cooling, in particular to a two-phase cold plate liquid cooling system dynamic flow control method based on multi-modal perception. BACKGROUND

[0002] With the rapid development of application fields such as high-performance computing, big data centers and cloud servers, the heat flux generated inside the system continues to increase, and the traditional cooling method gradually transforms to the high-efficiency and compact two-phase cold plate liquid cooling technology. The two-phase cold plate liquid cooling system has become one of the mainstream solutions for heat dissipation of key devices such as data centers, servers and power chips, due to its high heat flux heat exchange capacity and excellent temperature control performance.

[0003] For example, the invention patent with publication number CN117979662A discloses a two-phase cold plate liquid cooling system and a control method. According to the invention, for the components cooled by the two-phase cold plate liquid cooling system in series, after the first evaporator cools the previous component, the dryness of the refrigerant output from the first evaporator is adjusted according to the heat transfer requirement of the component cooled by the second evaporator, to avoid the refrigerant dryness being too low or too high. In this way, the refrigerant entering the second evaporator has a suitable dryness, i.e. a suitable boiling point and heat exchange capacity, so that the corresponding component can be cooled to a suitable temperature, avoiding a large temperature difference between the front and rear components. Through this scheme, the problem of the temperature of the components in the rear part being lower than that of the components in the front part after being connected in series is solved, the temperature uniformity of the multiple components cooled in series is improved, and the problems of internal system pressure oscillation, circulation stagnation and circulation backflow are solved, greatly improving the temperature uniformity of the series heat sources and the reliability of the series-parallel system.

[0004] For example, the invention patent with publication number CN118250982A discloses a two-phase cold plate liquid cooling system and method. The system includes a data acquisition module, a control module, a main branch and at least one parallel branch. The main branch includes a condenser and a main circulation module, and the parallel branch includes a liquid cooling module and a target heat dissipation device. The condenser is used for heat exchange with the external environment. The main circulation module includes a first constant flow pump and a liquid reservoir, and is used to drive the cooling liquid to flow in the system. The liquid cooling module includes a second constant flow pump and a preheater, and is used to adjust the flow of the parallel branch. The data acquisition module is used to acquire data information of the main circulation module and the liquid cooling module in each parallel branch. The control module is used to determine a control signal for adjusting the main circulation module and / or the liquid cooling module according to the data information, to dissipate heat for the target heat dissipation device in the parallel branch. The present application can improve the reliability and safety of the two-phase cold plate liquid cooling system.

[0005] However, in the process of implementing the technical scheme of the embodiments of the present application, the present application finds that the above-mentioned technology at least has the following technical problems:

[0006] Most of the existing two-phase cold plate liquid cooling systems mainly rely on single temperature detection, which has problems such as response lag, strong locality, easy to be disturbed and signal distortion, and it is difficult to accurately reflect the overall load of the server and the actual heat dissipation capacity of the working medium in time, and lacks a redundant verification mechanism, so the risk of failure is high, which leads to inaccurate cold energy distribution adjustment and cannot meet the needs of high heat density data centers.

[0007] Therefore, in view of the above problems, there is an urgent need for a two-phase cold plate liquid cooling system dynamic flow control method based on multi-modal perception. SUMMARY

[0008] Technical problems solved

[0009] In view of the deficiencies of the prior art, the present application provides a two-phase cold plate liquid cooling system dynamic flow control method based on multi-modal perception, which solves the problem that the existing two-phase cold plate liquid cooling system generally only relies on single temperature signal and lacks multi-source redundant verification mechanism, has problems such as response lag, weak local reflection ability, high risk when failure, etc., which leads to insufficient cold energy distribution adjustment accuracy and difficulty in meeting the actual needs of high heat density data centers for heat dissipation.

[0010] Technical scheme

[0011] To achieve the above purpose, the present application is realized by the following technical scheme: a two-phase cold plate liquid cooling system dynamic flow control method based on multi-modal perception, comprising the following steps: S1, real-time acquisition of liquid cooling monitoring data, data preprocessing of the liquid cooling monitoring data; S2, based on the preprocessed liquid cooling monitoring data, the cold plate state risk is judged, and signal cross verification is carried out, according to the cold plate state risk discrimination result, abnormal control measures are taken; S3, according to the cold plate state risk discrimination result, cooling liquid flow control is carried out, and according to the preprocessed liquid cooling monitoring data, the matching degree of flow control and chip heat load is evaluated, based on the matching degree evaluation result of heat load, flow control correction is carried out; S4, by analyzing the preprocessed liquid cooling monitoring data, the cold plate state risk discrimination result and the matching degree evaluation result of heat load, a big data modeling and simulation feedback platform is constructed, and the flow control self-healing effect is evaluated, the parameters and strategies are optimized, and self-learning evolution is carried out.

[0012] Further, the specific process of collecting liquid cooling monitoring data in real time and pre-processing the liquid cooling monitoring data is as follows: temperature sensors, pressure sensors, microwave sensors and flow sensors are arranged at the inlet and outlet of the cold plate and at key positions of the server; when the cold plate is first filled with liquid, the microwave attenuation value collected by the microwave sensor is recorded as the microwave reference attenuation, and the difference between the inlet and outlet temperatures of the cold plate is recorded to obtain the temperature difference reference value, and the specific heat capacity constant of the cooling liquid is obtained according to the type of the cooling liquid; the power supply current and voltage are read through the intelligent platform management interface of the server mainboard to calculate the input power of the chip in real time; the liquid cooling monitoring data is collected in real time, including the microwave reference attenuation, the temperature difference reference value, the specific heat capacity constant of the cooling liquid, the inlet temperature of the cold plate, the outlet temperature of the cold plate, the inlet pressure of the cold plate, the outlet pressure of the cold plate, the microwave attenuation value, the flow of the cooling liquid, the temperature of the chip and the input power of the chip; the signal amplitude and fluctuation trend of each sensor are checked to identify abnormal and outlier sampling points and eliminate them; the main and standby sensors are deployed at the flow, temperature and pressure measuring points, the main sensor is switched to the standby sensor when an abnormality is collected, if multiple sensors fail simultaneously, the adjacent nodes and historical data are used to generate virtual sensor data compensation by using the multiple linear regression algorithm; the collected liquid cooling monitoring data is subjected to low-pass filtering and moving average processing to remove working condition noise and environmental interference; the microwave attenuation value is converted into a standard physical quantity through interpolation and calibration algorithm; all the liquid cooling monitoring data are time-aligned and normalized by the minimum and maximum normalization method; the liquid cooling monitoring database is constructed, and the liquid cooling monitoring data is stored in the liquid cooling monitoring database.

[0013] Further, based on the pre-processed liquid cooling monitoring data, the specific process of judging the risk of the cold plate state is as follows: the liquid cooling monitoring data is obtained, the inlet temperature of the cold plate is subtracted from the outlet temperature of the cold plate to obtain the inlet and outlet temperature difference of the cold plate; the current cold plate outlet pressure is divided by the current cold plate inlet pressure to obtain the flow pressure drop ratio, the current microwave attenuation value is divided by the microwave reference attenuation to obtain the microwave dryness ratio, and the current cold plate inlet and outlet temperature difference is divided by the temperature difference reference value to obtain the temperature difference normalization ratio; the flow pressure drop ratio, the microwave dryness ratio and the temperature difference normalization ratio are multiplied, and a constant one is added to perform natural logarithm operation to obtain the cold plate state comprehensive dryness value.

[0014] Further, the specific process of signal cross-validation is as follows: cross-validation is performed on the temperature signal, pressure signal and microwave attenuation signal collected by the sensor: for each signal sample at a time, the historical mean of the current signals is calculated, and the relative change amplitude between each type of signal and its historical mean is calculated by comparison; then the relative change amplitudes of the signals are cross-compared, if the deviation between the normalized value of the relative change amplitude of one signal and the remaining signals exceeds the deviation threshold, it is determined that the signal has drift; and through the isolation forest algorithm, the time series fluctuation and overall correlation of each signal are comprehensively analyzed to identify the drift signal most relevant to the current anomaly, and temporarily reduce the weight of the drift signal in the cold plate state comprehensive dryness value.

[0015] Further, according to the cold plate state risk discrimination result, the specific process of taking abnormal control measures is as follows: the cold plate state comprehensive dryness value is calculated in real time, and the cold plate state comprehensive dryness value is compared with the risk threshold value, when the cold plate state comprehensive dryness value is less than the risk threshold value, it is determined that the cold plate is running in the safe interval; the current flow control strategy is maintained, and the operating parameters do not need to be adjusted, only the monitoring and data acquisition are performed at the normal frequency; when the cold plate state comprehensive dryness value is greater than or equal to the risk threshold value, it is determined that there is a potential flow deterioration and heat exchange performance decline trend; the cooling flow is increased, the short-time high-flow pulse is started, and the flow ratio of each sub-cold plate is adjusted; and the liquid cooling monitoring data sampling and monitoring frequency are increased; the flow, pressure, temperature and microwave attenuation value are detected in real time to determine whether the cold plate is locally blocked, the sensor is failed or the overall is out of adjustment, and temporary isolation is performed on the local abnormal flow channel, while the sensor signal is checked; the current working condition and the corresponding cold plate state comprehensive dryness value are written into the liquid cooling monitoring database.

[0016] Further, according to the cold plate state risk discrimination result, the cooling liquid flow is regulated, and according to the pretreated liquid cooling monitoring data, the specific process of evaluating the matching degree of the flow regulation and the chip thermal load is as follows: real-time regulation of the cooling liquid flow is performed, a sliding time window is set, and a time step is divided, and in the flow regulation process, the liquid cooling monitoring data is acquired in real time; the difference between the cold plate outlet temperature and the cold plate inlet temperature is multiplied by the cooling liquid flow and the specific heat constant of the cooling liquid to obtain the cold plate outlet heat flow, the derivative of the cold plate outlet heat flow with respect to time is calculated at each time step by using the multi-point difference method to obtain the cold plate outlet heat flow rate; the derivative of the chip temperature with respect to time is calculated at each time step to obtain the chip temperature rate; the difference between the chip temperature and the cold plate inlet temperature is divided by the chip input power to obtain the thermal resistance; the chip equivalent heat flow rate is obtained by multiplying the chip temperature rate by the inverse of the thermal resistance; the heat flow response error is obtained by subtracting the chip equivalent heat flow rate from the cold plate outlet heat flow rate and taking the absolute value; the heat flow response error cumulative amount is obtained by integrating and accumulating the heat flow response error of all time steps in the sliding time window; at the same time, the cold plate outlet heat flow rate at all time steps in the sliding time window is integrated and accumulated, and a minimum constant value is added to obtain the heat flow change reference amount; the cold plate load response matching value is obtained by dividing the response error cumulative amount by the heat flow change reference amount.

[0017] Further, based on the matching degree evaluation result of the thermal load, the specific process of flow regulation correction is as follows: the cold plate load response matching value is compared in real time with the load matching threshold value, when the cold plate load response matching value is less than or equal to the load matching threshold value, it is determined that the flow regulation matches well with the change of the chip load, the current flow regulation strategy is maintained, and adjustment is not needed, and only routine frequency continuous monitoring is performed; when the cold plate load response matching value is greater than the load matching threshold value, it is determined that the cold plate regulation response lags behind and does not match the load change, the cooling liquid flow is redistributed, and the rotation speed of the fluid pump is adjusted; based on the thermal load distribution, local abnormal nodes are identified, each cold plate flow is adjusted, and partitioned flow accurate compensation is performed, and monitoring and sampling frequency is preferentially increased; when it is monitored that the chip temperature is not improved after the flow regulation, a chip frequency reduction and dynamic power limit request is sent; a redundant flow channel is designed, when it is detected that the local blockage of the main flow channel lasts and does not recover, part of the cooling liquid flow is switched to the standby channel, and the abnormal channel is isolated; at the same time, the cold plate load response matching value, the flow regulation strategy and the corresponding working condition are written into the liquid cooling monitoring database, and the cold plate load response matching value parameter and the flow regulation strategy are optimized according to the continuous monitoring result until the cold plate load response matching value is less than or equal to the load matching threshold value.

[0018] Further, the specific process of constructing the big data modeling and simulation feedback platform by analyzing the pre-processed liquid cooling monitoring data, the cold plate state risk discrimination result and the matching degree evaluation result of the thermal load is as follows: the liquid cooling monitoring data, the corresponding cold plate state comprehensive dryness value and the cold plate load response matching value are combined to construct a multi-source simulation feature set, the multi-source simulation feature set is uploaded to the natural language big model for data simulation, and a plurality of abnormal simulation scenarios are constructed according to the historical and real-time multi-source simulation feature set; through the natural language big model, combined with the current environmental parameters including external temperature and humidity, power state, cooling liquid type and multi-load working condition configuration, the actual response process under various abnormal simulation scenarios is simulated and deduced; in the simulation process, the natural language big model quantitatively evaluates the occurrence frequency, time sequence relationship and propagation path of various abnormal liquid cooling monitoring data, and based on the Bayesian network algorithm, infers the causal relationship between variables and identifies the abnormal source path.

[0019] Further, the specific process of evaluating the flow regulation self-healing effect is as follows: at the same time, the actual monitored signal type number and the corresponding liquid cooling monitoring data including the cold plate outlet temperature, the cold plate outlet pressure and the microwave attenuation value are obtained; for each abnormal simulation scenario, based on the sliding time window, the mean value of the ith signal before the flow regulation execution is calculated, and the mean value of the ith signal after the flow regulation execution is calculated; the mean value of the ith signal after the flow regulation execution is subtracted from the mean value of the ith signal before the flow regulation execution, and the absolute value is obtained to obtain the self-healing change amplitude of the ith signal; the allowable fluctuation threshold of each signal is set, the self-healing change amplitude of the ith signal is subtracted from the allowable fluctuation threshold of the ith signal to obtain the over-allowed recovery deviation of the ith signal, the maximum function operation of the over-allowed recovery deviation of the ith signal is performed, that is, the larger one between the over-allowed recovery deviation of the ith signal and 0 is taken, if the over-allowed recovery deviation of the ith signal is positive, the actual calculation result is retained, otherwise 0 is taken; based on the signal type number, the over-allowed recovery deviations of all signals are accumulated and averaged to obtain the average over-standard recovery difference value, and the liquid cooling self-healing ability evaluation value is obtained by subtracting the constant one from the average over-standard recovery difference value.

[0020] Further, the specific process of optimizing each parameter and strategy and self-learning evolution is as follows: the liquid cooling self-healing ability evaluation value is calculated in real time as a core health index for feedback; if the liquid cooling self-healing ability evaluation value is greater than or equal to the self-healing threshold, it is considered that the current flow regulation strategy is effective, and the existing regulation is maintained; if the liquid cooling self-healing ability evaluation value is less than the self-healing threshold, the monitoring signal with the super-permitted recovery deviation greater than the recovery deviation threshold is identified, and the weight is reduced; the optimization regulation measure is triggered according to the hierarchical response mechanism: the cooling flow of the recovery poor cold plate is mainly improved; the sensitivity, threshold and PID parameter of the flow regulation are adjusted to improve the recovery effect; the short-time pulse flow flushing is started to improve the recovery ability of the local signal; the liquid cooling self-healing ability evaluation value, the optimization regulation measure and the optimization effect are written into the liquid cooling monitoring database, the flow regulation strategy and the liquid cooling self-healing ability evaluation value parameter are iteratively optimized, self-adaptive evolution is carried out, and the multi-signal self-healing ability is improved.

[0021] Beneficial effects

[0022] The present application has the following beneficial effects:

[0023] (1) The present application collects temperature, pressure, flow and microwave attenuation multi-source signals in real time through multi-modal, fuses multiple types of physical information, constructs a cold plate state comprehensive dryness value, improves the sensitivity and identification ability to complex working conditions, local dryness and abnormal risks in the cold plate, and effectively overcomes the traditional shortcomings of single signal being easily disturbed and response lag.

[0024] (2) The present application can identify signal drift and failure in real time through the signal cross-validation and intelligent anomaly detection mechanism of the isolated forest algorithm, reduce the weight of abnormal signals, improve the data robustness of the cold plate state comprehensive dryness value, and ensure that the flow regulation and health evaluation are always based on a high credibility data basis.

[0025] (3) The present application realizes dynamic and partitioned accurate regulation of the cold plate flow by constructing a load response matching index, can correct the flow regulation parameter according to the actual heat load and abnormal condition, and improves the response speed and regulation precision to complex working conditions of high heat density and sudden load change.

[0026] (4) The present application continuously evaluates the self-healing effect of the flow regulation by integrating a big data modeling and simulation feedback platform, archives and analyzes historical optimization cases, realizes online self-learning iteration of parameters and strategies, and effectively improves the self-healing ability, long-term stability and intelligent operation and maintenance level of the liquid cooling system.

[0027] Of course, any product implementing the present application does not necessarily need to achieve all the advantages described above at the same time. BRIEF DESCRIPTION OF DRAWINGS

[0028] Figure 1A flow chart of a dynamic flow control method for a two-phase cold plate liquid cooling system based on multi-modal perception;

[0029] Figure 2 A schematic diagram of a two-phase cold plate liquid cooling dynamic flow control principle based on multi-modal perception;

[0030] Figure 3 An internal structure diagram of multi-modal perception monitoring of a two-phase cold plate;

[0031] Figure 4 A trend chart of a cold plate state comprehensive dryness value.

[0032] In the figure, 1 is a heat exchange device, 2 is a fluid pump, 3 is a server, 4 is a pressure sensor, 5 is a temperature sensor, 6 is a microwave sensor, 7 is a two-phase cold plate, and 8 is a chip. DETAILED DESCRIPTION

[0033] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present application. As understood by those skilled in the art, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0034] Please refer to Figures 1-4 The embodiments of the present application provide a technical solution: a dynamic flow control method for a two-phase cold plate liquid cooling system based on multi-modal perception, as shown in Figure 1 The method comprises the following steps: S1, real-time acquisition of liquid cooling monitoring data, data preprocessing of the liquid cooling monitoring data; S2, based on the preprocessed liquid cooling monitoring data, determination of a cold plate state risk, signal cross-validation, and according to the cold plate state risk determination result, taking abnormal control measures; S3, according to the cold plate state risk determination result, cooling liquid flow control, and according to the preprocessed liquid cooling monitoring data, evaluation of the matching degree of flow control and chip 8 thermal load, based on the matching degree evaluation result of the thermal load, flow control correction; S4, through analysis of the preprocessed liquid cooling monitoring data, the cold plate state risk determination result and the matching degree evaluation result of the thermal load, a big data modeling and simulation feedback platform is constructed, and the self-healing effect of flow control is evaluated, the parameters and strategies are optimized, and self-learning evolution is performed.

[0035] Specifically, the liquid cooling monitoring data is collected in real time, and the specific process of data preprocessing of the liquid cooling monitoring data is as follows: temperature sensors 5, pressure sensors 4, microwave sensors 6 and flow sensors are arranged at the inlet and outlet of the cold plate and the server 3 key positions, which are respectively used to collect the key thermal and fluid parameters of the cold plate and the pipeline, so as to realize all-round perception of the thermal load, flow state and phase distribution; when the cold plate is first filled with liquid, the microwave attenuation value collected by the microwave sensor 6 is recorded as the microwave reference attenuation, wherein the microwave sensor 6 detects the vapor-liquid ratio in the cold plate based on the electromagnetic characteristics of the working medium, and the microwave reference attenuation reflects the initial state when the cold plate is filled with liquid; and the difference between the inlet and outlet temperatures of the cold plate is recorded to obtain the temperature difference reference value, and the specific heat capacity constant of the cooling liquid is obtained according to the type of the cooling liquid; and the power supply current and voltage are read through the intelligent platform management interface of the server 3 mainboard to calculate the chip input power in real time; the liquid cooling monitoring data is collected in real time, including: microwave reference attenuation, temperature difference reference value, specific heat capacity constant of cooling liquid, cold plate inlet temperature, cold plate outlet temperature, cold plate inlet pressure, cold plate outlet pressure, microwave attenuation value, cooling liquid flow, chip temperature and chip input power; the signal amplitude and fluctuation trend of each sensor are checked, the abnormal and outlier sampling points are identified and removed, through statistical analysis and abnormal detection of the collected signals, the abnormal data caused by sensor failure, external interference and communication abnormality can be excluded, and the accuracy of subsequent data processing is improved; the master and standby sensors are deployed at the flow, temperature and pressure measuring points, the master sensor is switched to the standby sensor when abnormal data is collected, the redundancy check and fault self-recovery at the hardware level are realized, if multiple sensors fail at the same time, the adjacent nodes and historical data are used, and the virtual sensor data compensation is generated by using the multiple linear regression algorithm, that is, the method based on spatial correlation and time series regression, the adjacent measuring points and historical working condition data are used to fit and generate the missing or failed signals, so as to ensure the continuity and integrity of the data link; the collected liquid cooling monitoring data is subjected to low-pass filtering and moving average processing to remove working condition noise and environmental interference, the low-pass filtering can filter out high-frequency noise, and the moving average further smooths the signal to improve the data quality; the microwave attenuation value is converted into standard physical quantities by interpolation and calibration algorithm, including amplitude correction, frequency calibration and mapping with known vapor-liquid working conditions of the original signal; all the liquid cooling monitoring data is time-aligned, that is, the data collected at different sampling frequencies and different steps are unified to the same time stamp, and the minimum maximum normalization method is used for normalization processing; a liquid cooling monitoring database is constructed, and the liquid cooling monitoring data is stored in the liquid cooling monitoring database.

[0036] As Figure 2The diagram illustrates the principle of dynamic flow control for two-phase cold plate liquid cooling based on multimodal sensing. It showcases the core working principle and main component structure of the dynamic flow control method. The system primarily consists of heat exchanger 1, fluid pump 2, server 3, pressure sensor 4, temperature sensor 5, and microwave sensor 6, forming a closed-loop liquid cooling control circuit for high heat density server heat dissipation. In actual operation, fluid pump 2 drives the coolant circulation. The coolant first undergoes preliminary cooling via heat exchanger 1, then flows through the two-phase cold plate 7 within server 3, efficiently exchanging heat with the chip 8 and key heat-generating components. Pressure sensor 4, temperature sensor 5, and microwave sensor 6 are respectively positioned at the cold plate inlet / outlet and key pipeline nodes to collect real-time fluid pressure, temperature, and microwave attenuation data. All sensor signals undergo collaborative acquisition and preprocessing. The control unit performs multimodal signal fusion to calculate the comprehensive dryness value of the cold plate state and implements cross-validation, achieving comprehensive perception of the cold plate's operating condition and anomaly risk identification. Based on the calculation results, the speed and flow rate of fluid pump 2 are dynamically adjusted to achieve precise distribution and adaptive control of coolant flow, timely response to changes in server 3 load and potential anomalies, and ensure efficient and reliable heat dissipation of the cold plate in high heat density scenarios.

[0037] like Figure 3 The diagram shows the internal structure of the two-phase cold plate multimodal sensing and monitoring system. Pressure sensor 4, installed on the liquid-cooled flow channel, collects real-time inlet and outlet pressures of the cold plate, reflecting flow resistance, blockages, and abnormal fluid distribution. It is one of the core physical quantities for judging flow anomalies and changes in zoned resistance. Temperature sensor 5 is located at key nodes in the liquid-cooled pipeline to detect temperature changes at the inlet and outlet of the cold plate, accurately monitoring heat exchange efficiency and cold plate operating conditions. This is a crucial foundation for dynamic load tracking and energy efficiency analysis. Microwave sensor 6, based on the principle of microwave signal attenuation, is installed on the cold plate pipeline. By utilizing the difference in microwave attenuation during propagation in the vapor-liquid two-phase working fluid, it monitors the vapor-liquid phase ratio, dryness changes, and localized drying risks within the cold plate in real time, enhancing the sensitivity to complex phase changes and abnormal operating conditions. The two-phase cold plate 7 is the core heat dissipation unit, efficiently transferring the waste heat from the chip 8. Through its internal multi-channel structure and multi-point parameter acquisition, it achieves comprehensive monitoring and control of heat distribution and flow state. Chip 8 is a high heat density electronic component that is directly cooled by the two-phase cold plate 7 and is the fundamental load source for thermal management.

[0038] In this embodiment, comprehensive real-time collection and high-quality preprocessing of multi-source physical parameters are achieved, and the continuity and accuracy of monitoring data are effectively guaranteed through collaborative deployment of multiple types of sensors and master-slave redundancy and virtual compensation technology. Signal anomaly detection and multi-level filtering improve the data anti-interference ability, and interpolation calibration and normalization processing ensure the comparability and fusion of different physical quantities of data, ultimately building a high-reliability liquid cooling monitoring database, which provides a solid data foundation for subsequent intelligent criterion analysis, flow regulation and health management functions.

[0039] Specifically, based on the preprocessed liquid cooling monitoring data, the specific process of discriminating the cold plate state risk is as follows: obtaining the liquid cooling monitoring data, subtracting the cold plate outlet temperature from the cold plate inlet temperature to obtain the cold plate inlet and outlet temperature difference, which can directly reflect the heat exchange capacity of the cold plate and the heat change absorbed or released by the fluid passing through the cold plate; dividing the current cold plate outlet pressure by the current cold plate inlet pressure to obtain the flow pressure drop ratio, which reflects the pressure drop degree of the fluid flowing through the cold plate and is a key indicator for measuring the flow resistance and existing blockage, dryness anomaly in the cold plate; dividing the current microwave attenuation value by the microwave reference attenuation to obtain the microwave dryness ratio, which detects the change of vapor-liquid phase content by using the microwave sensor 6, and by comparing with the reference state, the internal phase change distribution and dryness risk of the cold plate can be discriminated; dividing the current cold plate inlet and outlet temperature difference by the temperature difference reference value to obtain the temperature difference normalization ratio, which is used to eliminate the influence of initial conditions under different working conditions, so that the state criterion has good comparability and engineering universality; multiplying the flow pressure drop ratio, the microwave dryness ratio and the temperature difference normalization ratio, the three indicators cooperatively reflect the comprehensive working condition of the cold plate in flow, phase state and heat exchange, and the constant is added to the natural logarithm operation to obtain the cold plate state comprehensive dryness value, which helps to improve the sensitivity of the criterion to abnormal state and weaken the influence of extreme value. The cold plate state comprehensive dryness value can be used as a core risk indicator to represent the internal partial dryness, flow deterioration and heat exchange performance change of the cold plate.

[0040] Among them, the specific formula of the cold plate state comprehensive dryness value is:

[0041] ;

[0042] In the formula, represents the cold plate state comprehensive dryness value at the current time, which realizes the comprehensive quantitative discrimination of the cold plate dryness, flow anomaly and heat exchange deterioration state by coupling the three core physical links of flow, structure and heat exchange, and using logarithmic amplification to enhance the sensitivity to mutation; the larger the cold plate state comprehensive dryness value, the higher the internal dryness of the cold plate and the greater the abnormal risk, which is an important indicator for flow regulation, alarm and early warning, and health assessment; represents the cold plate outlet pressure, reflecting the pressure of the cooling liquid when it leaves the cold plate; represents the cold plate inlet pressure, reflecting the pressure of the cooling liquid when it enters the cold plate; represents the current microwave attenuation value, reflecting the real-time attenuation value of the microwave signal passing through the flow channel under the current working condition of the cold plate; represents the microwave reference attenuation, reflecting the measured microwave signal reference attenuation value when the cold plate is full of liquid; represents the current cold plate inlet and outlet temperature difference, indicating the temperature difference between the two ends of the cold plate, reflecting the actual heat exchange effect; represents the reference value of the temperature difference, reflecting the measured inlet and outlet temperature difference when the cold plate is full of liquid; represents the flow pressure drop ratio, representing the flow resistance and state of the cooling liquid flowing through the cold plate; when the cold plate is dry, blocked, or flow is abnormal, the outlet pressure increases relatively, the ratio increases, the flow resistance increases, and the risk of flow abnormality increases; represents the microwave dryness ratio, reflecting the distribution of fluid vapor and liquid; if the current microwave signal attenuation is greater than the reference value, it indicates that the gas phase proportion increases, the dryness increases, and the risk of cold plate drying increases; represents the temperature difference normalization ratio, reflecting the heat exchange capacity and abnormal change; when bubbles, dryness, and flow deterioration occur, the temperature difference between the two ends of the cold plate increases, the ratio increases, indicating that the heat exchange performance deteriorates and the abnormal risk increases.

[0043] In this embodiment, Table 1 is a cold plate state comprehensive dryness value data table, which records in detail the cold plate outlet pressure, cold plate inlet pressure, microwave attenuation value, microwave reference attenuation, cold plate inlet and outlet temperature difference, temperature difference reference value, and cold plate state comprehensive dryness value at different times. The cold plate inlet pressure is uniformly set to 100.0, the microwave reference attenuation is set to 8.5, and the temperature difference reference value is set to 4.0; wherein, the cold plate outlet pressure corresponding to time number 1 is 103.2, the microwave attenuation value is 8.7, the cold plate inlet and outlet temperature difference is 4.2, and the cold plate state comprehensive dryness value is 0.7465; the cold plate outlet pressure corresponding to time number 2 is 101.0, the microwave attenuation value is 8.5, the cold plate inlet and outlet temperature difference is 4.1, and the cold plate state comprehensive dryness value is 0.7104; the cold plate outlet pressure corresponding to time number 3 is 107.5, the microwave attenuation value is 9.2, the cold plate inlet and outlet temperature difference is 4.8, and the cold plate state comprehensive dryness value is 0.8742; the cold plate outlet pressure corresponding to time number 4 is 109.0, the microwave attenuation value is 9.7, the cold plate inlet and outlet temperature difference is 5.2, and the cold plate state comprehensive dryness value is 0.9616; the cold plate outlet pressure corresponding to time number 5 is 100.5, the microwave attenuation value is 8.4, the cold plate inlet and outlet temperature difference is 4.0, and the cold plate state comprehensive dryness value is 0.6896.

[0044] Table 1 Cold plate state comprehensive dryness value data table

[0045]

[0046] As Figure 4The figure shows the trend chart of the cold plate state comprehensive dryness value. The trend of the change of the cold plate state comprehensive dryness value at five different monitoring times is shown; the horizontal axis is the time number, the vertical axis is the cold plate state comprehensive dryness value, and each point represents the cold plate state comprehensive dryness value at a specific monitoring time. The broken line shows the fluctuation of the dryness value over time. According to Table 1 and Figure 4 It can be seen that the cold plate state comprehensive dryness value reaches a lower point at time 2, then rises rapidly, reaches a peak at time 4, and then reaches a minimum at time 5; the overall trend is first decreased, then greatly increased, and then quickly fallen; especially at time 3 and 4, the cold plate state comprehensive dryness value greatly increases, indicating that the cold plate may have the risk of flow deterioration, local dryout and heat transfer performance decline.

[0047] In this embodiment, through the cooperative calculation of the flow pressure drop ratio, the microwave dryness ratio and the temperature difference normalized ratio, the comprehensive discrimination of the cold plate flow state, the phase distribution and the heat transfer capacity is realized. The normalization and natural logarithm operation are adopted to improve the sensitivity to abnormal risks and the engineering universality of the index, so that the cold plate state comprehensive dryness value can accurately characterize the key risks of local dryout, flow deterioration and heat transfer performance change, and provide a scientific and reliable criterion basis for subsequent flow regulation and abnormal adaptive adjustment.

[0048] Specifically, the specific process of signal cross verification is as follows: the temperature signal, the pressure signal and the microwave attenuation signal collected by the sensor are cross verified: for each signal sample at each time, the historical mean value of the current signal is calculated, the historical mean value can be dynamically updated in real time based on the sliding window algorithm, which reflects the baseline level of the signal under normal operating conditions, and the relative change amplitude between each type of signal and its historical mean value is calculated by comparison, the relative change amplitude is used to measure the deviation degree of the current fluctuation of the signal from its normal state, the calculation formula is to take the absolute value of the current sampling value minus the historical mean value, and normalize it with the historical mean value; then the relative change amplitudes of the signals are cross compared, that is, the relative change amplitude of each signal is compared with the average change amplitude of the other signals to identify the abnormal state of each signal, if the deviation between the normalized value of the relative change amplitude of one signal and the average change amplitude of the other signals exceeds the deviation threshold, it is determined that the signal has drift, the deviation threshold can be set to a reasonable tolerance range according to statistical analysis; and through the isolation forest algorithm, the time sequence fluctuation and overall correlation of each signal are comprehensively analyzed to identify the drift signal most related to the current abnormality, wherein the isolation forest is an unsupervised anomaly detection algorithm based on random splitting and ensemble learning, which is suitable for automatic discovery of abnormal points in multi-dimensional data. Through the isolation forest, the sampling points with the most abnormal characteristics in the multi-dimensional signal can be mined, the discrimination accuracy is improved, and the weight of the drift signal in the cold plate state comprehensive dryness value is temporarily reduced. The dynamic weight adjustment mechanism can reduce the interference of the abnormal signal on the cold plate state comprehensive dryness value, and enhance the stability and reliability of the cold plate state risk discrimination.

[0049] In this embodiment, by dynamic cross-validation and anomaly detection of temperature, pressure and microwave attenuation multi-class signals, signal drift and failure can be identified in time, and combined with isolated forest algorithm to identify and weight abnormal signals, which effectively improves the robustness of data fusion and the reliability of the criterion. The dynamic weight adjustment mechanism further improves the fault tolerance and adaptability to multi-source signal anomalies, ensuring the accurate perception of the cold plate state comprehensive dryness value criterion to the working condition risk and stable output, providing high credibility data support for subsequent flow regulation and fault self-healing.

[0050] Specifically, according to the cold plate state risk discrimination result, the specific process of taking abnormal regulation measures is: the cold plate state comprehensive dryness value is calculated in real time, and the cold plate state comprehensive dryness value is compared with the risk threshold value, the risk threshold value can be dynamically set through historical working condition statistical analysis, which is used to distinguish the safe and risk running interval of the cold plate; when the cold plate state comprehensive dryness value is less than the risk threshold value, it is determined that the cold plate is running in the safe interval; maintain the current flow regulation strategy, do not need to adjust the operating parameters, only monitor and data acquisition at the normal frequency, ensure efficient use of resources, reduce energy consumption and equipment wear and tear; when the cold plate state comprehensive dryness value is greater than or equal to the risk threshold value, it is determined that there is a potential flow deterioration and heat exchange performance decline trend, and the active risk response state is entered; improve the cooling flow, start the short-time high-flow pulse, the short-time high-flow pulse refers to temporarily increasing the flow of cooling liquid through the fluid pump 2, so as to flush the internal bubbles, deposited impurities and micro-blockage of the cold plate, and quickly improve the local flow condition; adjust the flow ratio of each sub-cold plate, adjust the flow of each loop through intelligent valve, accurately match the local load and abnormal distribution; and improve the liquid cooling monitoring data sampling and monitoring frequency, increasing the sampling frequency helps to capture abnormal fluctuations and subtle changes in the regulation effect, and enhances the response sensitivity to abnormal state; real-time detection of flow, pressure, temperature and microwave attenuation value, to judge whether it is cold plate local blockage, sensor failure or global imbalance, through multi-parameter joint analysis, the type of abnormal source can be identified, the local working condition abnormality and global abnormality can be effectively distinguished, and temporary isolation of local abnormal flow channel is implemented, temporary isolation refers to closing part of the loop to prevent abnormal spread and affect the overall operation, at the same time, through signal redundancy, data verification and virtual compensation technology, the sensor signal is verified; the current working condition and the corresponding cold plate state comprehensive dryness value are written into the liquid cooling monitoring database.

[0051] In this embodiment, through the dynamic threshold discrimination based on the cold plate state comprehensive dryness value and the multi-level response mechanism, the rapid identification and intelligent control of the cold plate flow deterioration and heat transfer performance decline are realized. The adaptive flow rate increase, pulse flushing, precise adjustment of partition flow, combined with multi-parameter real-time monitoring, abnormal type discrimination and temporary isolation measures effectively improve the risk response ability and local fault self-healing ability, ensuring the continuous and efficient and safe operation of the cold plate under complex working conditions.

[0052] Specifically, according to the risk discrimination result of the cold plate state, the cooling liquid flow is regulated, and according to the pretreated liquid cooling monitoring data, the specific process of evaluating the matching degree of flow regulation and chip 8 thermal load is as follows: by adjusting the rotation speed of the fluid pump 2 and the opening degree of each partition valve, the cooling liquid flow is dynamically adjusted to realize the adaptive distribution of the heat dissipation capacity under different load working conditions, and the real-time regulation of the cooling liquid flow is carried out. At the same time, a sliding time window is set according to the working condition response requirement and the sampling rate, the range is 30 seconds to 5 minutes, and the time step is divided based on the sampling interval, and the window length and the step can be adaptively optimized according to the actual operation effect; in the flow regulation process, the liquid cooling monitoring data is obtained in real time; the difference between the outlet temperature of the cold plate and the inlet temperature of the cold plate is multiplied by the cooling liquid flow and the specific heat capacity constant of the cooling liquid to obtain the outlet heat flow of the cold plate, and the outlet heat flow of the cold plate is the heat carried away by the cold plate per unit time, and the calculation formula is: ; wherein, is the cooling liquid flow, is the specific heat capacity of the cooling liquid, , respectively, the cold plate outlet temperature and the cold plate inlet temperature; the derivative of the cold plate outlet heat flow with respect to time is calculated by using a multi-point difference method for each time step, to obtain the cold plate outlet heat flow rate of change; the multi-point difference method is a numerical differentiation method, which calculates the derivative of the signal with respect to time by using multiple groups of sampling points, thereby improving the accuracy and noise immunity of the change rate calculation; the derivative of the chip temperature with respect to time is calculated for each time step, to obtain the chip temperature rate of change; the thermal resistance is obtained by dividing the difference between the chip temperature and the cold plate inlet temperature by the chip input power, and the thermal resistance is used to measure the heat conduction capacity between the chip and the cold plate, and reflects the actual heat exchange performance; the chip equivalent heat flow rate of change is obtained by multiplying the chip temperature rate of change by the reciprocal of the thermal resistance; the heat flow response error is obtained by subtracting the chip equivalent heat flow rate of change from the cold plate outlet heat flow rate of change, and taking the absolute value, and the heat flow response error quantifies the dynamic matching degree of the cold plate flow regulation to the chip thermal load response, and the smaller the heat flow response error is, the more accurate the regulation is; the response error cumulative amount is obtained by integrating and accumulating the heat flow response errors of all time steps in the sliding time window; meanwhile, the heat flow change reference amount is obtained by integrating and accumulating the cold plate outlet heat flow rates of change of all time steps in the sliding time window, and adding a minimum constant value, wherein the minimum constant value is 0.001, which is used to prevent the denominator from being zero, and to ensure the stability and non-dimensionalization of the calculation; the cold plate load response matching value is obtained by dividing the response error cumulative amount by the heat flow change reference amount, which is used to determine the dynamic adaptation effect of the flow regulation to the actual thermal load, and is an important core criterion for the self-adaptive optimization and closed-loop regulation of the flow regulation.

[0053] wherein the specific formula of the cold plate load response matching value is:

[0054] ;

[0055] in the formula, represents the cold plate load response matching value, reflects the matching degree of the cold plate flow regulation strategy to the chip thermal load dynamic response, measures the flow regulation effect, and the smaller the cold plate load response matching value is, the more timely and accurate the regulation response is; the increase of the cold plate load response matching value indicates that there is a response lag and an abnormality between the regulation and the thermal load; represents the cold plate outlet heat flow, reflects the heat taken away from the cold plate by the cooling liquid per unit time; represents the cold plate outlet heat flow rate of change, reflects the speed and intensity of the cooling capacity change under the cold plate flow regulation, and is an important quantitative index of the response speed and activity; represents the thermal resistance, reflects the equivalent thermal resistance of all heat conduction paths between the chip and the cooling liquid; represents the chip temperature, reflects the temperature of the target chip cooled by the cold plate; represents the chip equivalent heat flow rate of change, reflects the speed and trend of the current temperature change of the chip, and is a direct embodiment of the dynamic thermal load and the cooling effect change; represents a minimum constant value, taking a value of 0.001, to prevent the denominator from being zero; represents a response error accumulation amount, measuring the cumulative error of dynamic mismatch between the cooling flow output of the cold plate and the actual temperature change of the chip within a period of time, reflecting the amplitude and duration of control lag, control deficiency and control overshoot; the greater the error, the less accurate the flow control covers the changes in the heat load of the chip 8, and the poorer the matching degree; represents a heat flow change reference amount, indicating the total amount of change in the cold plate heat flow output within the same time window, ensuring that the cold plate load response matching value under different working conditions can be directly used as a criterion for control effect.

[0056] In the present embodiment, real-time adaptive control of the cooling liquid flow is achieved, and numerical algorithms such as multi-point difference and sliding window are used to accurately evaluate the dynamic matching degree of the cold plate flow control and the heat load of the chip 8. Not only can the difference between the heat flow and the chip 8 response in the control process be quantified, but also the calculation stability and universality can be ensured through the normalization criterion, thereby improving the intelligence, precision and engineering robustness of the flow control, and providing a solid data and criterion basis for the safe and efficient operation and automatic closed-loop optimization of the liquid cooling system under high heat density scenarios.

[0057] Specifically, based on the matching degree evaluation result of the heat load, the specific process of flow regulation correction is as follows: the cold plate load response matching value is compared with the load matching threshold value in real time. When the cold plate load response matching value is less than or equal to the load matching threshold value, it is determined that the flow regulation matches well with the change of the chip 8 load, the current flow regulation strategy is maintained, and there is no need to adjust, and only the conventional frequency continues to be monitored. When the cold plate load response matching value is greater than the load matching threshold value, it is determined that the cold plate regulation response lags behind and does not match the load change, the cooling liquid flow is redistributed, and the speed of the fluid pump 2 is adjusted. By adjusting the pump speed and the opening degree of each partition valve, the flow is flexibly controlled to quickly respond to the change of the heat load. Based on the heat load distribution, local abnormal nodes are identified, and the flow of each cold plate is adjusted to compensate for the partition flow accurately. The accurate compensation means that the high-heat area and abnormal nodes obtain more cooling resources through independent regulation of the partition flow, so as to improve the overall heat dissipation uniformity and efficiency. The monitoring and sampling frequency is preferentially increased, the encrypted sampling frequency facilitates real-time capture of the rapid change of each signal, and the sensitivity of the criterion and regulation is improved. When it is monitored that the chip temperature does not improve after the flow regulation, the chip 8 frequency reduction and dynamic power limit request is sent out, that is, the power output of the high-heat chip 8 is temporarily limited by the mainboard management to reduce the cooling load and avoid the risk of overheating. A redundant flow channel is designed. When it is detected that the local blockage of the main flow channel continues to be not recovered, part of the cooling liquid flow is switched to the standby channel, and the abnormal channel is isolated. The redundant flow channel means a reserved standby flow path that can share part of the cooling task when the main channel is abnormal, thereby enhancing the fault tolerance and self-healing ability. At the same time, the cold plate load response matching value, the flow regulation strategy and the corresponding working condition are written into the liquid cooling monitoring database to provide data support for subsequent health management, intelligent operation and self-learning. According to the continuous monitoring result, the cold plate load response matching value parameter and the flow regulation strategy are optimized until the cold plate load response matching value is less than or equal to the load matching threshold value, thereby forming a criterion-driven closed-loop adaptive regulation system to continuously improve the dynamic optimization capability and robustness.

[0058] In the embodiment, through the dynamic criterion driving and hierarchical flow correction mechanism, the cooling liquid flow efficiently and adaptively responds to the change of the heat load of the chip 8. The cooling resource can be accurately identified and distributed, and the load fluctuation and abnormal working condition can be actively responded to. Through the redundant channel design, the chip 8 power limit and the frequency optimization multi-strategy cooperation, the regulation precision, the fault self-healing ability and the safety and reliability of the overall operation of the liquid cooling of the cold plate are improved, thereby laying a solid data foundation and closed-loop optimization capability for subsequent health management and intelligent operation.

[0059] Specifically, by analyzing the pre-processed liquid cooling monitoring data, the cold plate state risk discrimination result and the matching degree evaluation result of the thermal load, the specific process of building a big data modeling and simulation feedback platform is as follows: the liquid cooling monitoring data and its corresponding cold plate state comprehensive dryness value and cold plate load response matching value are combined to build a multi-source simulation feature set, which is uploaded to a natural language big model such as deepseek for data simulation. Not only can it process text information, but also has multi-modal data fusion and scene reasoning capabilities. Through training and reasoning, it can simulate complex physical and working condition changes, realize deep understanding and automatic scene construction of multi-dimensional monitoring data, and build multiple abnormal simulation scenarios based on historical and real-time multi-source simulation features. Abnormal simulation scenarios refer to automatically simulating and generating multiple high-risk and extreme operating conditions including sensor failure, flow channel blockage, insufficient flow, and extreme load impact based on big model analysis capabilities, providing a virtual experimental environment for risk resistance capability evaluation and strategy optimization. Through the natural language big model, combined with current environmental parameters including external temperature and humidity, power state, cooling liquid type and multi-load working condition configuration, the actual response process under various abnormal simulation scenarios is simulated and deduced. The system automatically calls external environmental data and device configuration parameters to interactively simulate liquid cooling monitoring data and various criteria under various abnormal environments, deduce dynamic responses, working condition changes and risk transmission processes under various abnormal environments. During the simulation process, the natural language big model quantitatively evaluates the frequency, timing relationship and propagation path of various abnormal liquid cooling monitoring data, i.e. frequency statistics, timing causal chain analysis and multi-variable dynamic trajectory tracking of abnormal signal points in historical and simulation data. Based on the Bayesian network algorithm, the optimal variable dependency structure is generated based on historical and real-time liquid cooling monitoring data through a greedy algorithm, the conditional probability distribution of each variable is calculated, the causal structure between variables is learned and updated, and the causal relationship between variables is inferred. The Bayesian network algorithm is a graph structure modeling method based on probability theory, which can express the conditional dependency relationship and causal inference mechanism between multiple variables, and is widely used in intelligent diagnosis and risk prediction. When an abnormal signal is detected, the abnormal source path is identified by backtracking along the causal chain of the Bayesian network. The abnormal source path is used to lock the core node and causal chain that caused the abnormality, supporting subsequent operation intervention and model optimization.

[0060] In this embodiment, through the fusion of multi-source simulation features and the deep reasoning of intelligent big models, automatic simulation, interaction and source analysis of multi-dimensional monitoring data and abnormal criteria are realized. Multiple complex abnormal scenarios can be constructed, and working condition changes and risk transmission can be automatically deduced. The Bayesian network algorithm is used to accurately identify abnormal influence chains and core fault nodes, greatly improving the risk assessment, abnormal diagnosis and adaptive optimization capabilities under extreme working conditions, and providing a solid data foundation and technical support for efficient decision-making, intelligent operation and maintenance and long-term health management.

[0061] Specifically, the specific process of evaluating the flow regulation self-healing effect is: at the same time, the actual monitored signal type number and the corresponding liquid cooling monitoring data are obtained, including the cold plate outlet temperature, the cold plate outlet pressure and the microwave attenuation value, which respectively reflect the thermal state, fluid pressure and phase distribution of the cold plate outlet, and are key physical quantities for measuring the self-healing recovery effect; for each abnormal simulation scene, based on a sliding time window, the mean value of the ith signal before the execution of the flow regulation is calculated, and the mean value of the ith signal after the execution of the flow regulation is also calculated; the mean value of the ith signal after the execution of the flow regulation is subtracted from the mean value of the ith signal before the execution of the flow regulation, and the absolute value is taken to obtain the self-healing change amplitude of the ith signal, which is used to measure the actual recovery degree of each key signal after the implementation of the self-healing measure; the allowable fluctuation threshold of each signal is set, which can be adaptively set through the historical working condition standard deviation, and reflects the natural fluctuation range of each signal under normal operation; the self-healing change amplitude of the ith signal is subtracted from the allowable fluctuation threshold of the ith signal to obtain the over-allowed recovery deviation of the ith signal, which represents the part of the actual recovery effect after self-healing that exceeds the normal fluctuation range, i.e., the abnormal amplitude that is not completely recovered; the over-allowed recovery deviation of the ith signal is subjected to a maximum function operation, i.e., the larger one between the over-allowed recovery deviation of the ith signal and 0 is taken, if the over-allowed recovery deviation of the ith signal is positive, the actual calculation result is retained, otherwise 0 is taken, and the maximum function can prevent normal fluctuations and over-recovery from being counted as abnormal, effectively improving the engineering robustness of the criterion; based on the signal type number, the over-allowed recovery deviations of all signals are accumulated and averaged to obtain the average over-standard recovery difference, which is used to comprehensively represent the joint self-healing effect of multiple signals, and the smaller the value is, the better the overall recovery effect is, and the liquid cooling self-healing ability evaluation value is obtained by subtracting the average over-standard recovery difference by a constant one, and the closer to 1 it is, the stronger the flow regulation self-healing ability is, which can be used as a core criterion for subsequent adaptive optimization and health management.

[0062] wherein the specific formula of the liquid cooling self-healing ability evaluation value is:

[0063] ;

[0064] in the formula, represents the liquid cooling self-healing ability evaluation value, which measures whether multiple monitoring signals can be effectively restored to the normal level before and after the implementation of the self-healing measure under different abnormal scenes, whether the flow regulation is effective, and reflects the multi-modal self-healing ability; the liquid cooling self-healing ability evaluation value ranges from 0 to 1, the closer to 1 it is, the better the recovery effect of each key signal is, and the stronger the self-healing ability is; the closer to 0 it is, the worse the recovery is, and there is an unsolved abnormality; represents the signal type number, which is used to jointly judge the self-healing effect of multiple signals and reflects the multi-modal health; represents the mean of the ith signal before the execution of flow regulation, representing the baseline health state before the occurrence of abnormality; represents the mean of the ith signal after the execution of flow regulation, representing the true signal level after the recovery of abnormality; represents the allowable fluctuation threshold of the ith signal, based on the monitored signal in the historical sliding time window, the standard deviation of the signal in the real-time window is calculated, and the allowable fluctuation threshold is set to twice the standard deviation in real time; represents the self-healing change amplitude of the ith signal, quantifying the signal recovery amplitude, the smaller the difference is, the better; represents the super-permitted recovery deviation of the ith signal, wherein represents the maximum function operation, that is, only the calculation results of the super-permitted recovery deviation greater than 0 are retained, representing only the statistics of the over-standard part of the abnormality that has not been completely recovered, quantifying the details of the self-healing deficiency.

[0065] In this embodiment, by quantitatively analyzing the self-healing change amplitude and the historical fluctuation threshold of the key signals of the cold plate outlet temperature, pressure and microwave attenuation, the actual recovery effect of the flow regulation measures under abnormal conditions can be accurately evaluated. The multi-signal joint criterion and dynamic threshold setting improve the accuracy and engineering robustness of the self-healing capability evaluation, provide a scientific and quantifiable criterion basis for the adaptive optimization and health management of flow regulation, and enhance the intelligent response and long-term operation stability under complex working conditions.

[0066] Specifically, the specific process of optimizing each parameter and strategy and self-learning evolution is as follows: a liquid cooling self-healing capability evaluation value is calculated in real time as a core health index for feedback; if the liquid cooling self-healing capability evaluation value is greater than or equal to a self-healing threshold, it is considered that the current flow regulation strategy is effective, and the existing regulation is maintained; the self-healing threshold is dynamically set through historical data, and is used to distinguish the effective and ineffective intervals of the regulation strategy; if the liquid cooling self-healing capability evaluation value is less than the self-healing threshold, a monitoring signal with a recovery deviation greater than a recovery deviation threshold is identified, and the weight is reduced; the weight reduction refers to temporarily reducing the weight of the abnormal signal in the liquid cooling self-healing capability evaluation value criterion, avoiding the abnormal signal from having too great an impact on the overall regulation decision, and improving the robustness of multi-modal fusion; an optimized regulation measure is triggered according to a hierarchical response mechanism, the hierarchical response mechanism selects the optimal regulation strategy according to the liquid cooling self-healing capability evaluation value, and takes into account energy consumption and recovery speed: the cooling flow of the cold plate with poor recovery is mainly improved, that is, the implementation flow of the cold plate with poor self-healing effect is weighted and distributed, and the heat dissipation demand of the abnormal cold plate is preferentially guaranteed; the sensitivity, threshold and PID parameter of the flow regulation are adjusted to improve the recovery effect, wherein the PID parameter adjustment refers to optimizing the proportional, integral and differential coefficients of the flow regulator to improve the response speed and stability of the regulation; a short-time pulse flow flushing is started, the short-time pulse is used to remove local blockage, bubbles and deposition of the cold plate by temporarily increasing the flow, which helps to quickly recover the heat dissipation performance of the abnormal node and improve the recovery capability of the local signal; the liquid cooling self-healing capability evaluation value, the optimized regulation measure and the optimization effect are written into the liquid cooling monitoring database to provide big data support for the self-learning algorithm and historical case backtracking, and the flow regulation strategy and the liquid cooling self-healing capability evaluation value parameter are iteratively optimized, through continuous collection, analysis, feedback and parameter adjustment, the dynamic self-optimization and online learning of the regulation strategy and algorithm are realized, the self-adaptive evolution is carried out, the multi-signal self-healing capability is improved, and high robustness and intelligent self-healing capability are ensured in different running stages and working conditions.

[0067] In the embodiment, through the real-time feedback based on the self-healing capability evaluation value and the hierarchical response mechanism, the continuous self-adaptive optimization of the flow regulation parameter and strategy is realized. The weight of the abnormal signal can be dynamically adjusted, the flow can be finely distributed, the PID parameter can be optimized, and the pulse flushing measure can be combined, so that the self-healing response to complex abnormalities and the robustness of multi-signal fusion are effectively improved. Through the self-learning evolution driven by big data, intelligent, self-optimization and high-reliability operation are ensured in various working conditions, which provides a solid data and algorithm foundation for health management and fault prevention.

[0068] It is to be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting; it is not intended to exclude myriad other embodiments of the present application that other present or future devices, platforms, technologies, and methodologies can utilize as appropriate for particular situations. For example, the term "and / or" includes any and all combinations of one or more of the associated listed items. Expressions such as "at least one of," when preceding the syllables of a list of elements, modify the entire list of elements and do not modify the elements individually. It is to be understood that where the application, or portions thereof, is implemented using hardware, software, or a combination thereof, the various embodiments are not limited to any particular computing system or programming language. It is to be understood that the software implemented aspects of the application can be implemented in a variety of programming languages such as C, C++, Java, and the like. It is also to be understood that the application is not limited to any particular software or hardware implementation that can be employed in carrying out the application. It is to be understood that the terms "including," "comprising," or "having" contain for the purposes of interpretation only, an open term so as to cover a singular as well as plural, unless the context clearly dictates otherwise. It is to be understood that the phraseology "and others" is a term of art that is used herein to encompass a list of items that is not limited to the items specifically named, but can also include other items not specifically named. It is to be understood that the phrase "consisting essentially of" is a term of art that is used herein to encompass a list of items that is not limited to the items specifically named, but can also include other items not specifically named so long as the other items do not materially change the basic and novel characteristics of the claimed application. It is to be understood that the phrase "consisting of" is a term of art that is used herein to encompass a list of items that is limited to the items specifically named.

[0069] The preferred embodiments of the application disclosed above are only for helping to explain the application. The preferred embodiments do not describe all the details of the application, nor limit the application to the specific embodiments described. As understood by those skilled in the art, many modifications and variations can be made in light of the teachings above. The embodiments are chosen and described in order to best explain the principles and practical application of the application, so that others skilled in the art can best utilize the application. The application is limited only by the claims and their full scope and equivalents.

Claims

1. A method for dynamic flow control of a two-phase cold plate liquid cooling system based on multi-modal sensing, the method comprising: The method comprises the following steps: ​ S1, real-time acquisition of liquid cooling monitoring data, data preprocessing of the liquid cooling monitoring data; S2, based on the preprocessed liquid cooling monitoring data, the state risk of the cold plate is judged, and signal cross-validation is performed, according to the cold plate state risk judgment result, abnormal control measures are taken; S3, according to the cold plate state risk judgment result, the cooling liquid flow is regulated, and according to the preprocessed liquid cooling monitoring data, the matching degree of the flow regulation and the chip (8) thermal load is evaluated, and based on the matching degree evaluation result of the thermal load, the flow regulation is corrected; The specific process of the matching degree evaluation result of the thermal load for flow regulation correction is: Real-time comparison of cold plate load response matching value and load matching threshold value, when the cold plate load response matching value is less than or equal to the load matching threshold value, it is determined that the flow regulation matches well with the change of the chip (8) load, the current flow regulation strategy is maintained, and there is no need to adjust, only the normal frequency is continuously monitored; When the cold plate load response matching value is greater than the load matching threshold value, it is determined that the cold plate regulation response lags behind and does not match the load change, the cooling liquid flow is redistributed, the speed of the fluid pump (2) is adjusted; based on the thermal load distribution, the local abnormal nodes are identified, the flow of each cold plate is adjusted, the partition flow is accurately compensated, and the monitoring and sampling frequency is preferentially increased; when the chip temperature is not improved after monitoring the flow regulation, the chip (8) frequency reduction and dynamic power limit request is sent; a redundant flow channel is designed, when it is detected that the local blockage of the main flow channel continues to be not restored, part of the cooling liquid flow is switched to the standby channel, and the abnormal channel is isolated; At the same time, the cold plate load response matching value, the flow regulation strategy and the corresponding working condition are written into the liquid cooling monitoring database, and the cold plate load response matching value parameters and the flow regulation strategy are optimized according to the continuous monitoring result, until the cold plate load response matching value is less than or equal to the load matching threshold value; S4, through analyzing the preprocessed liquid cooling monitoring data, the cold plate state risk judgment result and the matching degree evaluation result of the thermal load, a big data modeling and simulation feedback platform is constructed, the flow regulation self-healing effect is evaluated, the parameters and strategies are optimized, and self-learning evolution is performed.

2. The method of claim 1, wherein, The specific process of real-time acquisition of liquid cooling monitoring data and data preprocessing of the liquid cooling monitoring data is: Temperature sensors (5), pressure sensors (4), microwave sensors (6) and flow sensors are arranged at the inlet and outlet of the cold plate and the key positions of the server (3), when the cold plate is first filled with liquid, the microwave attenuation value collected by the microwave sensor (6) is recorded as the microwave reference attenuation, and the difference between the inlet and outlet temperatures of the cold plate is recorded to obtain the temperature difference reference value, and the specific heat constant of the cooling liquid is obtained according to the type of the cooling liquid; The input power of the chip is calculated in real time by reading the power supply current and voltage through the intelligent platform management interface of the server (3) mainboard; Real-time acquisition of liquid cooling monitoring data, the liquid cooling monitoring data includes: microwave reference attenuation, temperature difference reference value, specific heat constant of cooling liquid, cold plate inlet temperature, cold plate outlet temperature, cold plate inlet pressure, cold plate outlet pressure, microwave attenuation value, cooling liquid flow, chip temperature and chip input power; The signal amplitude and fluctuation trend of each sensor are checked, abnormal and outlier sampling points are identified and removed; the master and standby sensors are deployed at the flow, temperature and pressure measurement points, the main control switches to the standby sensor when an anomaly occurs, if multiple sensors fail simultaneously, the adjacent nodes and historical data are used, and the virtual sensor data compensation is generated by using the multiple linear regression algorithm; the liquid cooling monitoring data collected is subjected to low-pass filtering and moving average processing to remove working condition noise and environmental interference; The microwave attenuation value is converted into a standard physical quantity by interpolation and calibration algorithm; all liquid cooling monitoring data are time-aligned and normalized by minimum-maximum normalization method; a liquid cooling monitoring database is constructed, and the liquid cooling monitoring data are stored in the liquid cooling monitoring database.

3. The method for dynamic flow control of multi-modal perception based two-phase cold plate liquid cooling system of claim 1, wherein, The specific process of judging the cold plate state risk based on the preprocessed liquid cooling monitoring data is as follows: The liquid cooling monitoring data are obtained, the cold plate inlet temperature is subtracted from the cold plate outlet temperature to obtain the cold plate inlet and outlet temperature difference, the current cold plate outlet pressure is divided by the current cold plate inlet pressure to obtain the flow pressure drop ratio, the current microwave attenuation value is divided by the microwave reference attenuation to obtain the microwave dryness ratio, and the current cold plate inlet and outlet temperature difference is divided by the temperature difference reference value to obtain the temperature difference normalization ratio; The flow pressure drop ratio, microwave dryness ratio and temperature difference normalization ratio are multiplied, and a constant one is added to perform natural logarithm operation to obtain the cold plate state comprehensive dryness value.

4. The method for dynamic flow control of multi-modal perception based two-phase cold plate liquid cooling system of claim 1, wherein, The specific process of performing signal cross-validation is as follows: The temperature signal, pressure signal and microwave attenuation signal collected by the sensor are subjected to cross-validation: for each time signal sample, the historical mean value of the current signal is calculated, and the relative change amplitude between each type of signal and its historical mean value is calculated by comparison; the relative change amplitudes of the signals are cross-compared, if the deviation between the normalized value of the relative change amplitude of one signal and the remaining signals exceeds the deviation threshold, it is determined that the signal drifts; and through the isolation forest algorithm, the time sequence fluctuation and overall correlation of each signal are comprehensively analyzed, the drift signal most related to the current anomaly is identified, and the weight of the drift signal in the cold plate state comprehensive dryness value is temporarily reduced.

5. The method for dynamic flow control of multi-modal perception based two-phase cold plate liquid cooling system of claim 1, wherein, The specific process of taking abnormal control measures according to the cold plate state risk judgment result is as follows: The cold plate state comprehensive dryness value is calculated in real time, and the cold plate state comprehensive dryness value is compared with the risk threshold value, when the cold plate state comprehensive dryness value is less than the risk threshold value, it is determined that the cold plate is in a safe interval; The current flow control strategy is maintained, and the operating parameters do not need to be adjusted, and only the monitoring and data collection are performed at the regular frequency; When the cold plate state comprehensive dryness value is greater than or equal to the risk threshold value, it is determined that there is a potential flow deterioration and heat exchange performance decline trend; The cooling flow is increased, the short-time high-flow pulse is started, and the flow ratio of each sub-cold plate is adjusted; The liquid cooling monitoring data sampling and monitoring frequency are increased; the flow, pressure, temperature and microwave attenuation value are detected in real time, it is judged whether the cold plate is locally blocked, the sensor fails or the overall is out of adjustment, the local abnormal flow channel is temporarily isolated, and the sensor signal is verified; The current working condition and the corresponding cold plate state comprehensive dryness value are written into the liquid cooling monitoring database.

6. The method for dynamic flow control of multi-modal perception based two-phase cold plate liquid cooling system of claim 1, wherein, According to the cold plate state risk discrimination result, the cooling liquid flow is regulated, and according to the pretreated liquid cooling monitoring data, the matching degree of the flow regulation and the chip (8) thermal load is evaluated. The specific process is as follows: Real-time regulation of cooling liquid flow is performed, a sliding time window is set, and time steps are divided. During the flow regulation process, real-time liquid cooling monitoring data is obtained. The difference between the cold plate outlet temperature and the cold plate inlet temperature is multiplied by the cooling liquid flow and the specific heat capacity constant of the cooling liquid to obtain the cold plate outlet heat flow. The derivative of the cold plate outlet heat flow with respect to time is calculated using the multi-point difference method for each time step to obtain the cold plate outlet heat flow rate. The derivative of the chip temperature with respect to time is also calculated for each time step to obtain the chip temperature rate. The difference between the chip temperature and the cold plate inlet temperature is divided by the chip input power to obtain the thermal resistance; The chip equivalent heat flow rate is obtained by multiplying the chip temperature rate by the inverse of the thermal resistance. The chip equivalent heat flow rate is subtracted from the cold plate outlet heat flow rate, and the absolute value is taken to obtain the heat flow response error. The heat flow response error of all time steps in the sliding time window is integrated to obtain the response error accumulation; At the same time, the cold plate outlet heat flow rate of all time steps in the sliding time window is integrated and added to a minimum constant value to obtain the heat flow change reference value. The cold plate load response matching value is obtained by dividing the response error accumulation by the heat flow change reference value.

7. The dynamic flow control method for multi-modal perception based two-phase cold plate liquid cooling system of claim 1, wherein, The specific process of constructing a big data modeling and simulation feedback platform by analyzing the pretreated liquid cooling monitoring data, the cold plate state risk discrimination result, and the matching degree evaluation result of the thermal load is as follows: Liquid cooling monitoring data, its corresponding cold plate state comprehensive dryness value, and cold plate load response matching value are combined to construct a multi-source simulation feature set. The multi-source simulation feature set is uploaded to a natural language big model for data simulation. Based on historical and real-time multi-source simulation features, multiple abnormal simulation scenarios are constructed. Through the natural language big model, combined with current environmental parameters including external temperature and humidity, power state, cooling liquid type, and multi-load working condition configuration, the actual response process under various abnormal simulation scenarios is simulated and deduced. During the simulation process, the natural language big model quantitatively evaluates the occurrence frequency, time sequence relationship, and propagation path of various abnormal liquid cooling monitoring data, and infers the causal relationship between variables based on the Bayesian network algorithm to identify the abnormal source path.

8. The method for dynamic flow control of multi-modal perception based two-phase cold plate liquid cooling system of claim 1, wherein, The specific process of evaluating the self-healing effect of flow regulation is as follows: At the same time, the actual monitoring signal type number and corresponding liquid cooling monitoring data, including cold plate outlet temperature, cold plate outlet pressure, and microwave attenuation value, are obtained. For each abnormal simulation scenario, based on the sliding time window, the mean of the ith signal before flow regulation is executed is calculated, and the mean of the ith signal after flow regulation is executed is calculated; The self-healing change amplitude of the ith signal is obtained by subtracting the mean of the ith signal before flow regulation from the mean of the ith signal after flow regulation and taking the absolute value. The allowable fluctuation threshold of each signal is set, the self-recovery change amplitude of the i-th signal is subtracted from the allowable fluctuation threshold of the i-th signal to obtain an over-permitted recovery deviation of the i-th signal, a maximum function operation is performed on the over-permitted recovery deviation of the i-th signal, that is, the greater one between the over-permitted recovery deviation of the i-th signal and 0 is taken, if the over-permitted recovery deviation of the i-th signal is a positive value, the actual calculation result is reserved, otherwise 0 is taken; based on the number of signal types, the over-permitted recovery deviations of all signals are accumulated and averaged to obtain an average over-standard recovery difference value, and the liquid cooling self-recovery ability evaluation value is obtained by subtracting a constant from the average over-standard recovery difference value.

9. The method for dynamic flow control of multi-modal perception based two-phase cold plate liquid cooling system of claim 1, wherein, The specific process of optimizing each parameter and strategy and performing self-learning evolution is as follows: The liquid cooling self-recovery ability evaluation value is calculated in real time and is fed back as a core health indicator; if the liquid cooling self-recovery ability evaluation value is greater than or equal to a self-recovery threshold, it is considered that the current flow regulation strategy is effective, and the existing regulation is maintained; If the liquid cooling self-recovery ability evaluation value is less than the self-recovery threshold, the monitoring signals with over-permitted recovery deviations greater than a recovery deviation threshold are identified and are given a weight reduction; an optimization regulation measure is triggered according to a hierarchical response mechanism: the cooling flow of the cold plate with poor recovery is mainly improved; The sensitivity, threshold and PID parameters of the flow regulation are adjusted to improve the recovery effect; A short-time pulse flow flushing is started to improve the recovery ability of local signals; The liquid cooling self-recovery ability evaluation value, the optimization regulation measure and the optimization effect are written into a liquid cooling monitoring database, the flow regulation strategy and the liquid cooling self-recovery ability evaluation value parameters are iteratively optimized, self-adaptive evolution is performed, and the multi-signal self-recovery ability is improved.

Citation Information

Patent Citations

  • Two-phase cold plate liquid cooling system and control method

    CN117979662A

  • Two-phase cold plate liquid cooling system and method

    CN118250982A

  • Control method and control system of data center liquid cooling heat dissipation system

    CN119597054A

  • Liquid cooling resource intelligent distribution regulation and control method and system

    CN120264673A