A method and system for monitoring and controlling the temperature of a liquid cooling plate

By acquiring the operating parameters of the liquid cooling system, calculating the deviation between the ideal coolant flow rate and the actual flow rate, and performing differentiated adjustments and trend analysis, the problems of coolant performance degradation and flow channel fouling in the liquid cooling system are solved, thereby improving the system's operational reliability and stability.

CN121568376BActive Publication Date: 2026-04-03HANGZHOU ZHONGTAI CRYOGENIC TECH CORP
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-26
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

In the long-term operation of existing liquid cooling systems, the performance of the coolant deteriorates and the accumulation of fouling in the flow channels leads to a decrease in heat dissipation efficiency and an increased risk of local overheating. Traditional control methods cannot actively predict or intervene in the decline in cooling efficiency, which may lead to premature failure of core components.

Method used

By acquiring the operating parameters of multiple independent cooling circuits, calculating the deviation between the ideal coolant flow rate and the actual flow rate, and performing differentiated adjustments and monitoring trend analysis, precise control and early warning of cooling intensity can be achieved.

Benefits of technology

It effectively avoids the risk of reduced heat dissipation efficiency and localized overheating caused by coolant performance degradation and flow channel fouling, improves the operational reliability and stability of the liquid cooling system, extends equipment service life and reduces maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121568376B_ABST
    Figure CN121568376B_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for monitoring and controlling the temperature of a liquid-cooled plate, belonging to the field of liquid cooling control. The method includes: acquiring operating parameters of multiple independent cooling loops; calculating, based on the operating parameters, the ideal coolant flow rate required for each cooling loop to reach a preset target outlet temperature under the current heat load; obtaining a heat dissipation efficiency index by comparing the ideal coolant flow rate with the actual coolant flow rate; differentially adjusting the cooling intensity of each cooling loop according to its heat dissipation efficiency index; monitoring the changing trend of the heat dissipation efficiency index of each cooling loop, and issuing early warnings based on the changing trend. This invention solves the problems of coolant performance degradation, decreased heat dissipation efficiency due to flow channel fouling, and increased risk of localized overheating in existing liquid cooling systems during long-term operation, thereby improving the operational reliability of the liquid cooling system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of liquid cooling control, specifically relating to a method and system for monitoring and controlling the temperature of a liquid cooling plate. Background Technology

[0002] In modern industrial production, data processing, and the operation of high-tech equipment, core functional components generate enormous amounts of heat during prolonged, high-intensity operation. To ensure the stable and reliable operation of these devices and extend their service life, efficient and reliable liquid cooling systems are indispensable. However, as the system operates over time, the performance of the coolant inside the liquid cooling system gradually deteriorates. Simultaneously, tiny impurities or decomposition products in the coolant accumulate in the microchannels of the liquid cooling plates, forming a fouling layer. These changes silently reduce heat dissipation efficiency and increase resistance to coolant flow.

[0003] Traditional liquid cooling systems typically employ a control strategy based on coolant outlet temperature feedback. This involves installing a temperature sensor at the outlet of the liquid cooling plate, transmitting the collected temperature data to a central control unit. The control unit then adjusts the coolant pump speed or the cooling capacity of the heat exchanger according to a preset safe temperature threshold to ensure the temperature of the heat-generating components remains within a safe range. This method is effective in controlling temperature during the initial, smooth operation of the system.

[0004] However, during long-term operation, the physicochemical properties of the coolant will slowly deteriorate, such as changes in parameters like thermal conductivity, specific heat capacity, and viscosity. Simultaneously, prolonged contact between the coolant and metal components can lead to corrosion, generating trace metal ions or particulate matter, accelerating coolant degradation. This degradation does not occur uniformly; the degradation rate varies across different liquid cooling plates or flow channel regions. This means that even if the outlet temperature of all liquid cooling plates is controlled at the same set value, the power pumps in different cooling circuits may be operating under different conditions, with some circuits potentially operating close to full load and experiencing severely insufficient cooling margin.

[0005] Furthermore, tiny particles or corrosion products from coolant deterioration can deposit on the inner walls of the microchannels of the liquid cooling plate, forming a fouling layer. This fouling layer significantly increases coolant flow resistance and reduces heat transfer efficiency. To maintain normal flow rate and outlet temperature, the power pump needs to operate at higher power, leading to increased energy consumption and accelerated pump wear. Uneven fouling deposition can cause localized narrowing or even blockage of flow channels, causing a sharp rise in the temperature of locally heated components and posing an overheating risk. Traditional control systems cannot accurately detect this risk of localized "insufficient cooling margin."

[0006] Ultimately, when coolant deterioration and flow channel fouling reach severe levels, the operational stability of the liquid cooling system will be seriously threatened. Traditional monitoring and control methods based on a single outlet temperature can only passively respond to temperature increases that have already occurred, and cannot predict or actively intervene in the continuous decline in cooling efficiency and the potential risk of localized overheating, which may cause expensive core components to fail prematurely due to prolonged unhealthy operating conditions. Summary of the Invention

[0007] The purpose of this invention is to overcome the defects in the prior art, so as to at least solve the problems of coolant performance degradation, flow channel fouling leading to decreased heat dissipation efficiency, increased risk of local overheating, and the inability of traditional control methods to predict or actively intervene in the decrease in cooling efficiency during long-term operation of existing liquid cooling systems, and to provide a liquid cooling plate temperature monitoring and control method and system.

[0008] The specific technical solution adopted in this invention is as follows:

[0009] In a first aspect, the present invention provides a method for monitoring and controlling the temperature of a liquid cooling plate, as follows:

[0010] S1: Obtain the operating parameters of multiple independent cooling circuits, wherein the operating parameters include the coolant inlet temperature, coolant outlet temperature, coolant flow rate, and heat load of the heat-generating components of each cooling circuit;

[0011] S2: Based on the operating parameters, calculate the ideal coolant flow rate required for each cooling circuit to reach the preset target outlet temperature under the current heat load;

[0012] S3: Based on the comparison between the ideal coolant flow rate and the actual coolant flow rate, obtain the heat dissipation performance index;

[0013] S4: Adjust the cooling intensity of each cooling circuit differently according to the heat dissipation performance index of each cooling circuit;

[0014] S5: Monitor the changing trend of the heat dissipation efficiency index of each cooling circuit, and issue an early warning based on the changing trend.

[0015] Preferably, S2 is as follows:

[0016] S21: Divide the heating component into multiple sub-regions and obtain the local temperature change rate and local power consumption data of each sub-region;

[0017] S22: Based on the local temperature change rate and local power consumption data of each sub-region, infer the local heat load of each sub-region;

[0018] S23: Weighted average of the local heat load of each sub-region to obtain a refined heat load;

[0019] S24: Using the refined heat load as the input of the heat load of the heat-generating component, calculate the ideal coolant flow rate required for each cooling circuit to reach the preset target outlet temperature under the current heat load.

[0020] Preferably, S5 is as follows:

[0021] S51: Obtain coolant inlet temperature data for multiple independent cooling circuits, and calculate the consensus inlet temperature based on the multiple coolant inlet temperature data;

[0022] S52: Compare the first deviation between each of the coolant inlet temperature data and the group consensus inlet temperature, and construct an inlet temperature stability index based on the first deviation;

[0023] S53: When the first deviation exceeds the first preset threshold, the group consensus inlet temperature is used as the calibration coolant inlet temperature of the cooling circuit;

[0024] S54: Obtain real-time cooling data for each of the cooling circuits, and infer the effective specific heat capacity of the coolant in each of the cooling circuits based on the real-time cooling data; wherein, the real-time cooling data includes the heat load of the heat-generating component, the inlet temperature of the coolant, the outlet temperature of the coolant, and the flow rate of the coolant;

[0025] S55: Based on the effective specific heat capacity and the preset coolant design specific heat capacity, construct the coolant performance index;

[0026] S56: Calculate the ideal heat transfer temperature difference of each cooling circuit based on the heat load of the heating component, the coolant flow rate, and the effective specific heat capacity;

[0027] S57: Based on the ideal heat transfer temperature difference and the second deviation between the coolant outlet temperature and the coolant inlet temperature, calculate the heat transfer efficiency deviation index of each cooling circuit, and construct the flow channel efficiency index based on the heat transfer efficiency deviation index.

[0028] S58: Obtain the estimated heat load of the heating element;

[0029] S59: Calculate the actual heat removed based on the coolant flow rate, the effective specific heat capacity, and the second deviation, and correct the estimated heat load based on the actual heat removed to obtain the calibrated heat load;

[0030] S510: Construct a heat load input reliability index based on the calibrated heat load and the estimated heat load;

[0031] S511: Using the inlet temperature of the calibration coolant, the effective specific heat capacity, and the calibration heat load as calibration factors, calibrate the heat dissipation efficiency index of each cooling circuit;

[0032] S512: The inlet temperature stability index, the coolant performance index, the flow channel efficiency index, and the heat load input reliability index are used as cooling health sub-indices that affect the heat dissipation performance index. The changing trends and magnitudes of each cooling health sub-index are analyzed to trace the root causes of the decline in the heat dissipation performance index and to issue warnings for the root causes.

[0033] Preferably, S512 is as follows:

[0034] S5121: When each of the cooling health sub-indices exceeds the third preset threshold, abnormal cooling health sub-indices are identified and obtained.

[0035] S5122: Based on the deviation degree and historical influence weight of each of the abnormal cooling health sub-indices, assess its contribution to the decline of the heat dissipation efficiency index, and obtain the contribution assessment result.

[0036] S5123: Based on the criticality level of the cooling circuit and the current load status of the heat-generating components, assess the maintenance urgency of each of the aforementioned root causes of the anomaly, and obtain the maintenance urgency assessment results;

[0037] S5124: Based on the contribution assessment results and the maintenance urgency assessment results, generate a maintenance suggestion list, and issue an early warning for the abnormal root causes according to the maintenance suggestion list; wherein, the maintenance suggestion list includes the dominant root causes, secondary root causes, and maintenance priority ranking.

[0038] Preferably, S5122 is as follows:

[0039] S51221: Identify the causal relationships between the various abnormal cooling health sub-indices;

[0040] S51222: Assess the first independent contribution of the abnormal cooling health sub-index with causal relationship, as well as synergistic or inhibitory effects;

[0041] S51223: Based on the first independent contribution and the synergistic effect or the inhibitory effect, evaluate the contribution of each of the abnormal cooling health sub-indices to the decrease in the heat dissipation efficiency index, and obtain the contribution evaluation result.

[0042] Preferably, S5123 is as follows:

[0043] S51231: Identify the types and quantities of currently available maintenance resources, and preliminarily assign maintenance tasks based on the maintenance urgency level of each of the aforementioned anomaly root causes and the types and quantities of maintenance resources required thereto;

[0044] S51232: When the resource requirements of the initially allocated maintenance task exceed the available maintenance resources or a resource conflict occurs, resource scheduling optimization is performed;

[0045] The resource scheduling optimization includes:

[0046] S512321: Prioritize the resource needs of high-urgency tasks that have a significant impact on system stability;

[0047] S512322: Adjust the execution order and time window of maintenance tasks according to the maintenance urgency level, the contribution of the anomaly root cause, and the estimated maintenance time.

[0048] S512323: Generate maintenance work orders and resource scheduling plans;

[0049] S512324: Monitor the maintenance effect and adjust subsequent maintenance plans and resource allocation based on the maintenance effect and actual maintenance progress.

[0050] Preferably, S512324 is as follows:

[0051] S5123241: Collect the cooling health sub-indices before and after maintenance, and identify the causal relationship between each cooling health sub-indice;

[0052] S5123242: Assess the second independent contribution of each of the aforementioned cooling health sub-indices that have a causal relationship, as well as synergistic or inhibitory effects;

[0053] S5123243: Based on the second independent contribution and the synergistic effect or the inhibitory effect, calculate the comprehensive contribution of each of the cooling health sub-indices to the decrease in heat dissipation efficiency index;

[0054] S5123244: Compare the changes in the overall contribution of each of the cooling health sub-indices before and after maintenance, and adjust the subsequent maintenance plan and resource allocation based on the maintenance effect and actual maintenance progress.

[0055] In a second aspect, the present invention provides a liquid-cooled plate temperature monitoring and control system, comprising:

[0056] The parameter acquisition module is used to acquire the operating parameters of multiple independent cooling circuits, wherein the operating parameters include the coolant inlet temperature, coolant outlet temperature, coolant flow rate, and heat load of the heat-generating components of each cooling circuit.

[0057] The efficiency calculation module is used to calculate the ideal coolant flow rate required for each of the cooling circuits to reach the preset target outlet temperature under the current heat load based on the operating parameters; and to obtain the heat dissipation efficiency index based on the comparison between the ideal coolant flow rate and the actual coolant flow rate.

[0058] An intensity adjustment module is used to adjust the cooling intensity of each cooling circuit differently according to the heat dissipation efficiency index of each cooling circuit.

[0059] The monitoring and early warning module is used to monitor the changing trends of the heat dissipation efficiency indicators of each cooling circuit and issue early warnings based on the changing trends.

[0060] Thirdly, the present invention provides a computer device including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the liquid cooling plate temperature monitoring and control method as described in any of the first aspects.

[0061] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the liquid cooling plate temperature monitoring and control method as described in any of the first aspects.

[0062] Compared with the prior art, the present invention has the following advantages:

[0063] By acquiring the operating parameters of multiple independent cooling loops, the real-time operating status of each loop is comprehensively understood. Based on these parameters, the ideal coolant flow rate required for each cooling loop to reach the preset target outlet temperature under the current heat load is calculated and compared with the actual coolant flow rate to obtain the heat dissipation efficiency index. The heat dissipation efficiency index directly reflects the deviation between the actual heat dissipation capacity of each cooling loop and the ideal state. Subsequently, based on the heat dissipation efficiency index of each cooling loop, the cooling intensity of each cooling loop is adjusted differentially. The cooling intensity is increased for loops with lower heat dissipation efficiency and appropriately decreased for loops with higher heat dissipation efficiency, thereby achieving optimal resource allocation. Finally, the changing trend of the heat dissipation efficiency index of each cooling loop is monitored, and early warnings are issued based on this trend. By analyzing the changing trend, a continuous decline in cooling efficiency and potential local overheating risks can be detected in advance, thereby achieving proactive prediction and intervention. This overcomes the problem that traditional methods can only passively respond to temperature increases that have already occurred, effectively preventing expensive core components from failing prematurely due to prolonged unhealthy operating conditions, significantly extending equipment lifespan, and reducing maintenance costs.

[0064] In summary, the method and system provided by this invention effectively solve the problems of coolant performance degradation, reduced heat dissipation efficiency caused by flow channel fouling, and increased risk of local overheating in existing liquid cooling systems during long-term operation by acquiring refined parameters, adjusting differentially based on heat dissipation performance indicators, and providing proactive early warning of changing trends. This improves the operational reliability and stability of the liquid cooling system.

[0065] Details of one or more embodiments of the present invention are set forth in the following drawings and description, so that other features, objects and advantages of the invention will be more readily understood. Attached Figure Description

[0066] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this invention, illustrate exemplary embodiments of the invention and are used to explain the invention, but do not constitute an undue limitation of the invention. In the drawings:

[0067] Figure 1 This is a flowchart illustrating a liquid cooling plate temperature monitoring and control method according to an exemplary embodiment.

[0068] Figure 2 This is a flowchart illustrating step S2 according to an exemplary embodiment.

[0069] Figure 3 This is a partial flowchart illustrating step S512 according to an exemplary embodiment.

[0070] Figure 4 This is a flowchart illustrating step S5122 according to an exemplary embodiment.

[0071] Figure 5 This is a flowchart illustrating step S5123 according to an exemplary embodiment.

[0072] Figure 6 This is a block diagram illustrating a liquid-cooled plate temperature monitoring and control system according to an exemplary embodiment. Detailed Implementation

[0073] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to provide a thorough understanding of the present invention. However, the present invention can be practiced in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below. Technical features in various embodiments of the present invention can be combined accordingly without mutual conflict.

[0074] This invention provides a method for monitoring and controlling the temperature of a liquid cooling plate. The method mainly includes: S1: acquiring the operating parameters of multiple independent cooling circuits, wherein the operating parameters include the coolant inlet temperature, coolant outlet temperature, coolant flow rate, and heat load of the heat-generating components of each cooling circuit; S2: calculating the ideal coolant flow rate required for each cooling circuit to reach a preset target outlet temperature under the current heat load based on the operating parameters; S3: obtaining the heat dissipation efficiency index by comparing the ideal coolant flow rate with the actual coolant flow rate; S4: adjusting the cooling intensity of each cooling circuit differently according to the heat dissipation efficiency index of each cooling circuit; S5: monitoring the changing trend of the heat dissipation efficiency index of each cooling circuit and issuing an early warning based on the changing trend.

[0075] The following will describe in detail with reference to specific embodiments and accompanying drawings.

[0076] Example 1

[0077] This embodiment provides a method for monitoring and controlling the temperature of a liquid-cooled plate, such as... Figure 1 The diagram shown is a flowchart of a liquid cooling plate flow control method. This method includes the following steps:

[0078] S1. Obtain the operating parameters of multiple independent cooling circuits; among which, the operating parameters include the coolant inlet temperature, coolant outlet temperature, coolant flow rate, and heat load of heat-generating components for each cooling circuit.

[0079] In this embodiment, independent cooling loops refer to cooling channels in the liquid cooling system that are independent of each other and do not interfere with each other. Each loop is typically responsible for cooling one or more specific heat-generating components, and these loops can independently monitor parameters and adjust cooling intensity. Operating parameters refer to physical quantities that can be acquired or calculated in real time during the operation of the liquid cooling system, including coolant inlet temperature, coolant outlet temperature, coolant flow rate, and heat load of the heat-generating components. Specifically, the coolant inlet and outlet temperatures are acquired in real time by installing high-precision temperature sensors at the inlet and outlet of each cooling loop. The coolant flow rate is acquired by flow meters installed in the loop. The heat load of the heat-generating components is obtained by relying on the heat-generating component's own power consumption sensor or by software estimation models. As a preferred implementation, these sensors are connected to a data acquisition unit, which converts the acquired analog signals into digital signals and transmits them to the central processing unit for further processing via a communication interface.

[0080] S2. Based on the operating parameters, calculate the ideal coolant flow rate required for each cooling circuit to reach the preset target outlet temperature under the current heat load.

[0081] In this embodiment, a thermodynamic model is first established. This model can calculate the theoretically required coolant flow rate, i.e., the ideal coolant flow rate, based on the heat load of the heat-generating component, the coolant inlet temperature, and the preset target outlet temperature. For example, the heat generated by the heat-generating component can be balanced with the heat carried away by the coolant based on the law of conservation of energy.

[0082] S3. Based on the comparison between the ideal coolant flow rate and the actual coolant flow rate, obtain the heat dissipation performance index.

[0083] In this embodiment, the heat dissipation performance index is a comprehensive indicator that measures the heat dissipation capacity and efficiency of the cooling circuit. It reflects the cooling circuit's ability to effectively remove heat generated by heat-generating components under current operating conditions. A decrease in the heat dissipation performance index usually indicates a potential problem with the cooling system. The actual measured coolant flow rate is compared with the calculated ideal coolant flow rate. If the actual flow rate is much greater than the ideal flow rate, it may mean a decrease in cooling efficiency; if the actual flow rate is close to or less than the ideal flow rate, it may indicate insufficient cooling capacity. Through this comparison, the heat dissipation performance index can be quantitatively obtained. For example, the ratio of the ideal coolant flow rate to the actual coolant flow rate can be used as the heat dissipation performance index, or the difference between the two can be used to construct the heat dissipation performance index.

[0084] S4. Adjust the cooling intensity of each cooling circuit differently based on the heat dissipation efficiency index of each cooling circuit.

[0085] In this embodiment, differentiated cooling intensity adjustment refers to targeted adjustments to the cooling capacity of different cooling loops based on their heat dissipation efficiency indicators. This can be achieved by changing factors such as coolant flow rate, pump speed, or heat exchanger power to achieve optimal cooling for each loop. The cooling intensity of each independent cooling loop is finely adjusted according to its heat dissipation efficiency indicators. For example, if the heat dissipation efficiency indicator of a certain cooling loop shows a decrease in cooling efficiency, the coolant flow rate of that loop is increased, for example, by increasing the speed of the corresponding cooling pump or adjusting the opening of the coolant valve in that loop. Conversely, if the heat dissipation efficiency indicator of a certain loop is good, or even if it is overcooled, its cooling intensity is appropriately reduced to save energy. Differentiated adjustment ensures that each heat-generating component receives adequate cooling, avoiding the potential for localized overheating or overcooling that can result from traditional "one-size-fits-all" control methods.

[0086] S5. Monitor the changing trends of the heat dissipation efficiency indicators of each cooling circuit and issue early warnings based on the changing trends.

[0087] In this embodiment, historical data on the heat dissipation performance indicators of each cooling loop are continuously recorded and analyzed. By analyzing the trends of this historical data, it is possible to identify whether the heat dissipation performance indicators are continuously declining, whether the rate of decline is accelerating, or whether there are abnormal patterns such as periodic fluctuations. For example, time series analysis methods, such as moving averages and exponential smoothing, can be used to predict future trends in heat dissipation performance indicators. When the trend of changes in heat dissipation performance indicators exceeds a preset safety threshold or shows a continuous deterioration trend, the system will immediately issue an early warning. An early warning means that when an abnormal trend in heat dissipation performance indicators occurs, the system issues a warning in advance, indicating the potential risk of failure or performance degradation, so that timely intervention measures can be taken. The warning information includes the identification of the abnormal loop, the type of abnormality, and suggested inspection or maintenance measures, so that operators can intervene in a timely manner.

[0088] The technical solution described above achieves quantitative assessment and early warning of the cooling system's health status by acquiring heat dissipation efficiency indicators and performing real-time calculations and trend monitoring. By acquiring multiple operating parameters such as coolant inlet temperature, coolant outlet temperature, coolant flow rate, and heat load of heat-generating components, a more comprehensive and in-depth understanding of the actual operating status of each independent cooling loop is obtained. Subsequently, comparing the ideal coolant flow rate calculated based on these operating parameters with the actual coolant flow rate directly reflects whether the cooling loop's heat dissipation efficiency has decreased, and to what extent. This heat dissipation efficiency indicator based on multi-parameter comprehensive analysis can more sensitively capture subtle changes within the cooling system, such as the slow decline in coolant performance or the gradual deposition of fouling in the flow channels—changes that are difficult to detect with traditional single outlet temperature feedback control. Furthermore, based on the independent heat dissipation efficiency indicators of each cooling loop, the cooling intensity is adjusted in a targeted manner, ensuring that each heat-generating component receives adequate cooling, thereby improving the overall system's energy efficiency and operational stability. Finally, through long-term monitoring and trend analysis of the heat dissipation efficiency indicators, the system can issue early warnings before problems worsen to the point of affecting equipment operation, indicating potential failure risks. Maintenance personnel can intervene in a timely manner to carry out preventive maintenance, such as replacing deteriorated coolant and cleaning the flow channels, thereby avoiding equipment downtime or damage to core components due to cooling system failure, extending the service life of the equipment, and reducing operating costs.

[0089] The method provided by this invention effectively solves the problems of coolant performance degradation, reduced heat dissipation efficiency caused by flow channel fouling and increased risk of local overheating in existing liquid cooling systems during long-term operation by acquiring refined parameters, adjusting differentially based on heat dissipation performance indicators and providing proactive early warning of changing trends, thereby improving the operational reliability and stability of liquid cooling systems.

[0090] In one feasible design, refer to the appendix Figure 2 Step S2 includes:

[0091] S21. Divide the heat-generating component into multiple sub-regions and obtain the local temperature change rate and local power consumption data of each sub-region.

[0092] In this embodiment, the heat-generating component is decomposed into several smaller, independently monitorable and analyzable units based on its physical structure, functional partitions, or thermal characteristics. For example, an integrated circuit chip can be divided into different functional modules (such as CPU cores, GPU cores, memory controllers, etc.) as sub-regions. The local temperature change rate and local power consumption data of each sub-region can be obtained by deploying temperature sensors and power consumption monitoring modules within each sub-region. This data can reflect the heat generation and accumulation in each sub-region in real time.

[0093] S22. Based on the local temperature change rate and local power consumption data of each sub-region, infer the local heat load of each sub-region.

[0094] In this embodiment, thermodynamic principles and empirical models are used to convert the collected local temperature change rate and local power consumption data into the actual heat generated in each sub-region. For example, the local temperature change rate can reflect the heat capacity and instantaneous heat accumulation of the sub-region, while the local power consumption data directly indicates the rate at which electrical energy is converted into heat energy. By combining this information, the local heat load of each sub-region can be estimated more accurately.

[0095] S23. The local heat load of each sub-region is weighted and averaged to obtain the refined heat load.

[0096] In this embodiment, considering that different sub-regions may have varying degrees of impact on the overall heat dissipation requirements, the local heat loads of each sub-region are weighted and summed. The weights are set based on factors such as the area, criticality, heat density, or impact on overall performance of the sub-region. For example, a higher weight can be assigned to the core region, which is more sensitive to temperature, to ensure that the accuracy of its heat load dominates the overall calculation. The refined heat load obtained in this way can more accurately reflect the overall heat generation of the heat-generating components.

[0097] S24. Using refined heat load as the input of heat load for heat-generating components, calculate the ideal coolant flow rate required for each cooling circuit to reach the preset target outlet temperature under the current heat load.

[0098] In this embodiment, by inputting more accurate heat load data, the accuracy of the ideal coolant flow rate calculation is improved, avoiding the errors that may be caused by using a single, coarse overall heat load, and making the calculation results closer to actual needs.

[0099] The technical solution of the above embodiments, by meticulously dividing the heat-generating component into multiple sub-regions and monitoring and analyzing the local temperature change rate and local power consumption data of each sub-region, can infer the local heat load of each sub-region. This refined heat load estimation method overcomes the problem that traditional methods, which rely solely on the overall heat load, cannot accurately reflect the uneven heat distribution within the heat-generating component. Subsequently, by weighted averaging these local heat loads, a more representative and accurate refined heat load can be obtained. Finally, using the refined heat load as input for calculating the ideal coolant flow rate makes the calculation results more accurately reflect the actual heat dissipation requirements of the current heat-generating component, effectively avoiding local overheating or overcooling, and further optimizing the overall heat dissipation performance and energy efficiency of the liquid cooling plate.

[0100] In one example, suppose a high-performance computing chip is used as a heat-generating component, which contains multiple CPU cores, GPU cores, and cache modules.

[0101] In implementing the technical solution of this embodiment, the chip is first divided into several sub-regions, for example, each CPU core, each GPU core, and each cache module are considered as an independent sub-region. Next, a miniature temperature sensor and power consumption monitoring circuit are deployed in each sub-region to acquire the local temperature change rate and local power consumption data of each sub-region in real time. For example, the local temperature change rate of CPU core 1 is 0.5℃ / s, and the local power consumption is 50W; the local temperature change rate of GPU core 1 is 0.8℃ / s, and the local power consumption is 80W.

[0102] Then, based on this data and the chip's thermal model, the local thermal load of CPU core 1 was inferred to be 55W, and the local thermal load of GPU core 1 was 90W. After obtaining the local thermal loads of all sub-regions, a weighted average was calculated based on the different weights assigned to each sub-region according to their varying sensitivity to overall chip performance and temperature. For example, the CPU cores and GPU cores might be given higher weights, while the cache module might have a relatively lower weight. Through this weighted average, a refined thermal load, such as 180W, reflecting the overall heat distribution and generation of the chip, can be obtained.

[0103] Finally, this refined heat load is used as input and substituted into the calculation model for the ideal coolant flow rate to calculate the ideal coolant flow rate required by each independent cooling loop under the current load, ensuring that the temperature of each area of ​​the chip reaches the preset target outlet temperature. For example, the calculation shows that cooling loop 1 requires a flow rate of 2.5 L / min, and cooling loop 2 requires a flow rate of 3.0 L / min. The method provided in this application significantly improves the accuracy of the ideal coolant flow rate calculation, making the temperature monitoring and control of the liquid cooling plate more precise and effective.

[0104] In one feasible design, step S5 includes:

[0105] S51. Obtain the coolant inlet temperature data of multiple independent cooling circuits, and calculate the consensus inlet temperature based on the multiple coolant inlet temperature data.

[0106] S52. Compare the first deviation between each coolant inlet temperature data and the group consensus inlet temperature, and construct an inlet temperature stability index based on the first deviation.

[0107] S53. When the deviation exceeds the first preset threshold, the group consensus inlet temperature is used as the calibration coolant inlet temperature of the cooling circuit.

[0108] In this embodiment, after acquiring coolant inlet temperature data for multiple independent cooling circuits, a reference inlet temperature can be provided for each cooling circuit by calculating a consensus inlet temperature. The consensus inlet temperature is the average or median of the inlet temperatures of all or most cooling circuits, and its purpose is to identify potential drift or anomalies in the inlet temperature sensors of individual cooling circuits. When the first deviation between the coolant inlet temperature data of a certain cooling circuit and the consensus inlet temperature exceeds a first preset threshold, it indicates a potential problem with the inlet temperature measurement of that circuit. In this case, the consensus inlet temperature is used as the calibration coolant inlet temperature for that cooling circuit to eliminate the impact of sensor error on the calculation of heat dissipation performance indicators, and an inlet temperature stability index is constructed to quantify this deviation.

[0109] S54. Obtain real-time cooling data for each cooling circuit, and inversely infer the effective specific heat capacity of the coolant in each cooling circuit based on the real-time cooling data. The real-time cooling data includes the heat load of the heat-generating components, the inlet temperature of the coolant, the outlet temperature of the coolant, and the flow rate of the coolant.

[0110] S55. Based on the effective specific heat capacity and the preset design specific heat capacity of the coolant, construct the coolant performance index.

[0111] S56. Calculate the ideal heat transfer temperature difference for each cooling circuit based on the heat load of the heating component, the coolant flow rate, and the effective specific heat capacity.

[0112] S57. Based on the ideal heat transfer temperature difference and the second deviation between the coolant outlet temperature and the coolant inlet temperature, calculate the heat transfer efficiency deviation index of each cooling circuit, and construct the flow channel efficiency index based on the heat transfer efficiency deviation index.

[0113] In this embodiment, the effective specific heat capacity refers to the heat transfer capacity of the coolant under actual operating conditions, reflecting the actual performance of the coolant, such as whether it has deteriorated or become contaminated. The coolant performance index is used to assess the health status of the coolant. The ideal heat transfer temperature difference refers to the temperature rise that the coolant should have when flowing through heat-generating components under ideal conditions. By comparing the ideal heat transfer temperature difference with a second deviation (i.e., the actual temperature rise) between the coolant outlet temperature and the coolant inlet temperature, the heat transfer efficiency deviation index of each cooling loop is calculated, and a flow channel efficiency index is constructed based on this. The flow channel efficiency index is used to assess whether there is blockage, scaling, or other problems affecting heat transfer efficiency in the internal flow channels of the cooling loop.

[0114] S58. Obtain the estimated heat load of the heating component.

[0115] S59. Calculate the actual heat removed based on the coolant flow rate, effective specific heat capacity, and second deviation, and correct the estimated heat load based on the actual heat removed to obtain the calibrated heat load.

[0116] S510. Based on the calibrated heat load and the estimated heat load, construct the heat load input reliability index.

[0117] In this embodiment, to improve the accuracy of heat load estimation, the estimated heat load of the heat-generating component is obtained. Based on the coolant flow rate, effective specific heat capacity, and second deviation, the actual heat removed can be calculated. By comparing the actual heat removed with the estimated heat load, the estimated heat load is corrected to obtain a calibrated heat load. Based on the calibrated heat load and the estimated heat load, a heat load input reliability index is constructed to evaluate the accuracy of the heat load estimation and avoid misjudgments of heat dissipation performance indicators due to inaccurate heat load input.

[0118] S511. Use the inlet temperature of the calibration coolant, the effective specific heat capacity, and the calibration heat load as calibration factors to calibrate the heat dissipation performance indicators of each cooling circuit.

[0119] In this embodiment, the inlet temperature of the calibrated coolant, the effective specific heat capacity, and the calibrated heat load are used as calibration factors to calibrate the heat dissipation performance indicators of each cooling circuit, thereby obtaining a more accurate heat dissipation performance evaluation.

[0120] S512. The inlet temperature stability index, coolant performance index, flow channel efficiency index, and heat load input reliability index are used as cooling health sub-indices that affect heat dissipation performance indicators. The changing trends and magnitudes of each cooling health sub-indice are analyzed to trace the root causes of the decline in heat dissipation performance indicators and to provide early warnings for the root causes.

[0121] In this embodiment, by analyzing the changing trends and magnitudes of the aforementioned cooling health sub-indices, the root causes leading to the decline in heat dissipation efficiency indicators can be traced more precisely, and early warnings can be issued for these root causes. For example, an abnormality in the inlet temperature stability index may indicate sensor failure, a decrease in the coolant performance index may indicate coolant degradation, a reduction in the flow channel efficiency index may indicate flow channel blockage, and an abnormality in the heat load input reliability index may indicate a problem with the heat load estimation model.

[0122] The technical solutions described above, by introducing multiple cooling health sub-indices and calibration mechanisms, overcome the limitations of traditional single heat dissipation performance monitoring in diagnosing the root causes of problems. Specifically, by calculating the consensus inlet temperature and constructing an inlet temperature stability index, the measurement deviations of individual cooling circuit inlet temperature sensors are identified and corrected, ensuring the accuracy of input data. By inversely inferring the effective specific heat capacity of the coolant and constructing a coolant performance index, the actual performance of the coolant is evaluated in real time, thereby detecting problems such as coolant degradation or contamination. By calculating the ideal heat transfer temperature difference and constructing a flow channel efficiency index, the heat transfer efficiency of the internal flow channels of the cooling circuit can be quantified, and physical faults such as blockages or scaling can be detected in a timely manner. By calculating the actual heat removed and correcting the estimated heat load, a heat load input reliability index is constructed, effectively improving the accuracy of heat load input and avoiding misjudgments caused by inaccurate heat load estimation.

[0123] Finally, the calibration factors obtained above are used to calibrate the heat dissipation performance indicators, enabling them to more accurately reflect the actual operating status of the cooling system. Ultimately, by comprehensively analyzing the changing trends and magnitudes of these cooling health sub-indices, the system can perform refined diagnosis of the decline in heat dissipation performance indicators from multiple dimensions, thereby accurately tracing the specific root causes of the problem, such as sensor failure, coolant degradation, flow channel blockage, or thermal load estimation errors, significantly improving the accuracy and diagnostic capabilities of liquid cooling plate temperature monitoring and control.

[0124] In one example, suppose a liquid cooling plate system contains ten independent cooling loops. During a certain operation, the system detected a continuous downward trend in the heat dissipation efficiency of the third cooling loop. If only the above-described monitoring and control method is used, the system will issue a general warning indicating a decline in the heat dissipation performance of the third cooling loop, but the specific cause cannot be determined.

[0125] However, according to the present invention, the system further analyzes the cooling health sub-index of the third cooling circuit. Specifically:

[0126] First, the system acquires coolant inlet temperature data for all cooling circuits and calculates the consensus inlet temperature. If the first deviation between the coolant inlet temperature data of the third cooling circuit and the consensus inlet temperature continues to increase and exceeds a first preset threshold, the inlet temperature stability index will show an anomaly. This may indicate that the inlet temperature sensor of the third cooling circuit is drifting or malfunctioning. In this case, the system will use the consensus inlet temperature as the calibration coolant inlet temperature for the third cooling circuit.

[0127] Secondly, the system uses real-time cooling data from the third cooling circuit (heat load of heat-generating components, coolant inlet temperature, coolant outlet temperature, and coolant flow rate) to infer the effective specific heat capacity of the coolant. If the effective specific heat capacity is found to be significantly lower than the preset design specific heat capacity of the coolant, the coolant performance index will show a downward trend. This may indicate that the coolant in the third cooling circuit has deteriorated or become contaminated.

[0128] Next, the system calculates the ideal heat transfer temperature difference for the third cooling loop and compares it with the actual temperature difference (the second deviation between the coolant outlet temperature and the inlet temperature) to calculate the heat transfer efficiency deviation index, thereby constructing the flow channel efficiency index. If the flow channel efficiency index continues to decrease, it may indicate that there is blockage or scaling in the internal flow channels of the third cooling loop.

[0129] Finally, the system obtains the estimated heat load of the heat-generating components and corrects it based on the actual heat removed by the third cooling circuit, constructing a heat load input reliability index. If this index shows an anomaly, it may indicate a deviation in the heat load estimation model for the heat-generating components.

[0130] By comprehensively analyzing the trends and magnitudes of these cooling health sub-indices, the root causes of the decline in heat dissipation efficiency can be traced. For example, if the inlet temperature stability index is abnormal while other sub-indices are normal, the warning message will clearly indicate "faulty inlet temperature sensor in the third cooling circuit." If the coolant performance index drops significantly, the warning message will indicate "coolant degradation in the third cooling circuit." This refined diagnostic capability allows maintenance personnel to directly inspect and repair the root cause of the problem, such as replacing sensors, replacing coolant, or cleaning the flow channels, thereby greatly improving maintenance efficiency and accuracy.

[0131] In one feasible design, refer to the appendix Figure 3 In step S512, "analyzing the changing trends and magnitudes of each cooling health sub-index, tracing the root causes of the decline in heat dissipation efficiency indicators, and issuing warnings for the root causes of the abnormalities" specifically includes the following steps:

[0132] S5121. When each cooling health sub-index exceeds the third preset threshold, abnormal cooling health sub-index is identified and obtained.

[0133] In this embodiment, when the values ​​of various cooling health sub-indices, such as the inlet temperature stability index, coolant performance index, flow channel efficiency index, or heat load input reliability index, deviate from the normal range and exceed a preset third threshold, the system will be triggered to identify and acquire these abnormal cooling health sub-indices. The third preset threshold is set based on historical data, expert experience, or system design requirements, and is used to define what degree of deviation is considered abnormal.

[0134] S5122. Based on the deviation degree and historical influence weight of each abnormal cooling health sub-index, assess its contribution to the decline in heat dissipation efficiency index and obtain the contribution assessment result.

[0135] In this embodiment, for each identified abnormal cooling health sub-index, the degree to which the individual abnormal cooling health sub-index deviates from its normal value is analyzed, as well as the weight of the individual abnormal cooling health sub-index's impact on the decline in heat dissipation performance in historical data. The historical impact weight is learned from historical fault data through machine learning models (such as regression analysis, decision trees, etc.), reflecting the relative importance of different sub-indices in causing the decline in heat dissipation performance. By combining the degree of deviation and the historical impact weight, the contribution of each abnormal sub-index is quantified, thereby obtaining the contribution evaluation result.

[0136] S5123. Based on the criticality level of the cooling circuit and the current load status of the heat-generating components, assess the maintenance urgency of each abnormality source and obtain the maintenance urgency assessment results.

[0137] In this embodiment, to more effectively conduct early warning and maintenance, it is necessary to assess the maintenance urgency of each anomaly root cause. The assessment of maintenance urgency is based on the criticality level of the cooling circuit and the current load status of the heat-generating components. The criticality level of the cooling circuit is classified according to factors such as its importance in the overall system and its impact on system stability, for example, into high, medium, and low. The load status of the heat-generating components reflects the current system workload and heat dissipation requirements. For example, under high load conditions, even minor anomalies can lead to serious consequences, thus increasing maintenance urgency. By comprehensively considering these factors, the maintenance urgency assessment result is obtained.

[0138] S5124. Based on the contribution assessment results and maintenance urgency assessment results, generate a maintenance suggestion list and issue warnings for abnormal root causes according to the maintenance suggestion list. The maintenance suggestion list includes the dominant root cause, secondary root cause, and maintenance priority ranking.

[0139] In this embodiment, based on the contribution assessment results and the maintenance urgency assessment results, the system generates a maintenance suggestion list. This list not only identifies the dominant and secondary root causes of the decline in heat dissipation performance indicators but also provides a detailed maintenance priority ranking. The dominant root cause refers to the physical problem corresponding to the anomaly sub-index that contributes the most to the decline in heat dissipation performance, while secondary root causes are other anomalies that contribute less. The maintenance priority ranking combines contribution and urgency, guiding maintenance personnel to prioritize which issues to address to maximize maintenance efficiency and system stability. Based on this maintenance suggestion list, precise early warnings are issued for anomaly root causes.

[0140] The technical solution described above addresses the shortcomings of traditional early warning mechanisms in providing specific maintenance guidance by introducing an assessment of the contribution of abnormal cooling health sub-indices and a maintenance urgency assessment. Specifically, when a heat dissipation efficiency index shows a downward trend, the system first identifies subsystems or parameters that may have problems by identifying abnormal cooling health sub-indices that exceed a third preset threshold. Then, by quantifying the deviation degree and historical impact weight of each abnormal sub-index, its actual contribution to the overall decrease in heat dissipation efficiency is objectively assessed, avoiding potential biases from relying solely on experience. Simultaneously, the maintenance urgency is assessed by combining the criticality level of the cooling circuit and the load status of heat-generating components, ensuring that the early warning not only focuses on the severity of the problem but also considers its potential impact on the current system operation and the necessity of maintenance, thereby ensuring the reasonable allocation of maintenance resources and timely response. Finally, by generating a maintenance recommendation list containing the primary root cause, secondary root causes, and maintenance priority ranking, clear and actionable maintenance guidance is provided to operators, transforming the early warning into a concrete action plan.

[0141] In one example, suppose a liquid cooling plate system is monitored to show a downward trend in its heat dissipation efficiency during operation. Based on the method described above, cooling health sub-indices such as the inlet temperature stability index, coolant performance index, flow channel efficiency index, and heat load input reliability index are first calculated and analyzed.

[0142] Furthermore, it was found that both the coolant performance index and the flow channel efficiency index exceeded the preset third threshold, and were identified as abnormal cooling health sub-indices. Based on historical data analysis, it was found that the coolant performance index deviated significantly and had a high weight in terms of its impact on the decline in heat dissipation efficiency in historical failures. Therefore, its contribution to the current decline in heat dissipation efficiency was assessed as 60%. In contrast, the flow channel efficiency index deviated less significantly and had a lower weight in terms of historical impact, with its contribution assessed as 30%.

[0143] Simultaneously assess maintenance urgency. Suppose the cooling circuit is responsible for cooling a critical heat-generating component, and the heat-generating component is currently operating under high load, the system will determine that the maintenance urgency of the root cause of the anomaly is "high".

[0144] Based on the above contribution assessment and maintenance urgency assessment results, the system generates a maintenance recommendation list. This list may indicate that the primary cause is "coolant performance degradation (such as coolant deterioration or contamination)," and the secondary cause is "local blockage or scaling in the flow channels." The maintenance priority is: first, replace or treat the coolant; second, inspect and clean the relevant flow channels. Based on this list, an alert is issued to operators, recommending immediate and appropriate maintenance measures to prevent further deterioration of heat dissipation performance and ensure stable system operation.

[0145] In one feasible design, refer to the appendix Figure 4 Step S5122 includes:

[0146] S51221. Identify the causal relationships between various abnormal cooling health sub-indices.

[0147] In this embodiment, by analyzing historical operating data, combining expert experience, or utilizing causal inference models (such as Granger causality tests, Bayesian networks, etc.), it is determined whether there are direct or indirect causal chains between different cooling health sub-indices. For example, a decrease in the coolant performance index may lead to a decrease in the flow channel efficiency index, because coolant degradation affects its heat transfer capacity, thereby affecting the overall heat transfer efficiency within the flow channel.

[0148] S51222, assess the first independent contribution of the abnormal cooling health sub-index with causal relationship, as well as synergistic or inhibitory effects.

[0149] In this embodiment, for the identified anomalous cooling health sub-indices with causal relationships, it is first necessary to quantify the independent contribution of each sub-index after excluding the influence of other related sub-indices. Simultaneously, it is necessary to analyze whether multiple anomalous cooling health sub-indices mutually reinforce (synergistic effect) or mutually weaken (inhibitory effect) when they occur simultaneously. Synergistic effect refers to the total effect of multiple factors acting together being greater than the sum of the effects of their individual actions; inhibitory effect, on the contrary, refers to the presence of one factor reducing the influence of another factor. The synergistic or inhibitory effects are quantified using statistical regression models, machine learning algorithms, or simulations based on physical models.

[0150] S51223. Based on the first independent contribution and the synergistic or inhibitory effects, assess the contribution of each abnormal cooling health sub-index to the decline in heat dissipation performance index, and obtain the contribution assessment results.

[0151] In this embodiment, the independent impact of each abnormal cooling health sub-index, its synergistic effect with other sub-indexes, and any potential inhibitory effects are comprehensively considered to calculate its final, comprehensive contribution to the decline in heat dissipation performance. For example, a sub-index may contribute 10% independently, but has a 5% synergistic effect with another sub-index; therefore, its total contribution will be 10% plus the additional impact of the synergistic effect.

[0152] The technical solutions of the above embodiments, by identifying the causal relationships between various abnormal cooling health sub-indices, enable a deeper understanding of the underlying logic of the problem. For example, when the inlet temperature stability index decreases, if it can be identified that this is due to a decrease in the coolant performance index, then maintaining the coolant performance index will be more fundamental. Furthermore, by evaluating the first independent contribution, as well as synergistic or inhibitory effects, the true impact of each abnormal sub-index is quantified. The evaluation of synergistic effects helps identify those "exacerbating" combinations, while the evaluation of inhibitory effects avoids overemphasizing some sub-indices that appear abnormal but whose actual impact is offset by other factors. Ultimately, the solution of this application can avoid contribution evaluation bias caused by ignoring complex interactions, making the root cause of the decline in heat dissipation performance indicators more accurate, thereby improving the overall effectiveness and maintenance efficiency of liquid cooling plate temperature monitoring and control.

[0153] In one feasible design, refer to the appendix Figure 5 Step S5123 includes:

[0154] S51231. Identify the types and quantities of currently available maintenance resources, and preliminarily allocate maintenance tasks based on the maintenance urgency level of each anomaly root cause and the types and quantities of maintenance resources required.

[0155] In this embodiment, the system or management platform performs real-time inventory and classification statistics of various resources currently available for maintenance operations, such as manpower, equipment, tools, and spare parts. This includes, for example, the number and professional skills of maintenance engineers, the availability of specialized testing equipment, and the inventory of specific spare parts, thus providing accurate resource information for subsequent maintenance task allocation and scheduling. During the initial allocation of maintenance tasks, preliminary matching and allocation are performed based on the maintenance urgency level of each anomaly root cause (e.g., high, medium, low) and the specific type and quantity of maintenance resources required for that anomaly root cause. For example, an anomaly root cause requiring the replacement of a specific coolant pump will be initially allocated to maintenance personnel with relevant skills and the required model of coolant pump spare parts.

[0156] S51232. When the resource requirements of the initially allocated maintenance tasks exceed the available maintenance resources or a resource conflict occurs, resource scheduling optimization shall be performed, wherein resource scheduling optimization includes.

[0157] S51232-A: Prioritize the resource needs of high-urgency tasks that have a significant impact on system stability.

[0158] S51232-B: Adjust the execution order and time window of maintenance tasks according to the maintenance urgency level, the contribution of the anomaly root cause, and the estimated maintenance time.

[0159] S51232-C generates maintenance work orders and resource scheduling plans.

[0160] S5123-D: Monitor maintenance effectiveness and adjust subsequent maintenance plans and resource allocation based on maintenance effectiveness and actual maintenance progress.

[0161] In this embodiment, resource scheduling optimization is triggered when the resource requirements of the initially allocated maintenance tasks exceed the total amount of currently available maintenance resources, or when multiple maintenance tasks compete for the same scarce resource. Resource scheduling optimization aims to resolve resource conflicts and ensure the priority execution of critical tasks through intelligent algorithms or preset strategies. As a preferred implementation, resource scheduling optimization prioritizes the resource requirements of highly urgent tasks that have a significant impact on system stability. That is, maintenance tasks for anomalies that could lead to system downtime, severe performance degradation, or increased security risks will be given the highest resource allocation priority. For example, if the flow channel efficiency index of a cooling circuit drops sharply, potentially causing overheating of heat-generating components, the resource requirements of its maintenance task will be prioritized.

[0162] Furthermore, resource scheduling optimization adjusts the execution order and time window of maintenance tasks based on the urgency level, the contribution of the root cause of the anomaly, and the estimated maintenance time. For example, for tasks with equal urgency, a task with a higher contribution (a greater impact on heat dissipation performance) and a shorter estimated time may be prioritized. The adjustment of the time window considers the available time period of resources and the task's deadline. Subsequently, after resource scheduling optimization, the system generates detailed maintenance work orders and resource scheduling plans. The maintenance work order specifies the maintenance content, required resources, executors, and estimated completion time; the resource scheduling plan details the allocation of various resources in different time periods to guide actual maintenance operations.

[0163] In addition, to form a closed-loop management system, the system will continuously track the changes in relevant cooling health sub-indices during and after the maintenance task is executed, assess whether the maintenance operation has effectively resolved the root cause of the anomaly, and dynamically adjust the subsequent maintenance plan and resource allocation according to the actual maintenance progress (e.g., whether it was completed on time or whether unexpected situations were encountered) to adapt to changes in the actual situation.

[0164] The technical solution of the above embodiments improves the maintenance management capability of the liquid cooling plate temperature monitoring and control system by introducing the identification of available maintenance resources, preliminary task allocation, and an intelligent scheduling optimization mechanism in case of resource conflicts or shortages. Specifically, firstly, the types and quantities of currently available maintenance resources are identified, providing a clear resource profile for the actual execution of maintenance tasks. Secondly, preliminary allocation is performed based on the maintenance urgency level of the root cause of the anomaly and the required resources, ensuring basic resource matching.

[0165] When resource demands exceed available resources or conflicts occur, resource scheduling optimization prioritizes resource needs for high-urgency tasks that significantly impact system stability. The execution order and time windows of tasks are adjusted by comprehensively considering maintenance urgency levels, the contribution of anomaly root causes, and estimated maintenance time. This ensures that critical maintenance tasks receive timely and prioritized resource support and execution. Finally, by monitoring maintenance effectiveness and actual progress, a feedback loop is formed, enabling dynamic adjustments to maintenance plans and resource allocation based on actual conditions, further enhancing the flexibility of maintenance management.

[0166] In one example, suppose a data center deploys multiple liquid cooling plates, each corresponding to an independent cooling loop. The system uses the method described above to detect three root causes of anomalies:

[0167] 1. The coolant performance index of cooling circuit A has decreased, indicating that the coolant may be deteriorating. The maintenance urgency level is "high". It is expected that the coolant will need to be replaced and the system will be flushed, which will take 8 hours and require 2 senior maintenance engineers and a set of special coolant replacement equipment.

[0168] 2. The flow efficiency index of cooling circuit B has decreased slightly, indicating that there may be a slight blockage. The maintenance urgency level is "medium". It is expected that flow channel flushing will be required, which will take 4 hours and require 1 intermediate maintenance engineer and a set of general flushing equipment.

[0169] 3. The inlet temperature stability index of cooling circuit C fluctuates greatly, which may indicate a sensor failure. The maintenance urgency level is "medium". It is estimated that the sensor needs to be checked and calibrated, which will take 2 hours and requires one junior maintenance engineer and a set of sensor calibration tools.

[0170] The currently available maintenance resources include: 2 senior maintenance engineers, 1 intermediate maintenance engineer, 1 junior maintenance engineer, 1 set of dedicated coolant replacement equipment, 1 set of general flushing equipment, and 1 set of sensor calibration tools.

[0171] In terms of specific operations, firstly, the system identifies the types and quantities of currently available maintenance resources. Next, it performs initial allocation:

[0172] Cooling circuit A: Requires 2 senior engineers and specialized equipment, matching available resources.

[0173] Cooling Loop B: Requires 1 intermediate engineer and general equipment, matching available resources.

[0174] Cooling loop C: Requires 1 junior engineer and calibration tools, matching available resources.

[0175] In this initial allocation, the resource requirements of all tasks did not exceed the available resources, and no resource conflicts occurred.

[0176] However, if another abnormal root cause D occurs at this time: the coolant pump in cooling circuit D fails, causing the temperature of the heat-generating components to rise sharply, the maintenance urgency level is "urgent", and it is expected that the coolant pump will need to be replaced, which will take 6 hours and require 2 senior maintenance engineers and a spare coolant pump.

[0177] At this point, the required number of senior maintenance engineers has increased to four (two for cooling circuit A and two for cooling circuit D), but only two are available, resulting in a resource conflict. The system will then perform resource scheduling optimization.

[0178] Since the root cause of the anomaly in cooling circuit D has the greatest impact on system stability (urgent), the system will prioritize ensuring the resource needs of cooling circuit D. Therefore, two senior maintenance engineers and a spare coolant pump will be allocated to cooling circuit D first.

[0179] While the maintenance task for cooling circuit A (replacing the coolant) is also of high urgency, its immediate impact on system stability is slightly lower than that of cooling circuit D. Therefore, the maintenance task for cooling circuit A will be postponed, pending the availability of senior maintenance engineers.

[0180] The maintenance tasks for cooling circuits B and C can be carried out as originally planned because the required resources do not conflict and the urgency level is low.

[0181] Finally, the system generates maintenance work orders and resource scheduling plans:

[0182] First priority: Replacement of the coolant pump in cooling circuit D, to be performed by two senior engineers, expected to take 6 hours.

[0183] Second priority: Flushing of the flow channels in cooling circuit B, to be performed by one intermediate engineer, expected to take 4 hours.

[0184] Third priority: Sensor inspection of cooling circuit C, to be performed by a junior engineer, expected to take 2 hours.

[0185] Fourth priority: Coolant replacement for cooling circuit A. This will be performed by two senior engineers after the task for cooling circuit D is completed, and is expected to take 8 hours.

[0186] During maintenance, the system continuously monitors the heat dissipation efficiency indicators and cooling health sub-indices of each cooling circuit. For example, if the heat dissipation efficiency indicator quickly returns to normal after the coolant pump in cooling circuit D is replaced, it indicates a good maintenance result. If the flow channel efficiency index does not recover satisfactorily after flushing the flow channels in cooling circuit B, subsequent plans may need to be adjusted, such as scheduling a more thorough cleaning or inspection. In this way, maintenance plans and resource allocation can be dynamically adjusted based on actual maintenance results and progress, ensuring the effectiveness of maintenance work and the rational use of resources.

[0187] In a feasible design, step S5123-D (i.e., S512324) includes:

[0188] S5123241. Collect cooling health sub-indices before and after maintenance, and identify the causal relationship between each cooling health sub-indice.

[0189] In this embodiment, the system collects data on various cooling health sub-indices before and after maintenance activities. These sub-indices include the inlet temperature stability index, coolant performance index, flow channel efficiency index, and heat load input reliability index, each reflecting different aspects of the liquid cooling plate's health. By comparing these sub-indices before and after maintenance, the impact of maintenance activities on the performance of various parts of the system can be intuitively understood. Identifying the causal relationships between the various cooling health sub-indices means that in complex liquid cooling systems, these sub-indices are not independent; they may have causal chains that influence each other. For example, a decline in coolant performance may lead to a decrease in flow channel efficiency. By establishing or learning these causal relationship models, it is possible to more accurately understand how maintenance activities affect one sub-indice and subsequently other sub-indices, ultimately impacting the overall heat dissipation performance indicators.

[0190] S5123242, assess the second independent contribution of each cooling health sub-index with causal relationship, as well as synergistic or inhibitory effects.

[0191] In this embodiment, based on the identification of causal relationships, the degree of independent influence of each sub-index on the decline of heat dissipation performance index (second independent contribution) is quantified, while considering the synergistic effect (i.e. the enhancement effect produced when multiple sub-indexes work together) or the inhibitory effect (i.e. the negative impact of one sub-index on another sub-index).

[0192] S5123243. Based on the second independent contribution and the synergistic or inhibitory effects, calculate the comprehensive contribution of each cooling health sub-index to the decline in heat dissipation efficiency index.

[0193] In this embodiment, the overall contribution is a quantitative indicator that combines independent contribution, synergistic effect, and inhibitory effect to comprehensively reflect the overall impact of each cooling health sub-index on the decline of heat dissipation performance index before and after maintenance.

[0194] S5123244. Compare the changes in the overall contribution of each cooling health sub-index before and after maintenance, and adjust the subsequent maintenance plan and resource allocation according to the maintenance effect and actual maintenance progress.

[0195] In this embodiment, by comparing the changes in the overall contribution of each cooling health sub-index before and after maintenance, it is clear which sub-indexes improved after maintenance, which showed little improvement, and which may even have worsened. Based on this quantitative change, the system can more accurately assess the maintenance effect and actual maintenance progress, thereby adjusting subsequent maintenance plans and resource allocation in a targeted manner. For example, resources can be prioritized for maintenance tasks related to sub-indexes that show little improvement but have a significant impact on heat dissipation performance, ensuring continuous optimization of maintenance activities and efficient use of resources.

[0196] Therefore, the liquid cooling plate temperature monitoring and control method provided in this embodiment of the invention comprehensively grasps the real-time operating status of each circuit by acquiring the operating parameters of multiple independent cooling circuits. Based on these operating parameters, it calculates the ideal coolant flow rate required for each cooling circuit to reach the preset target outlet temperature under the current heat load and compares it with the actual coolant flow rate to obtain the heat dissipation efficiency index. The heat dissipation efficiency index directly reflects the degree of deviation between the actual heat dissipation capacity of each cooling circuit and the ideal state. Subsequently, according to the heat dissipation efficiency index of each cooling circuit, the cooling intensity of each cooling circuit is adjusted differently, increasing the cooling intensity of circuits with low heat dissipation efficiency and appropriately decreasing the cooling intensity of circuits with high heat dissipation efficiency, thereby achieving optimal resource allocation. Finally, it monitors the changing trend of the heat dissipation efficiency index of each cooling circuit and issues early warnings based on this trend. By analyzing the changing trend, it can detect the continuous decline in cooling efficiency and potential local overheating risks in advance, thereby achieving proactive prediction and intervention. This overcomes the problem that traditional methods can only passively respond to temperature increases that have already occurred, effectively preventing expensive core components from failing prematurely due to long-term unhealthy working conditions, significantly extending equipment lifespan, and reducing maintenance costs.

[0197] Example 2

[0198] This embodiment provides a liquid-cooled plate temperature monitoring and control system, the system comprising:

[0199] The parameter acquisition module 01 is used to acquire the operating parameters of multiple independent cooling circuits, including the coolant inlet temperature, coolant outlet temperature, coolant flow rate, and heat load of the heat-generating components for each circuit.

[0200] The efficiency calculation module 02 is used to calculate the ideal coolant flow rate required for each cooling circuit to reach the preset target outlet temperature under the current heat load based on the operating parameters, and to obtain the heat dissipation efficiency index by comparing the ideal coolant flow rate with the actual coolant flow rate.

[0201] The intensity adjustment module 03 is used to adjust the cooling intensity of each cooling circuit differently according to the heat dissipation efficiency index of each cooling circuit.

[0202] The monitoring and early warning module 04 is used to monitor the changing trends of the heat dissipation efficiency indicators of each cooling circuit and issue early warnings based on these trends.

[0203] Therefore, the method and system provided by this invention effectively solve the problems of coolant performance degradation, reduced heat dissipation efficiency and increased risk of local overheating caused by flow channel fouling in existing liquid cooling systems during long-term operation by acquiring refined parameters, adjusting differentially based on heat dissipation performance indicators and actively warning of changing trends, thereby improving the operational reliability and stability of liquid cooling systems.

[0204] Example 3

[0205] This embodiment provides a computer device, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps of the method provided in Embodiment 1.

[0206] Example 4

[0207] This embodiment provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the method provided in Embodiment 1.

[0208] The readable storage medium may be more specifically adopted, including but not limited to: portable disk, hard disk, random access memory, read-only memory, erasable programmable read-only memory, optical storage device, magnetic storage device, or any suitable combination thereof.

[0209] In a possible implementation, the present invention can also be implemented as a program product comprising program code that, when the program product is run on a terminal device, causes the terminal device to perform the steps of implementing the method provided in Embodiment 1.

[0210] The program code for executing the present invention can be written in any combination of one or more programming languages. The program code can be executed entirely on the user device, partially on the user device, as a standalone software package, partially on the user device and partially on a remote device, or entirely on a remote device.

[0211] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the invention. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the invention. Therefore, all technical solutions obtained through equivalent substitution or transformation fall within the protection scope of the present invention.

Claims

1. A method for monitoring and controlling the temperature of a liquid-cooled plate, characterized in that, Specifically as follows: S1: Obtain the operating parameters of multiple independent cooling circuits, among which, The operating parameters include the coolant inlet temperature, coolant outlet temperature, coolant flow rate, and heat load of the heat-generating components for each of the cooling circuits. S2: Based on the operating parameters, calculate the ideal coolant flow rate required for each cooling circuit to reach the preset target outlet temperature under the current heat load; S3: Based on the comparison between the ideal coolant flow rate and the actual coolant flow rate, obtain the heat dissipation performance index; S4: Adjust the cooling intensity of each cooling circuit differently according to the heat dissipation performance index of each cooling circuit; S5: Monitor the changing trend of the heat dissipation efficiency index of each cooling circuit, and issue an early warning based on the changing trend; S5 is specifically as follows: S51: Obtain coolant inlet temperature data for multiple independent cooling circuits, and calculate the consensus inlet temperature based on the multiple coolant inlet temperature data; S52: Compare the first deviation between each of the coolant inlet temperature data and the group consensus inlet temperature, and construct an inlet temperature stability index based on the first deviation; S53: When the first deviation exceeds the first preset threshold, the group consensus inlet temperature is used as the calibration coolant inlet temperature of the cooling circuit; S54: Obtain real-time cooling data for each of the cooling circuits, and infer the effective specific heat capacity of the coolant in each of the cooling circuits based on the real-time cooling data; wherein, the real-time cooling data includes the heat load of the heat-generating component, the inlet temperature of the coolant, the outlet temperature of the coolant, and the flow rate of the coolant; S55: Based on the effective specific heat capacity and the preset coolant design specific heat capacity, construct the coolant performance index; S56: Calculate the ideal heat transfer temperature difference of each cooling circuit based on the heat load of the heating component, the coolant flow rate, and the effective specific heat capacity; S57: Based on the ideal heat transfer temperature difference and the second deviation between the coolant outlet temperature and the coolant inlet temperature, calculate the heat transfer efficiency deviation index of each cooling circuit, and construct the flow channel efficiency index based on the heat transfer efficiency deviation index. S58: Obtain the estimated heat load of the heating element; S59: Calculate the actual heat removed based on the coolant flow rate, the effective specific heat capacity, and the second deviation, and correct the estimated heat load based on the actual heat removed to obtain the calibrated heat load; S510: Construct a heat load input reliability index based on the calibrated heat load and the estimated heat load; S511: Using the inlet temperature of the calibration coolant, the effective specific heat capacity, and the calibration heat load as calibration factors, calibrate the heat dissipation efficiency index of each cooling circuit; S512: The inlet temperature stability index, the coolant performance index, the flow channel efficiency index, and the heat load input reliability index are used as cooling health sub-indices that affect the heat dissipation performance index. The changing trends and magnitudes of each cooling health sub-index are analyzed to trace the root causes of the decline in the heat dissipation performance index and to issue warnings for the root causes.

2. The liquid cooling plate temperature monitoring and control method according to claim 1, characterized in that, S2 is specifically as follows: S21: Divide the heating component into multiple sub-regions and obtain the local temperature change rate and local power consumption data of each sub-region; S22: Based on the local temperature change rate and local power consumption data of each sub-region, infer the local heat load of each sub-region; S23: Weighted average of the local heat load of each sub-region to obtain a refined heat load; S24: Using the refined heat load as the input of the heat load of the heat-generating component, calculate the ideal coolant flow rate required for each cooling circuit to reach the preset target outlet temperature under the current heat load.

3. The liquid cooling plate temperature monitoring and control method according to claim 1, characterized in that, The specific details of S512 are as follows: S5121: When each of the cooling health sub-indices exceeds the third preset threshold, abnormal cooling health sub-indices are identified and obtained. S5122: Based on the deviation degree and historical influence weight of each of the abnormal cooling health sub-indices, assess its contribution to the decline of the heat dissipation efficiency index, and obtain the contribution assessment result. S5123: Based on the criticality level of the cooling circuit and the current load status of the heat-generating components, assess the maintenance urgency of each of the aforementioned root causes of the anomaly, and obtain the maintenance urgency assessment results; S5124: Based on the contribution assessment results and the maintenance urgency assessment results, generate a maintenance suggestion list, and issue an early warning for the abnormal root causes according to the maintenance suggestion list; wherein, the maintenance suggestion list includes the dominant root causes, secondary root causes, and maintenance priority ranking.

4. The liquid cooling plate temperature monitoring and control method according to claim 3, characterized in that, The specific details of S5122 are as follows: S51221: Identify the causal relationships between the various abnormal cooling health sub-indices; S51222: Assess the first independent contribution of the abnormal cooling health sub-index with causal relationship, as well as synergistic or inhibitory effects; S51223: Based on the first independent contribution and the synergistic effect or the inhibitory effect, evaluate the contribution of each of the abnormal cooling health sub-indices to the decrease in the heat dissipation efficiency index, and obtain the contribution evaluation result.

5. The liquid cooling plate temperature monitoring and control method according to claim 3, characterized in that, S5123 is as follows: S51231: Identify the types and quantities of currently available maintenance resources, and preliminarily assign maintenance tasks based on the maintenance urgency level of each of the aforementioned anomaly root causes and the types and quantities of maintenance resources required thereto; S51232: When the resource requirements of the initially allocated maintenance task exceed the available maintenance resources or a resource conflict occurs, resource scheduling optimization is performed; The resource scheduling optimization includes: S512321: Prioritize the resource needs of high-urgency tasks that have a significant impact on system stability; S512322: Adjust the execution order and time window of maintenance tasks according to the maintenance urgency level, the contribution of the anomaly root cause, and the estimated maintenance time. S512323: Generate maintenance work orders and resource scheduling plans; S512324: Monitor the maintenance effect and adjust subsequent maintenance plans and resource allocation based on the maintenance effect and actual maintenance progress.

6. The liquid cooling plate temperature monitoring and control method according to claim 5, characterized in that, The specific details of S512324 are as follows: S5123241: Collect the cooling health sub-indices before and after maintenance, and identify the causal relationship between each cooling health sub-indice; S5123242: Evaluate the second independent contribution of each of the aforementioned cooling health sub-indices that have a causal relationship, as well as synergistic or inhibitory effects; S5123243: Based on the second independent contribution and the synergistic effect or the inhibitory effect, calculate the comprehensive contribution of each of the cooling health sub-indices to the decrease in heat dissipation efficiency index; S5123244: Compare the changes in the overall contribution of each of the cooling health sub-indices before and after maintenance, and adjust the subsequent maintenance plan and resource allocation based on the maintenance effect and actual maintenance progress.

7. A liquid-cooled plate temperature monitoring and control system, characterized in that, include: The parameter acquisition module is used to acquire the operating parameters of multiple independent cooling circuits, wherein the operating parameters include the coolant inlet temperature, coolant outlet temperature, coolant flow rate, and heat load of the heat-generating components of each cooling circuit. The efficiency calculation module is used to calculate the ideal coolant flow rate required for each of the cooling circuits to reach the preset target outlet temperature under the current heat load based on the operating parameters; and to obtain the heat dissipation efficiency index based on the comparison between the ideal coolant flow rate and the actual coolant flow rate. An intensity adjustment module is used to adjust the cooling intensity of each cooling circuit differently according to the heat dissipation efficiency index of each cooling circuit. The monitoring and early warning module is used to monitor the changing trends of the heat dissipation efficiency indicators of each cooling circuit and issue early warnings based on the changing trends. The specific operation method of the monitoring and early warning module is as follows: S51: Obtain coolant inlet temperature data for multiple independent cooling circuits, and calculate the consensus inlet temperature based on the multiple coolant inlet temperature data; S52: Compare the first deviation between each of the coolant inlet temperature data and the group consensus inlet temperature, and construct an inlet temperature stability index based on the first deviation; S53: When the first deviation exceeds the first preset threshold, the group consensus inlet temperature is used as the calibration coolant inlet temperature of the cooling circuit; S54: Obtain real-time cooling data for each of the cooling circuits, and infer the effective specific heat capacity of the coolant in each of the cooling circuits based on the real-time cooling data; wherein, the real-time cooling data includes the heat load of the heat-generating component, the inlet temperature of the coolant, the outlet temperature of the coolant, and the flow rate of the coolant; S55: Based on the effective specific heat capacity and the preset coolant design specific heat capacity, construct the coolant performance index; S56: Calculate the ideal heat transfer temperature difference of each cooling circuit based on the heat load of the heating component, the coolant flow rate, and the effective specific heat capacity; S57: Based on the ideal heat transfer temperature difference and the second deviation between the coolant outlet temperature and the coolant inlet temperature, calculate the heat transfer efficiency deviation index of each cooling circuit, and construct the flow channel efficiency index based on the heat transfer efficiency deviation index. S58: Obtain the estimated heat load of the heating element; S59: Calculate the actual heat removed based on the coolant flow rate, the effective specific heat capacity, and the second deviation, and correct the estimated heat load based on the actual heat removed to obtain the calibrated heat load; S510: Construct a heat load input reliability index based on the calibrated heat load and the estimated heat load; S511: Using the inlet temperature of the calibration coolant, the effective specific heat capacity, and the calibration heat load as calibration factors, calibrate the heat dissipation efficiency index of each cooling circuit; S512: The inlet temperature stability index, the coolant performance index, the flow channel efficiency index, and the heat load input reliability index are used as cooling health sub-indices that affect the heat dissipation performance index. The changing trends and magnitudes of each cooling health sub-index are analyzed to trace the root causes of the decline in the heat dissipation performance index and to issue warnings for the root causes.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the liquid cooling plate temperature monitoring and control method as described in any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the liquid cooling plate temperature monitoring and control method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Precise flow control method and system for two-phase cold plate cooling data center

    CN120835514A

  • Dynamic flow control method for two-phase cold plate liquid cooling system based on multi-mode perception

    CN120909404A