Fault monitoring method and device of liquid cooling system and server system
By monitoring the coolant temperature and flow rate of the liquid-cooled system in real time, calculating the current specific heat capacity and issuing an alarm, the problem of inability to monitor the coolant quality in the existing technology is solved, and the timely replacement of the coolant and the guarantee of heat dissipation efficiency are achieved.
Patent Information
- Application Number
- CN202510081660.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-01-17
AI Technical Summary
The prior art cannot fully cover various faults of liquid cooling systems, especially the inability to monitor the quality of the coolant in real time, resulting in inconvenient maintenance. Problems may only be discovered if there are serious problems in the system.
By obtaining the inlet, outlet temperature and flow rate of the coolant of the liquid cooling system in real time, calculate the current specific heat capacity, and issue an alarm based on the specific heat capacity threshold to remind the coolant to deteriorate. At the same time, a specific heat capacity-time curve is generated to predict the remaining service life of the coolant, and a warning is issued before the deterioration date.
Real-time monitoring of the quality of coolant is realized, and the deteriorated coolant can be replaced in time without shutdown inspection, to prevent the reduction of heat dissipation efficiency and cause machine failure, and solve the problem of the inability to check the quality of coolant in the prior art without shutting down.
Smart Images

Figure CN119958889A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the computer field, and in particular, to a fault monitoring method, device, computer-readable storage medium, and server system for a liquid cooling system. Background Art
[0002] Currently, most servers use air cooling design, about 10-20% use cold plate liquid cooling design, and the share of immersion and spray liquid cooling is almost negligible. The early warning and alarm of the current air cooling design's heat dissipation system are relatively complete, covering every link from cold source to temperature, but for the liquid cooling design, there are almost no other monitoring items except monitoring whether there is leakage.
[0003] Currently, liquid cooling only monitors and predicts leakage, usually using a leakage detection line. When leakage occurs, liquid reaches the leakage detection line, causing the level of the leakage detection line to change. The level change is monitored by the BMC, and the BMC issues an early warning and alarm for the leakage.
[0004] It can be seen that the existing technology cannot fully cover various faults of the liquid cooling system. There is no way to capture a large amount of alarm information other than leakage. At the same time, there is no good way to deal with situations such as coolant deterioration, which is not conducive to the maintenance of the liquid cooling system. Users may only be able to detect problems with the liquid cooling system after a real problem such as overheating or burning of the board occurs. Summary of the invention
[0005] The embodiments of the present application provide a liquid cooling system fault monitoring method, device, computer-readable storage medium and server system to at least solve the problem in the prior art that the quality of the cooling liquid cannot be checked without shutting down the system.
[0006] According to one embodiment of the present application, a fault monitoring method for a liquid cooling system is provided, including: step S102, acquiring inlet temperature, outlet temperature and flow rate of the coolant of the liquid cooling system in real time, the inlet temperature being the temperature of the coolant when entering the liquid cooling system, and the outlet temperature being the temperature of the coolant when flowing out of the liquid cooling system; step S104, calculating the current specific heat capacity of the coolant in the liquid cooling system according to the inlet temperature, the outlet temperature and the flow rate; step S106, when the difference between the initial specific heat capacity of the coolant and the current specific heat capacity is greater than or equal to a specific heat capacity threshold, issuing a first alarm to remind the liquid cooling system that the coolant has deteriorated.
[0007] In an exemplary embodiment, when the difference between the initial specific heat capacity of the coolant and the current specific heat capacity is less than the specific heat capacity threshold, the method further includes: recording the current time, and repeating the steps S102 and S104 in sequence once at intervals of a first predetermined time to obtain at least one detected specific heat capacity and a corresponding specific heat capacity detection time; generating a coolant specific heat capacity-time curve according to the current specific heat capacity, the current time, at least one detected specific heat capacity and the corresponding specific heat capacity detection time, the coolant specific heat capacity-time curve being a curve showing that the specific heat capacity of the coolant changes with increasing usage time; predicting the remaining service life of the coolant according to the coolant specific heat capacity-time curve, the initial specific heat capacity and the specific heat capacity threshold, the remaining service life being the shortest time from the current time to the time when the coolant deteriorates; determining the deterioration date of the coolant according to the remaining service life, and issuing a pre-warning before the deterioration date of the coolant to remind the user to replace the coolant of the liquid cooling system.
[0008] In an exemplary embodiment, when the difference between the initial specific heat capacity and the current specific heat capacity of the coolant is greater than or equal to the specific heat capacity threshold, after issuing the first alarm, the method further includes: controlling the filtering device of the liquid cooling system to filter the coolant; when the cooling liquid is filtered, repeating the steps S102 and S104 once in sequence to obtain an updated specific heat capacity; when the difference between the initial specific heat capacity and the updated specific heat capacity is less than the specific heat capacity threshold, withdrawing the first alarm; when the difference between the initial specific heat capacity and the updated specific heat capacity is greater than or equal to the specific heat capacity threshold, not withdrawing the first alarm to remind the liquid cooling system that the coolant has deteriorated.
[0009] In an exemplary embodiment, when the difference between the initial specific heat capacity and the updated specific heat capacity is less than the specific heat capacity threshold, after withdrawing the first alarm, the method further includes: updating the remaining service life of the coolant to obtain the latest remaining service life; determining the deterioration date of the coolant according to the latest remaining service life, and issuing a pre-alarm before the deterioration date of the coolant to remind replacement of the coolant of the liquid cooling system.
[0010] In an exemplary embodiment, before step S102, the method further includes: acquiring in real time an inlet pressure and an outlet pressure of the coolant of the liquid cooling system, the inlet pressure being the hydraulic pressure when the coolant enters the liquid cooling system, and the outlet pressure being the hydraulic pressure when the coolant flows out of the liquid cooling system; and issuing a second alarm when the difference between the inlet pressure and the outlet pressure is greater than or equal to a pressure difference threshold to remind that the liquid cooling system is clogged.
[0011] In an exemplary embodiment, when the difference between the liquid inlet pressure and the liquid outlet pressure is greater than or equal to a pressure difference threshold, before issuing a second alarm, the method further includes: when the liquid inlet pressure is less than a predetermined liquid inlet pressure, issuing a third alarm to remind that the coolant inlet pipe of the liquid cooling system is clogged, and the predetermined liquid inlet pressure is the minimum liquid inlet pressure of the liquid cooling system.
[0012] In an exemplary embodiment, when the difference between the liquid inlet pressure and the liquid outlet pressure is greater than or equal to the pressure difference threshold, after issuing a second alarm, the method further includes: increasing the liquid inlet pressure of the liquid cooling system; reacquiring the liquid inlet pressure and the liquid outlet pressure at an interval of a second predetermined time to obtain an updated liquid inlet pressure and an updated liquid outlet pressure; when the difference between the updated liquid inlet pressure and the updated liquid outlet pressure is less than the pressure difference threshold, withdrawing the second alarm; when the difference between the updated liquid inlet pressure and the updated liquid outlet pressure is greater than or equal to the pressure difference threshold, starting a backup cooling system to replace the liquid cooling system for cooling.
[0013] According to another embodiment of the present application, a fault monitoring device for a liquid cooling system is provided, including: a first acquisition module, used to execute step S102, to acquire in real time the inlet temperature, outlet temperature and flow rate of the coolant in the liquid cooling system, the inlet temperature being the temperature of the coolant when entering the liquid cooling system, and the outlet temperature being the temperature of the coolant when flowing out of the liquid cooling system; a calculation module, used to execute step S104, to calculate the current specific heat capacity of the coolant in the liquid cooling system according to the inlet temperature, the outlet temperature and the flow rate; a first alarm module, used to execute step S102, to issue a first alarm when the difference between the initial specific heat capacity of the coolant and the current specific heat capacity is greater than or equal to a specific heat capacity threshold, so as to remind the liquid cooling system that the coolant has deteriorated.
[0014] According to another embodiment of the present application, a computer-readable storage medium is provided, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above method embodiments when run.
[0015] According to another embodiment of the present application, a server system is provided, including a server, a liquid cooling system, a memory and a processor, wherein a computer program is stored in the memory, the liquid cooling system is used to dissipate heat from the server, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0016] Through the present application, the above method can calculate the current specific heat capacity of the coolant in real time by real-time monitoring the inlet temperature, outlet temperature and flow rate of the coolant in the liquid cooling system. If the difference between the initial specific heat capacity and the current specific heat capacity of the coolant is greater than or equal to the specific heat capacity threshold, indicating that the coolant has deteriorated, a first alarm is issued to achieve real-time monitoring of the coolant quality. There is no need to stop the machine to sample the coolant to check whether it has deteriorated. The deteriorated coolant can be replaced in time to prevent the heat dissipation efficiency from being reduced and causing machine failure, solving the problem in the prior art that the quality of the coolant cannot be checked without stopping the machine. In addition, since the specific heat capacity cannot be measured directly, the prior art does not have a standard for checking the quality of the coolant. It can only check whether the coolant has impurities. The inspection process is complicated and it is impossible to directly determine whether the coolant has deteriorated. It is usually replaced regularly, resulting in a large amount of coolant that has not deteriorated being wasted. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 A hardware structure block diagram of a terminal for executing a fault monitoring method for a liquid cooling system provided in an embodiment of the present application is shown;
[0018] Figure 2 is a flow chart of a fault monitoring method for a liquid cooling system according to an embodiment of the present application;
[0019] Figure 3 is a structural block diagram of a server system including a liquid cooling system according to an embodiment of the present application;
[0020] Figure 4 It is a structural schematic diagram of a fault monitoring device for a liquid cooling system according to an embodiment of the present application. DETAILED DESCRIPTION
[0021] The embodiments of the present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0022] It should be noted that the terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.
[0023] For the convenience of description, some nouns or terms involved in the embodiments of the present application are explained below:
[0024] BMC: Baseboard Management Controller, the full name of which is Baseboard Management Controller.
[0025] Cold plate liquid cooling system: abbreviated as cold plate system, which indirectly transfers the heat of the heating device to the cooling liquid enclosed in the circulation pipeline through a cold plate (usually a closed cavity made of heat-conducting metals such as copper and aluminum), and then takes the heat away through the cooling liquid;
[0026] Immersion and spray liquid cooling: The components are in direct contact with the coolant, and heat is dissipated by overall immersion or liquid spraying;
[0027] Air cooling: using fans to dissipate heat from heat-generating components;
[0028] CDU: The full name of CDU is Coolant Distribution Unit, which is a system used to distribute cooling liquid between liquid-cooled electronic equipment. It provides secondary side flow distribution, pressure control, physical isolation, anti-condensation and other functions.
[0029] Specific Heat Capacity: Indicated by the symbol c, also known as specific heat capacity, abbreviated as specific heat, is the heat capacity of a unit mass of a substance, that is, the heat absorbed or released when a unit mass of an object changes unit temperature.
[0030] The method embodiments provided in the embodiments of the present application can be executed in a computer device or a similar computing device. Taking running on a server device as an example, Figure 1 FIG. 1 is a hardware structure block diagram of a server device of a method for monitoring a fault of a liquid cooling system according to an embodiment of the present application. Figure 1 As shown, the computer device may include one or more ( Figure 1 Only one is shown in the figure) a processor 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data, wherein the above-mentioned computer device may also include a transmission device 106 and an input and output device 108 for communication functions. It can be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above-mentioned computer device. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations shown.
[0031] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the fault monitoring method of the liquid cooling system in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, to implement the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories may be connected to the server device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0032] The transmission device 106 is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of the computer device. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, referred to as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0033] In this embodiment, a fault monitoring method for a liquid cooling system is provided. Figure 2 is a flow chart of a fault monitoring method for a liquid cooling system according to an embodiment of the present application. Figure 2 As shown, the process includes the following steps:
[0034] Step S102, obtaining inlet temperature, outlet temperature and flow rate of the coolant of the liquid cooling system in real time, wherein the inlet temperature is the temperature of the coolant when it enters the liquid cooling system, and the outlet temperature is the temperature of the coolant when it flows out of the liquid cooling system;
[0035] Step S104, calculating the current specific heat capacity of the coolant in the liquid cooling system according to the liquid inlet temperature, the liquid outlet temperature and the flow rate;
[0036] Step S106, when the difference between the initial specific heat capacity of the coolant and the current specific heat capacity is greater than or equal to the specific heat capacity threshold, a first alarm is issued to remind the liquid cooling system that the coolant is deteriorating.
[0037] Through the above steps, the inlet temperature, outlet temperature and flow rate of the coolant in the liquid cooling system are monitored in real time, and the current specific heat capacity of the coolant can be calculated in real time. If the difference between the initial specific heat capacity and the current specific heat capacity of the coolant is greater than or equal to the specific heat capacity threshold, indicating that the coolant has deteriorated, the first alarm is issued to achieve real-time monitoring of the coolant quality. There is no need to stop the machine to sample the coolant to check whether it has deteriorated. The deteriorated coolant can be replaced in time to prevent the heat dissipation efficiency from being reduced and causing machine failure, solving the problem in the prior art that the quality of the coolant cannot be checked without stopping the machine. In addition, since the specific heat capacity cannot be measured directly, the prior art does not have a standard for checking the quality of the coolant. It can only check whether the coolant has impurities. The inspection process is complicated and it is impossible to directly determine whether the coolant has deteriorated. It is usually replaced regularly, resulting in a large amount of coolant that has not deteriorated being wasted.
[0038] The execution subject of the above steps may be BMC, etc., but is not limited thereto.
[0039] In addition, the hardware system needs to add sensors to detect the inlet temperature, outlet temperature, flow rate, hydraulic pressure, etc. of the coolant in the cold plate, including temperature sensors, pressure sensors and flow sensors. The installation position of the sensors is as follows: Figure 3 As shown in the figure, the solid line is the transmission line of the coolant, and the dotted line is the transmission line of the sensor data. These sensors are connected to the management controller BMC. The BMC is responsible for collecting sensor data, processing the data according to various strategies, and analyzing the status of the liquid cooling system. BMC can calculate the specific heat capacity of the current coolant by calculating the power consumption and the problem of inlet and outlet liquid. If the specific heat capacity is not much different from the specific heat capacity of the initial new coolant, the quality of the coolant is good and has not deteriorated. If the specific heat capacity drops significantly, the quality of the coolant may have dropped. The calculation formula and principle are as follows: 1. First calculate the mass of coolant flowing through per unit time based on the flow rate and the initial density of the coolant: m = Q (flow rate per second) * ρ (initial density of coolant); 2. BMC reads the power P of the current heating component, and reads the liquid inlet temperature ti and the outlet temperature to at the same time; 3. The power P is the number of joules of the heating component per unit time, and the current specific heat capacity of the coolant is calculated according to the specific heat capacity formula, c = P / (Q (flow rate per second) * ρ (initial density of coolant) * (to-ti)), and the specific heat capacity is compared with the standard specific heat capacity of the coolant. If the difference is too large, the BMC will alarm. At the same time, in order to adapt to different types of coolants, BMC should also provide a coolant type setting interface to compare the specific heat capacity calculated by BMC with the specific heat capacity of the coolant. For example, when the coolant is water, the specific heat capacity of water is 4.2×10 3 J / (kg×℃), ethylene glycol is 4.0×10 3 J / (kg×℃), when the BMC is in the water coolant mode, the current coolant specific heat capacity calculated in step 3 must be equal to 4.2×10 3When the difference is too large, BMC will issue an alarm.
[0040] In order to replace the coolant in time, in an optional implementation manner, when the difference between the initial specific heat capacity of the coolant and the current specific heat capacity is less than the specific heat capacity threshold, the method further includes:
[0041] Step S202, recording the current time, and repeating the above steps S102 and S104 once in sequence every first predetermined time interval, to obtain at least one detected specific heat capacity and a corresponding specific heat capacity detection time;
[0042] Step S204, generating a coolant specific heat capacity-time curve according to the current specific heat capacity, the current time, at least one of the detected specific heat capacities and the corresponding specific heat capacity detection time, wherein the coolant specific heat capacity-time curve is a curve showing that the specific heat capacity of the coolant changes with the increase of the usage time;
[0043] Step S206, predicting the remaining service life of the coolant according to the coolant specific heat capacity-time curve, the initial specific heat capacity and the specific heat capacity threshold, the remaining service life being the shortest time from the current moment to the moment when the coolant deteriorates;
[0044] Step S208, determining the deterioration date of the coolant according to the remaining service life, and issuing a pre-warning before the deterioration date of the coolant to remind the replacement of the coolant of the liquid cooling system.
[0045] In the above embodiment, sensor data is acquired once at every first predetermined time interval, thereby detecting the specific heat capacity of the coolant once. The specific heat capacity detected multiple times and the corresponding specific heat capacity detection moments can generate a coolant specific heat capacity-time curve. If the specific heat capacity of the coolant transforms linearly with time, the remaining service life of the coolant can be predicted based on the slope of the specific heat capacity change. Of course, if it does not transform linearly, the remaining service life of the coolant can be predicted by a fitted curve. When the specific heat capacity of the coolant does not change to the alarm level, the deterioration date of the coolant can be determined based on the predicted remaining service life, and a pre-alarm can be issued before the actual alarm to remind the user to pay attention to the quality of the coolant and replace the coolant in time to ensure the heat dissipation efficiency and avoid overheating failures of components that require heat dissipation.
[0046] It should be noted that the above sensors measure multiple times each time they collect data, and calculate the average value as the collected data to prevent accidental errors from causing false alarms.
[0047] In order to save coolant, in an optional implementation manner, when the difference between the initial specific heat capacity of the coolant and the current specific heat capacity is greater than or equal to the specific heat capacity threshold, after issuing the first alarm, the method further includes:
[0048] Step S302, controlling the filtering device of the liquid cooling system to filter the cooling liquid;
[0049] Step S304, when the cooling liquid is filtered, repeating the steps S102 and S104 one by one, to obtain an updated specific heat capacity;
[0050] Step S306, when the difference between the initial specific heat capacity and the updated specific heat capacity is less than the specific heat capacity threshold, withdrawing the first alarm;
[0051] Step S308: When the difference between the initial specific heat capacity and the updated specific heat capacity is greater than or equal to the specific heat capacity threshold, the first alarm is not withdrawn to remind the liquid cooling system that the coolant is deteriorating.
[0052] In the above implementation mode, the reduction in the specific heat capacity of the coolant may be caused by impurities mixed in. Directly replacing the coolant causes great waste. The coolant is first filtered through a filtering device, and the specific heat capacity is retested after filtering to obtain an updated specific heat capacity. If the difference between the initial specific heat capacity and the updated specific heat capacity is less than the specific heat capacity threshold, it indicates that filtration can remove impurities and increase the specific heat capacity. There is no need to replace the coolant, and the first alarm can be withdrawn. Otherwise, filtration cannot remove impurities. For example, water vapor mixed in cannot be filtered out, and the coolant has deteriorated. The first alarm is not withdrawn to remind the user to replace the coolant as soon as possible.
[0053] In order to replace the coolant in time, in an optional implementation manner, when the difference between the initial specific heat capacity and the updated specific heat capacity is less than the specific heat capacity threshold, after the first alarm is withdrawn, the method further includes:
[0054] Step S402, updating the remaining service life of the coolant to obtain the latest remaining service life;
[0055] Step S404, determining the deterioration date of the coolant according to the latest remaining service life, and issuing a pre-warning before the deterioration date of the coolant to remind the replacement of the coolant of the liquid cooling system.
[0056] In the above implementation mode, after the coolant is filtered, if the specific heat capacity meets the standard, there is no need to replace the coolant, and the remaining service life of the coolant is updated to obtain the latest remaining service life, thereby updating the deterioration date of the coolant, and issuing a timely warning to replace the coolant to avoid replacing the coolant too early and wasting the service life of the coolant.
[0057] In order to monitor the blocking fault in real time, in an optional implementation manner, before the above step S102, the above method further includes:
[0058] Step S502, obtaining inlet pressure and outlet pressure of the coolant of the liquid cooling system in real time, wherein the inlet pressure is the hydraulic pressure when the coolant enters the liquid cooling system, and the outlet pressure is the hydraulic pressure when the coolant flows out of the liquid cooling system;
[0059] Step S504: When the difference between the liquid inlet pressure and the liquid outlet pressure is greater than or equal to the pressure difference threshold, a second alarm is issued to remind that the liquid cooling system is clogged.
[0060] In the above embodiment, the BMC obtains the inlet pressure Fi and outlet pressure Fo of the coolant of the above liquid cooling system in real time through the pressure sensor. In a normal liquid cooling system, the coolant is almost incompressible. According to Pascal's law, Fi and Fo should be basically equal. If the difference between Fi and Fo is too large, there is a blockage between the Fi and Fo sensors, and a second alarm is issued to remind the above liquid cooling system to be blocked.
[0061] In order to narrow the scope of investigation of the blockage position, in an optional implementation manner, when the difference between the liquid inlet pressure and the liquid outlet pressure is greater than or equal to the pressure difference threshold, before issuing the second alarm, the method further includes:
[0062] Step S602, when the above-mentioned liquid inlet pressure is less than the predetermined liquid inlet pressure, a third alarm is issued to remind that the coolant inlet pipe of the above-mentioned liquid cooling system is blocked, and the above-mentioned predetermined liquid inlet pressure is the minimum liquid inlet pressure of the above-mentioned liquid cooling system.
[0063] In the above embodiment, if the BMC obtains that the liquid inlet pressure Fi is too small, there is a blockage in front of the Fi sensor, and a third alarm is issued to remind the coolant inlet pipe of the above liquid cooling system to be blocked, thereby narrowing the scope of investigation of the blockage position.
[0064] In order to facilitate solving the blockage fault, in an optional implementation manner, when the difference between the liquid inlet pressure and the liquid outlet pressure is greater than or equal to the pressure difference threshold, after issuing the second alarm, the method further includes:
[0065] Step S702, increasing the liquid inlet pressure of the liquid cooling system;
[0066] Step S704, reacquiring the liquid inlet pressure and the liquid outlet pressure at a second predetermined time interval to obtain an updated liquid inlet pressure and an updated liquid outlet pressure;
[0067] Step S706, when the difference between the updated liquid inlet pressure and the updated liquid outlet pressure is less than the pressure difference threshold, canceling the second alarm;
[0068] Step S708: When the difference between the updated liquid inlet pressure and the updated liquid outlet pressure is greater than or equal to the pressure difference threshold, the backup cooling system is started to replace the liquid cooling system for cooling.
[0069] In the above implementation mode, sometimes the blockage of the liquid cooling system is not serious. The blockage can be unblocked by appropriately increasing the liquid inlet pressure of the liquid cooling system without stopping for inspection and maintenance, which greatly reduces the maintenance workload. The liquid pressure and the liquid outlet pressure are updated at a second predetermined time interval. If the difference between the updated liquid inlet pressure and the updated liquid outlet pressure is less than the pressure difference threshold, it indicates that the blockage has been unblocked and the second alarm can be withdrawn. Otherwise, the blockage cannot be unblocked by simply increasing the liquid inlet pressure of the liquid cooling system. The standby cooling system can be started to replace the liquid cooling system for cooling until the liquid cooling system is repaired and the blockage is unblocked to maintain uninterrupted heat dissipation.
[0070] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, disk, CD), and includes a number of instructions for a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the above methods of each embodiment of the present application.
[0071] In this embodiment, a fault monitoring device for a liquid cooling system is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions that have been made will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.
[0072] Figure 4 is a structural block diagram of a fault monitoring device for a liquid cooling system according to an embodiment of the present application, such as Figure 4 As shown, the device includes
[0073] The first acquisition module 22 is used to execute step S102 to acquire the inlet temperature, outlet temperature and flow rate of the coolant of the liquid cooling system in real time, wherein the inlet temperature is the temperature of the coolant when it enters the liquid cooling system, and the outlet temperature is the temperature of the coolant when it flows out of the liquid cooling system;
[0074] The calculation module 24 is used to execute step S104, and calculate the current specific heat capacity of the coolant in the liquid cooling system according to the liquid inlet temperature, the liquid outlet temperature and the flow rate;
[0075] The first alarm module 26 is used to execute step S106, and when the difference between the initial specific heat capacity of the above-mentioned coolant and the above-mentioned current specific heat capacity is greater than or equal to the specific heat capacity threshold, issue a first alarm to remind the above-mentioned liquid cooling system that the above-mentioned coolant is deteriorating.
[0076] Through the above module, the inlet temperature, outlet temperature and flow rate of the coolant in the liquid cooling system are monitored in real time, and the current specific heat capacity of the coolant can be calculated in real time. If the difference between the initial specific heat capacity and the current specific heat capacity of the coolant is greater than or equal to the specific heat capacity threshold, it indicates that the coolant has deteriorated, and the first alarm is issued to achieve real-time monitoring of the coolant quality. There is no need to stop the machine to sample the coolant to check whether it has deteriorated. The deteriorated coolant can be replaced in time to prevent the heat dissipation efficiency from being reduced and causing machine failure, which solves the problem of the inability to check the quality of the coolant without stopping the machine in the prior art. In addition, since the specific heat capacity cannot be measured directly, the prior art does not have a standard for checking the quality of the coolant. It can only check whether the coolant has impurities. The inspection process is complicated and it is impossible to directly determine whether the coolant has deteriorated. It is usually replaced regularly, resulting in a large amount of coolant that has not deteriorated being wasted.
[0077] The execution subject of the above steps may be BMC, etc., but is not limited thereto.
[0078] In addition, the hardware system needs to add sensors to detect the inlet temperature, outlet temperature, flow rate, hydraulic pressure, etc. of the coolant in the cold plate, including temperature sensors, pressure sensors and flow sensors. The installation position of the sensors is as follows: Figure 3As shown in the figure, the solid line is the transmission line of the coolant, and the dotted line is the transmission line of the sensor data. These sensors are connected to the management controller BMC. The BMC is responsible for collecting sensor data, processing the data according to various strategies, and analyzing the status of the liquid cooling system. BMC can calculate the specific heat capacity of the current coolant by calculating the power consumption and the problem of inlet and outlet liquid. If the specific heat capacity is not much different from the specific heat capacity of the initial new coolant, the quality of the coolant is good and has not deteriorated. If the specific heat capacity drops significantly, the quality of the coolant may have dropped. The calculation formula and principle are as follows: 1. First calculate the mass of coolant flowing through per unit time based on the flow rate and the initial density of the coolant: m = Q (flow rate per second) * ρ (initial density of coolant); 2. BMC reads the power P of the current heating component, and reads the liquid inlet temperature ti and the outlet temperature to at the same time; 3. The power P is the number of joules of the heating component per unit time, and the current specific heat capacity of the coolant is calculated according to the specific heat capacity formula, c = P / (Q (flow rate per second) * ρ (initial density of coolant) * (to-ti)), and the specific heat capacity is compared with the standard specific heat capacity of the coolant. If the difference is too large, the BMC will alarm. At the same time, in order to adapt to different types of coolants, BMC should also provide a coolant type setting interface to compare the specific heat capacity calculated by BMC with the specific heat capacity of the coolant. For example, when the coolant is water, the specific heat capacity of water is 4.2×10 3 J / (kg×℃), ethylene glycol is 4.0×10 3 J / (kg×℃), when the BMC is in the water coolant mode, the current coolant specific heat capacity calculated in step 3 must be equal to 4.2×10 3 When the difference is too large, BMC will issue an alarm.
[0079] In order to replace the coolant in time, in an optional implementation manner, the above device further includes:
[0080] A first repeating module is used for recording the current time when the difference between the initial specific heat capacity of the coolant and the current specific heat capacity is less than the specific heat capacity threshold value, and repeating the steps S102 and S104 once at intervals of a first predetermined time, to obtain at least one detected specific heat capacity and a corresponding specific heat capacity detection time;
[0081] A generating module, for generating a coolant specific heat capacity-time curve according to the current specific heat capacity, the current time, at least one of the detected specific heat capacities and the corresponding specific heat capacity detection time, wherein the coolant specific heat capacity-time curve is a curve showing that the specific heat capacity of the coolant changes with the increase of the usage time;
[0082] A prediction module, for predicting the remaining service life of the coolant according to the coolant specific heat capacity-time curve, the initial specific heat capacity and the specific heat capacity threshold, wherein the remaining service life is the shortest time from the current moment to the moment when the coolant deteriorates;
[0083] The first determination module is used to determine the deterioration date of the coolant according to the remaining service life, and issue a pre-warning before the deterioration date of the coolant to remind the replacement of the coolant of the liquid cooling system.
[0084] In the above embodiment, sensor data is acquired once at every first predetermined time interval, thereby detecting the specific heat capacity of the coolant once. The specific heat capacity detected multiple times and the corresponding specific heat capacity detection moments can generate a coolant specific heat capacity-time curve. If the specific heat capacity of the coolant transforms linearly with time, the remaining service life of the coolant can be predicted based on the slope of the specific heat capacity change. Of course, if it does not transform linearly, the remaining service life of the coolant can be predicted by a fitted curve. When the specific heat capacity of the coolant does not change to the alarm level, the deterioration date of the coolant can be determined based on the predicted remaining service life, and a pre-alarm can be issued before the actual alarm to remind the user to pay attention to the quality of the coolant and replace the coolant in time to ensure the heat dissipation efficiency and avoid overheating failures of components that require heat dissipation.
[0085] It should be noted that the above sensors measure multiple times each time they collect data, and calculate the average value as the collected data to prevent accidental errors from causing false alarms.
[0086] In order to save coolant, in an optional implementation, the device further includes:
[0087] A first control module is used to control the filtering device of the liquid cooling system to filter the coolant after issuing a first alarm when the difference between the initial specific heat capacity of the coolant and the current specific heat capacity is greater than or equal to a specific heat capacity threshold;
[0088] A second repeating module is used to repeat the above steps S102 and S104 one time in sequence when the above cooling liquid is filtered, so as to obtain an updated specific heat capacity;
[0089] A first withdrawal module, configured to withdraw the first alarm when the difference between the initial specific heat capacity and the updated specific heat capacity is less than the specific heat capacity threshold;
[0090] The reminder module is used for not withdrawing the first alarm when the difference between the initial specific heat capacity and the updated specific heat capacity is greater than or equal to the specific heat capacity threshold, so as to remind the liquid cooling system that the coolant has deteriorated.
[0091] In the above implementation mode, the reduction in the specific heat capacity of the coolant may be caused by impurities mixed in. Directly replacing the coolant causes great waste. The coolant is first filtered through a filtering device, and the specific heat capacity is retested after filtering to obtain an updated specific heat capacity. If the difference between the initial specific heat capacity and the updated specific heat capacity is less than the specific heat capacity threshold, it indicates that filtration can remove impurities and increase the specific heat capacity. There is no need to replace the coolant, and the first alarm can be withdrawn. Otherwise, filtration cannot remove impurities. For example, water vapor mixed in cannot be filtered out, and the coolant has deteriorated. The first alarm is not withdrawn to remind the user to replace the coolant as soon as possible.
[0092] In order to replace the coolant in time, in an optional implementation manner, the above device further includes:
[0093] A first updating module is used for updating the remaining service life of the coolant to obtain the latest remaining service life after withdrawing the first alarm when the difference between the initial specific heat capacity and the updated specific heat capacity is less than the specific heat capacity threshold;
[0094] The second determination module is used to determine the deterioration date of the above-mentioned coolant according to the above-mentioned latest remaining service life, and to issue a pre-warning before the deterioration date of the above-mentioned coolant to remind the replacement of the above-mentioned coolant of the above-mentioned liquid cooling system.
[0095] In the above implementation mode, after the coolant is filtered, if the specific heat capacity meets the standard, there is no need to replace the coolant, and the remaining service life of the coolant is updated to obtain the latest remaining service life, thereby updating the deterioration date of the coolant, and issuing a timely warning to replace the coolant to avoid replacing the coolant too early and wasting the service life of the coolant.
[0096] In order to monitor the blocking fault in real time, in an optional implementation manner, the above device further includes:
[0097] A second acquisition module is used to obtain the inlet pressure and outlet pressure of the coolant of the liquid cooling system in real time before obtaining the inlet temperature, outlet temperature and flow rate of the coolant of the liquid cooling system in real time, wherein the inlet pressure is the hydraulic pressure when the coolant enters the liquid cooling system, and the outlet pressure is the hydraulic pressure when the coolant flows out of the liquid cooling system;
[0098] The second alarm module is used to issue a second alarm when the difference between the liquid inlet pressure and the liquid outlet pressure is greater than or equal to the pressure difference threshold to remind that the liquid cooling system is blocked.
[0099] In the above embodiment, the BMC obtains the inlet pressure Fi and outlet pressure Fo of the coolant of the above liquid cooling system in real time through the pressure sensor. In a normal liquid cooling system, the coolant is almost incompressible. According to Pascal's law, Fi and Fo should be basically equal. If the difference between Fi and Fo is too large, there is a blockage between the Fi and Fo sensors, and a second alarm is issued to remind the above liquid cooling system to be blocked.
[0100] In order to narrow the scope of investigation of the blockage position, in an optional implementation manner, the above device further includes:
[0101] The third alarm module is used to issue a third alarm when the difference between the above-mentioned liquid inlet pressure and the above-mentioned liquid outlet pressure is greater than or equal to the pressure difference threshold, and before issuing the second alarm, when the above-mentioned liquid inlet pressure is less than the preset liquid inlet pressure, to remind that the coolant inlet pipe of the above-mentioned liquid cooling system is blocked, and the above-mentioned preset liquid inlet pressure is the minimum liquid inlet pressure of the above-mentioned liquid cooling system.
[0102] In the above embodiment, if the BMC obtains that the liquid inlet pressure Fi is too small, there is a blockage in front of the Fi sensor, and a third alarm is issued to remind the coolant inlet pipe of the above liquid cooling system to be blocked, thereby narrowing the scope of investigation of the blockage position.
[0103] In order to facilitate solving the blocking fault, in an optional implementation manner, the above device further includes:
[0104] A second control module is used to increase the liquid inlet pressure of the liquid cooling system after issuing a second alarm when the difference between the liquid inlet pressure and the liquid outlet pressure is greater than or equal to a pressure difference threshold;
[0105] A second updating module, configured to re-acquire the liquid inlet pressure and the liquid outlet pressure at intervals of a second predetermined time to obtain an updated liquid inlet pressure and an updated liquid outlet pressure;
[0106] A second withdrawal module, configured to withdraw the second alarm when the difference between the updated liquid inlet pressure and the updated liquid outlet pressure is less than the pressure difference threshold;
[0107] The third control module is used to start the backup cooling system to replace the liquid cooling system for cooling work when the difference between the updated liquid inlet pressure and the updated liquid outlet pressure is greater than or equal to the pressure difference threshold.
[0108] In the above implementation mode, sometimes the blockage of the liquid cooling system is not serious. The blockage can be unblocked by appropriately increasing the liquid inlet pressure of the liquid cooling system without stopping for inspection and maintenance, which greatly reduces the maintenance workload. The liquid pressure and the liquid outlet pressure are updated at a second predetermined time interval. If the difference between the updated liquid inlet pressure and the updated liquid outlet pressure is less than the pressure difference threshold, it indicates that the blockage has been unblocked and the second alarm can be withdrawn. Otherwise, the blockage cannot be unblocked by simply increasing the liquid inlet pressure of the liquid cooling system. The standby cooling system can be started to replace the liquid cooling system for cooling until the liquid cooling system is repaired and the blockage is unblocked to maintain uninterrupted heat dissipation.
[0109] It should be noted that the above modules can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited to: the above modules are all located in the same processor; or the above modules are located in different processors in any combination.
[0110] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above method embodiments when running.
[0111] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0112] An embodiment of the present application also provides a server system, including a server, a liquid cooling system, a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the liquid cooling system is used to dissipate heat for the server, and the processor is configured to run the computer program to execute the steps in any one of the method embodiments.
[0113] In an exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.
[0114] For specific examples in this embodiment, reference may be made to the examples described in the above embodiments and exemplary implementation modes, and this embodiment will not be described in detail herein.
[0115] Obviously, those skilled in the art should understand that the above modules or steps of the present application can be implemented by a general computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, and in some cases, the steps shown or described can be executed in a different order from that herein, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. Thus, the present application is not limited to any specific combination of hardware and software.
[0116] The above description is only the preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the principles of the present application shall be included in the protection scope of the present application.
Claims
1. A fault monitoring method for a liquid cooling system, characterized in that: include: Step S102, obtaining inlet temperature, outlet temperature and flow rate of the coolant of the liquid cooling system in real time, wherein the inlet temperature is the temperature of the coolant when it enters the liquid cooling system, and the outlet temperature is the temperature of the coolant when it flows out of the liquid cooling system; Step S104, calculating the current specific heat capacity of the coolant in the liquid cooling system according to the liquid inlet temperature, the liquid outlet temperature and the flow rate; Step S106, when the difference between the initial specific heat capacity of the coolant and the current specific heat capacity is greater than or equal to a specific heat capacity threshold, a first alarm is issued to remind the liquid cooling system that the coolant is deteriorating.
2. The method according to claim 1, characterized in that In the case where the difference between the initial specific heat capacity of the coolant and the current specific heat capacity is less than the specific heat capacity threshold, the method further includes: Record the current time, and repeat the step S102 and the step S104 once in sequence every first predetermined time interval to obtain at least one detected specific heat capacity and a corresponding specific heat capacity detection time; Generate a coolant specific heat capacity-time curve according to the current specific heat capacity, the current time, at least one of the detected specific heat capacities and the corresponding specific heat capacity detection time, wherein the coolant specific heat capacity-time curve is a curve showing that the specific heat capacity of the coolant changes with the increase of the usage time; Predicting the remaining service life of the coolant according to the coolant specific heat capacity-time curve, the initial specific heat capacity and the specific heat capacity threshold, the remaining service life being the shortest time from the current moment to the moment when the coolant deteriorates; The deterioration date of the coolant is determined according to the remaining service life, and a pre-warning is issued before the deterioration date of the coolant to remind the user to replace the coolant of the liquid cooling system.
3. The method according to claim 1, characterized in that In the case where the difference between the initial specific heat capacity of the coolant and the current specific heat capacity is greater than or equal to the specific heat capacity threshold, after issuing the first alarm, the method further includes: Controlling the filtering device of the liquid cooling system to filter the cooling liquid; When the cooling liquid is filtered, the step S102 and the step S104 are repeated once to obtain an updated specific heat capacity; When the difference between the initial specific heat capacity and the updated specific heat capacity is less than the specific heat capacity threshold, withdrawing the first alarm; In the case where the difference between the initial specific heat capacity and the updated specific heat capacity is greater than or equal to the specific heat capacity threshold, the first alarm is not withdrawn to remind the liquid cooling system that the coolant is deteriorating.
4. The method according to claim 3, characterized in that In the case that the difference between the initial specific heat capacity and the updated specific heat capacity is less than the specific heat capacity threshold, after withdrawing the first alarm, the method further includes: Updating the remaining service life of the coolant to obtain the latest remaining service life; The deterioration date of the coolant is determined according to the latest remaining service life, and a pre-warning is issued before the deterioration date of the coolant to remind the user to replace the coolant of the liquid cooling system.
5. The method according to any one of claims 1 to 4, characterized in that Before step S102, the method further includes: acquiring inlet pressure and outlet pressure of the coolant of the liquid cooling system in real time, wherein the inlet pressure is the hydraulic pressure when the coolant enters the liquid cooling system, and the outlet pressure is the hydraulic pressure when the coolant flows out of the liquid cooling system; When the difference between the liquid inlet pressure and the liquid outlet pressure is greater than or equal to the pressure difference threshold, a second alarm is issued to remind that the liquid cooling system is clogged.
6. The method according to claim 5, characterized in that When the difference between the inlet pressure and the outlet pressure is greater than or equal to the pressure difference threshold, before issuing the second alarm, the method further includes: In the case where the liquid inlet pressure is less than a predetermined liquid inlet pressure, a third alarm is issued to remind that the coolant inlet pipe of the liquid cooling system is clogged, and the predetermined liquid inlet pressure is the minimum liquid inlet pressure of the liquid cooling system.
7. The method according to claim 5, characterized in that When the difference between the inlet pressure and the outlet pressure is greater than or equal to the pressure difference threshold, after issuing a second alarm, the method further includes: Increasing the liquid inlet pressure of the liquid cooling system; Reacquire the liquid inlet pressure and the liquid outlet pressure at a second predetermined time interval to obtain an updated liquid inlet pressure and an updated liquid outlet pressure; When the difference between the updated liquid inlet pressure and the updated liquid outlet pressure is less than the pressure difference threshold, withdrawing the second alarm; When the difference between the updated liquid inlet pressure and the updated liquid outlet pressure is greater than or equal to the pressure difference threshold, the backup cooling system is started to replace the liquid cooling system for cooling.
8. A fault monitoring device for a liquid cooling system, characterized in that: include: A first acquisition module is used to execute step S102 to acquire inlet temperature, outlet temperature and flow rate of the coolant of the liquid cooling system in real time, wherein the inlet temperature is the temperature of the coolant when it enters the liquid cooling system, and the outlet temperature is the temperature of the coolant when it flows out of the liquid cooling system; A calculation module, configured to execute step S104, and calculate the current specific heat capacity of the coolant in the liquid cooling system according to the liquid inlet temperature, the liquid outlet temperature and the flow rate; The first alarm module is used to execute step S106, and when the difference between the initial specific heat capacity of the coolant and the current specific heat capacity is greater than or equal to the specific heat capacity threshold, issue a first alarm to remind the liquid cooling system that the coolant is deteriorating.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program implements the steps of the method described in any one of claims 1 to 7 when executed by a processor.
10. A server system, comprising a server, a liquid cooling system, a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: The liquid cooling system is used to dissipate heat from the server, and the processor implements the steps of the method described in any one of claims 1 to 7 when executing the computer program.
Citation Information
Patent Citations
Liquid cooling system of server
CN105549702A
Liquid cooling monitoring method and device
CN113626291A
Fault detection method and device, aircraft and storage medium
CN118746722A
Liquid cooling system fault monitoring method and semiconductor test equipment
CN118913749A
Temperature regulator abnormality detector
JP2013170810A
Cited By
Sensor fault comprehensive identification method and system in liquid cooling system
CN120180044A
A Comprehensive Sensor Fault Identification Method and System in a Liquid Cooling System
CN120180044B
Immersed energy storage system energy efficiency improving method based on liquid cooling technology
CN120453576A
Server cold plate blockage detection method, device and equipment, storage medium and product
CN120763003A
Server cold plate blockage detection method, device, equipment, storage medium and product
CN120763003B