Server cold plate blockage detection method, device and equipment, storage medium and product

By obtaining the temperature and flow data of the liquid-cooled server and combining the mapping relationship to determine the cold plate blockage, the problem of the existing technology that cannot detect cold plate blockage in time is solved, ensuring the stable operation of the server.

CN120763003AActive Publication Date: 2025-10-10CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511262100.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-05
Publication Date
2025-10-10
Estimated Expiration
2045-09-05

AI Technical Summary

Technical Problem

Existing technologies are unable to detect server cold plate blockage in a timely manner, resulting in reduced heat dissipation, which may cause batches of server failures and affect business continuity and stability.

Method used

By obtaining the current single-server liquid inlet temperature, processing unit load rate and liquid flow rate of the liquid-cooled server, and combining the preset mapping relationship to determine the target reference temperature, and comparing the deviation between the current processing unit temperature and the target reference temperature, it is determined whether the cold plate pipeline is blocked.

Benefits of technology

It enables timely identification of cold plate blockage risks, reduces the risk of batch server failures, and ensures business continuity and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120763003A_ABST
    Figure CN120763003A_ABST
Patent Text Reader

Abstract

The invention discloses a server cold plate blockage detection method, device and equipment, a storage medium and a product, and the method comprises the steps: determining a target reference temperature matched with a current working condition according to the current single-service liquid inlet temperature, the current processing unit load rate and the current single-service liquid flow of a liquid cooling server, and comparing the target reference temperature with the current temperature of a processing unit; according to the method, whether temperature abnormity caused by cold plate blockage exists or not is analyzed, the cold plate blockage risk can be recognized in time, workers can take maintenance measures in advance, the risk that the servers in batches fail at the same time is effectively reduced, serious interference on services borne by the servers is greatly reduced, and service continuity and stability are guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a server cold plate blockage detection method, device, equipment, storage medium and product. Background Art

[0002] To meet the cooling needs of liquid-cooled servers, existing solutions mostly use cold plate liquid cooling solutions. After long-term operation, impurities accumulate in the pipes of the cold plate, and microorganisms grow in the water, causing the cold plate to become clogged, which in turn affects the server's heat dissipation.

[0003] Currently, cold plate blockage is detected with a significant lag. Typically, staff only detect a cold plate blockage when the server's processing unit casing temperature reaches a set upper temperature limit. Furthermore, this cold plate blockage is not an isolated incident; once it occurs, it often causes batches of servers to fail simultaneously, requiring them to be taken offline for cold plate replacement. This situation can severely disrupt the services carried by the servers, potentially leading to service interruptions and response delays.

[0004] Therefore, there is an urgent need to provide a method that can timely detect the risk of server cold plate blockage. Summary of the Invention

[0005] Based on this, the present invention provides a server cold plate blockage detection method, device, equipment, storage medium and product to solve the defect in the prior art that the risk of server cold plate blockage cannot be detected in time.

[0006] To achieve the above objectives, an embodiment of the present invention provides a server cold plate blockage detection method, comprising: Obtain the current single-server liquid inlet temperature, current processing unit load rate, current single-server liquid flow rate, and current processing unit temperature of the liquid-cooled server; wherein the current single-server liquid inlet temperature is the temperature of the coolant currently entering the corresponding cold plate of the liquid-cooled server, and the current single-server liquid flow rate is the coolant flow rate currently entering the corresponding cold plate of the liquid-cooled server; Based on a preset mapping relationship between liquid cooling server parameters and processing unit reference temperatures, a target reference temperature is determined according to the current single server liquid inlet temperature, the current processing unit load rate, and the current single server liquid flow rate; When the current temperature of the processing unit is greater than the target reference temperature, and the deviation between the current temperature of the processing unit and the target reference temperature is greater than a set deviation threshold, the cold plate pipe of the liquid cooling server is blocked.

[0007] To achieve the above objectives, an embodiment of the present invention further provides a server cold plate blockage detection device, comprising: A data acquisition module is used to obtain the current single-server liquid inlet temperature, the current processing unit load rate, the current single-server liquid flow rate, and the current temperature of the processing unit of the liquid-cooled server; wherein the current single-server liquid inlet temperature is the temperature of the coolant currently entering the corresponding cold plate of the liquid-cooled server, and the current single-server liquid flow rate is the coolant flow rate currently entering the corresponding cold plate of the liquid-cooled server; a reference temperature determination module, configured to determine a target reference temperature based on a mapping relationship between preset liquid cooling server parameters and processing unit reference temperatures, according to the current single server liquid inlet temperature, the current processing unit load rate, and the current single server liquid flow rate; The blockage detection module is used to detect that the cold plate pipe of the liquid cooling server is blocked when the current temperature of the processing unit is greater than the target reference temperature and the deviation between the current temperature of the processing unit and the target reference temperature is greater than a set deviation threshold.

[0008] To achieve the above-mentioned objectives, an embodiment of the present invention also provides a server cold plate blockage detection device, comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the server cold plate blockage detection method as described in any of the above embodiments.

[0009] To achieve the above-mentioned purpose, an embodiment of the present invention also provides a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the server cold plate blockage detection method as described in any of the above embodiments.

[0010] To achieve the above objectives, an embodiment of the present invention further provides a computer program product, including a computer program / instruction, which, when executed by a processing unit, implements the server cold plate blockage detection method as described in any of the above embodiments.

[0011] Compared with the prior art, the server cold plate blockage detection method, device, equipment, storage medium and product disclosed by the embodiment of the application first acquire the current single-server liquid inlet temperature, the current processing unit load rate, the current single-server liquid flow and the current processing unit temperature of the liquid cooling server; then, according to the preset mapping relationship between the liquid cooling server parameters and the processing unit reference temperature, the current single-server liquid inlet temperature, the current processing unit load rate and the current single-server liquid flow are combined to determine the target reference temperature adapted to the current working condition; when the current processing unit temperature is not only greater than the target reference temperature, but also the deviation degree between the two exceeds the set deviation threshold, it can be determined that the cold plate pipeline is blocked. As can be seen, the embodiment of the application determines the target reference temperature adapted to the current working condition according to the current single-server liquid inlet temperature, the current processing unit load rate and the current single-server liquid flow of the liquid cooling server, and then compares the current processing unit temperature to analyze whether there is temperature abnormality caused by cold plate blockage, which can timely identify the cold plate blockage risk, so that the staff can take maintenance measures in advance, effectively reduce the risk of simultaneous failure of batch servers, greatly reduce the serious interference to the business carried by the server, and protect the business continuity and stability. BRIEF DESCRIPTION OF DRAWINGS

[0012] In order to more clearly illustrate the technical solutions of the present application, the following will briefly introduce the drawings needed to be used in the embodiments. Obviously, the drawings described in the following are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0013] Figure 1 is a flow diagram of a server cold plate blockage detection method provided by an embodiment of the present application; Figure 2 is a structural diagram of a server cold plate blockage detection device provided by an embodiment of the present application; Figure 3 is a structural diagram of a server cold plate blockage detection device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0014] The technical solutions in the embodiments of the present application will be described clearly and completely in the following with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0015] Referring to Figure 1 , Figure 11 is a flow chart of a server cold plate blockage detection method provided by an embodiment of the present invention. Specifically, the server cold plate blockage detection method includes steps S11 to S13: S11. Obtaining a current single-server liquid inlet temperature, a current processing unit load rate, a current single-server liquid flow rate, and a current processing unit temperature of the liquid-cooled server; wherein the current single-server liquid inlet temperature is the temperature of the coolant currently entering the corresponding cold plate of the liquid-cooled server, and the current single-server liquid flow rate is the flow rate of the coolant currently entering the corresponding cold plate of the liquid-cooled server; S12. Determine a target reference temperature based on a preset mapping relationship between liquid cooling server parameters and processing unit reference temperatures, according to the current single server liquid inlet temperature, the current processing unit load rate, and the current single server liquid flow rate; S13. When the current temperature of the processing unit is greater than the target reference temperature, and the deviation between the current temperature of the processing unit and the target reference temperature is greater than a set deviation threshold, the cold plate pipe of the liquid cooling server is blocked.

[0016] Specifically, liquid-cooled servers are installed in cabinets. The cooling distribution units (CDUs) in a liquid-cooled data center deliver coolant to the server's cold plates, removing heat from the server's processing units. The coolant then returns from the cold plates to the CDUs. The processing units can be central processing units (CPUs), graphics processing units (GPUs), or other processing units, without limitation. The cold plates are located near the processing units to facilitate the removal of heat generated by their operation. The current single-server inlet liquid temperature refers to the current coolant temperature entering the corresponding cold plate of the liquid-cooled server and can be detected by a temperature sensor located at the inlet of the corresponding cold plate of the liquid-cooled server. The current single-server liquid flow rate refers to the current coolant flow rate entering the corresponding cold plate of the liquid-cooled server and can be detected by a flow meter located at the inlet of the corresponding cold plate of the liquid-cooled server. The current temperature of the processing unit can be detected by a temperature sensor located near the processing unit housing. It is worth noting that the current single-server liquid inlet temperature, the current processing unit load rate, the current single-server liquid flow rate and the current processing unit temperature can be obtained through conventional detection means and are not limited to the above-mentioned specific detection methods.

[0017] Specifically, the current single-server liquid inlet temperature, current processing unit load rate, current single-server liquid flow rate, and current processing unit temperature of the liquid-cooled server can be obtained at regular intervals (e.g., 12 hours, 24 hours, or 36 hours) or in real time. Then, based on a pre-stored server model or a server model reported by the liquid-cooled server, a mapping relationship between the corresponding liquid-cooled server parameters and the processing unit reference temperature is selected from a database. Combined with the current single-server liquid inlet temperature, current processing unit load rate, and current single-server liquid flow rate, a target reference temperature matching the current operating conditions is found. The current processing unit temperature is then compared with the target reference temperature. If the current processing unit temperature is greater than the target reference temperature, and the deviation between the two exceeds a set deviation threshold, it is determined that the cold plate pipe of the liquid-cooled server is at risk of blockage, resulting in the processing unit being unable to dissipate heat in a timely manner. Optionally, the deviation between the two refers to the difference between the current processing unit temperature and the target reference temperature, or the deviation between the two refers to the ratio of the current processing unit temperature to the target reference temperature. The deviation between the two can also be expressed in other ways, which are not limited here.

[0018] It is worth noting that the specific value of the deviation threshold can be set according to actual conditions and is not limited here.

[0019] Optionally, the mapping relationship between the liquid-cooled server parameters and the processing unit reference temperature can be a mapping relationship between the liquid-cooled server parameters and the processing unit reference temperature in an ideal standard scenario, or a mapping relationship between the liquid-cooled server parameters and the processing unit reference temperature in a prototype measurement scenario, without limitation. The mapping relationship in the ideal standard scenario can be obtained through simulation or through experiments in a highly controlled standard laboratory environment; the prototype measurement scenario is closer to the actual state of the liquid-cooled server before it leaves the factory, and the scenario setting will relax some constraints and actively incorporate non-ideal factors.

[0020] Optionally, the real-time collected processing unit temperature can be compared with the processing unit temperature of the liquid-cooled server in the initial online stage under the same working conditions. If the two are significantly different, for example, the real-time collected processing unit temperature increases by 5-10°C or 10%-20% compared to the processing unit temperature of the liquid-cooled server in the initial online stage under the same working conditions, it is considered that there is a risk of blockage in the cold plate pipeline.

[0021] Compared with the existing technology, the embodiment of the present invention determines the target reference temperature adapted to the current working conditions based on the current single-server liquid inlet temperature of the liquid-cooled server, the current processing unit load rate, and the current single-server liquid flow rate, and then compares it with the current temperature of the processing unit to analyze whether there is a temperature anomaly caused by cold plate blockage. It can identify the risk of cold plate blockage in a timely manner, allowing staff to take maintenance measures in advance, effectively reducing the risk of simultaneous failure of batches of servers, significantly reducing serious interference with the business carried by the server, and ensuring business continuity and stability.

[0022] In a preferred embodiment, based on steps S11 to S13, the method further includes: Acquire liquid flow sequence data of a server cluster; wherein the server cluster includes at least two of the liquid-cooled servers; Calculating a cluster flow rate change trend based on the liquid flow rate sequence data; When the cluster traffic change trend is a continuous decrease in traffic, it is determined that a systemic congestion problem exists in the server cluster.

[0023] For example, assuming the server cluster includes all liquid-cooled servers served by a cooling distribution unit, the total flow rate can be measured using a flow meter installed at the liquid supply or return port of the cooling distribution unit, generating a chronological sequence of liquid flow rate data. Under a constant pressure setting, if the overall flow rate continues to decrease over a certain period of time (e.g., one month), a preliminary determination can be made that a systemic blockage exists. Alternatively, the total flow rate can be the sum of the current single-server liquid flow rates of all liquid-cooled servers in the server cluster. Other methods can also be used to determine the total flow rate, which are not limited here.

[0024] Preferably, step S12 specifically includes: when the server cluster has the systematic blockage problem, based on the mapping relationship between the preset liquid cooling server parameters and the processing unit reference temperature, the target reference temperature is determined according to the current single server liquid inlet temperature, the current processing unit load rate and the current single server liquid flow.

[0025] It is understandable that, compared with directly analyzing the temperature of the processing unit, analyzing the temperature of the processing unit only after determining that there is a systemic blockage problem can reduce the amount of data calculation and reduce the computing pressure.

[0026] Optionally, the server cluster is composed of all the liquid-cooled servers served by the cooling distribution unit; or, the server cluster is composed of some of the liquid-cooled servers served by the cooling distribution unit; the cooling liquid is transported by the cooling distribution unit to the cold plate of the liquid-cooled server to serve the heat dissipation of the liquid-cooled server, and then returns to the cooling distribution unit from the cold plate.

[0027] It is understood that if the cooling distribution unit serves a small number of servers, for example, if the total number of servers served is less than a set upper limit (this limit can be set based on actual conditions), then the server cluster refers to all liquid-cooled servers served by the cooling distribution unit. If the cooling distribution unit serves a larger number of servers, then all liquid-cooled servers can be divided into multiple server clusters. Traffic analysis can be performed on each server cluster, and based on the traffic analysis results, it can be determined whether to further perform processing unit temperature analysis on each liquid-cooled server in the server cluster, thereby reducing the amount of data processing required.

[0028] In one embodiment, based on any of the above embodiments, the mapping relationship between the liquid cooling server parameters and the processing unit reference temperature includes a mapping relationship in a prototype measurement scenario; The mapping relationship in the prototype measurement scenario is determined by the following method: Obtain the mapping relationship between liquid cooling server parameters and processing unit reference temperature under ideal standard scenarios; Constructing a prototype measurement scenario, adjusting a reference single-server liquid inlet temperature, a reference single-server liquid flow rate, and a reference processing unit load rate of the liquid-cooled server in the prototype measurement scenario, and obtaining a reference processing unit temperature of the liquid-cooled server after each adjustment; According to the reference single-server liquid inlet temperature, the reference single-server liquid flow rate, the reference processing unit load rate and the processing unit reference temperature, the mapping relationship under the ideal standard scenario is adjusted to obtain the mapping relationship under the prototype measured scenario.

[0029] For example, it is assumed that the mapping relationship between the liquid-cooled server parameters and the processing unit reference temperature is the mapping relationship between the liquid-cooled server parameters and the processing unit reference temperature in the prototype actual measurement scenario. The processing unit temperature of the liquid-cooled server is mainly determined by the processing unit load rate and the single-server liquid flow rate. In an ideal standard scenario, taking a liquid-cooled server under normal conditions as an example: when the CPU load rate is 100% and the single-server liquid flow rate is 0.5L / min, if the single-server inlet liquid temperature is 40°C, then the corresponding CPU temperature is 68°C. The mapping relationship between the liquid-cooled server parameters and the processing unit reference temperature in an ideal standard scenario can be constructed in the following way: Under different single-server liquid flow rates and different single-server inlet liquid temperatures, such as 0.5~3L / min, the CPU load rate is adjusted, the corresponding CPU case temperature is detected, and the CPU case temperature is regarded as the processing unit reference temperature. The test results were then fitted to form an algorithmic function: Tce = f(CPUload, L) + (Tin - T); where Tce is the processing unit reference temperature (in degrees Celsius); f is a functional formula, CPUload is the CPU load rate, L is the single-server liquid flow rate (in L / min); Tin is the single-server liquid inlet temperature (in degrees Celsius), and T is an empirical parameter. The actual value of T can be set according to actual conditions, such as 35, 40, or 45, but is not limited here. After obtaining the mapping relationship between liquid-cooled server parameters and processing unit reference temperature under an ideal standard scenario, a prototype test scenario was set up. In this scenario, the corresponding CPU reference temperature under different reference single-server liquid inlet temperatures, different reference single-server liquid flow rates, and different reference processing unit load rates was tested. The mapping relationship under the ideal standard scenario was then adjusted based on this test data, resulting in the mapping relationship under the prototype test scenario: Processing unit reference temperature = f(CPUload, a × L + b) + (Tin - T), where a and b are coefficients.

[0030] In this implementation, a mapping relationship is first established for an ideal standard scenario, then adjusted to approximate the mapping relationship for the actual prototype measurement scenario, rather than directly fitting the relationship in the actual prototype measurement scenario. This avoids the influence of fluctuations in factors such as the processing unit load rate and the liquid flow rate per server, as well as non-core interference variables such as equipment assembly errors. By strictly controlling variables (such as constant ambient temperature, eliminating assembly errors, and stabilizing flow rates), the ideal standard scenario completely eliminates non-core interference, allowing processing unit temperature changes to be determined solely by the three core variables: CPU load, L, and Tin. For example, in the ideal standard scenario, at a preset single-server inlet temperature, CPU load = 50% and L = 2 L / min, the CPU case temperature is 58°C. When L is adjusted to 1.9 L / min, the CPU case temperature rises to 59.5°C. This clearly defines the core relationship: "For every 0.1 L / min decrease in flow rate, the case temperature increases by 1.5°C." This relationship serves as a benchmark for subsequent analysis of deviations from actual scenarios. Direct fitting in the actual prototype measurement scenario would simply not yield such a clear core logic. Furthermore, the mapping relationship under the ideal standard scenario is a universal framework for different real-world test scenarios, adapting to individual differences between different batches of prototypes and environmental variations between different data centers. The mapping relationship under the ideal standard scenario only requires adjusting coefficients a and b based on the measured data of each batch of prototypes, eliminating the need to retest the core variable relationship. In one embodiment, based on any of the above embodiments, the mapping relationship between the liquid cooling server parameters and the processing unit reference temperature includes a mapping relationship under an ideal standard scenario and a mapping relationship under a prototype measurement scenario.

[0031] The method of determining the target reference temperature based on the mapping relationship between the preset liquid cooling server parameters and the processing unit reference temperature and according to the current single server liquid inlet temperature, the current processing unit load rate, and the current single server liquid flow rate includes: Determining a first reference temperature based on the mapping relationship in the prototype measurement scenario and the current single-server liquid inlet temperature, the current processing unit load rate, and the current single-server liquid flow rate; Based on the mapping relationship in the ideal standard scenario, determining a second reference temperature according to the current single server liquid inlet temperature, the current processing unit load rate, and the current single server liquid flow rate; Determining a target initial processing unit temperature based on initial processing unit temperature data of the liquid-cooled server and the current processing unit load rate; wherein the initial processing unit temperature data is a set of initial processing unit temperatures collected by the liquid-cooled server at an initial stage of operation under different parameter conditions; the parameter conditions include an initial single-server liquid inlet temperature, an initial single-server liquid flow rate, and an initial processing unit load rate, and a variable in the parameter conditions includes the initial processing unit load rate; subtracting the second reference temperature from the target processing unit initial temperature to obtain an error adjustment value; The first reference temperature is added to the error adjustment value to obtain a target reference temperature.

[0032] Exemplarily, the target parameter temperature is determined by the following formula: f(CPUload, a×L+b)+(Tin―T)+ε; wherein f is a function, CPUload is the CPU load rate, L is the single-server liquid flow rate; Tin is the single-server inlet temperature, T is an empirical parameter, and the actual value of T can be set according to actual conditions; ε is the error adjustment value, which is the difference between the initial temperature of the target processing unit and the second reference temperature.

[0033] For example, initial processing unit temperature data comes from initial testing of liquid-cooled servers. During testing, initial processing unit temperatures can be obtained using two variable settings: 1. The initial single-server liquid inlet temperature and initial single-server liquid flow rate are fixed, and only the processing unit load factor is adjusted as a variable to detect the corresponding initial processing unit temperature. 2. The processing unit load factor and at least one of the initial single-server liquid inlet temperature and initial single-server liquid flow rate are simultaneously adjusted as variables to detect the corresponding initial processing unit temperature.

[0034] It is worth noting that, in this embodiment, the core function of the error adjustment value is to capture the difference between the actual state of the liquid-cooled server when it is first run on-site in the computer room and the theoretical benchmark, thereby reducing interference.

[0035] In one embodiment, based on any of the above embodiments, the method is applied to a cloud monitoring platform; the current single-server liquid flow is detected by a flow meter of a computer room monitoring system and uploaded to the cloud monitoring platform by the computer room monitoring system; the current temperature of the processing unit is monitored by a server monitoring system and uploaded to the cloud monitoring platform by the server monitoring system.

[0036] Exemplarily, flow meters are provided at the liquid supply and return ports of the cooling capacity distribution unit, a temperature sensor is provided at the input end of the cold plate of the liquid-cooled server, a flow meter is provided at the cold plate of the liquid-cooled server, the flow meter establishes a communication connection with the computer room monitoring switch, the computer room monitoring switch establishes a communication connection with the computer room monitoring system, each flow meter uploads the detected flow data to the computer room monitoring system via the computer room monitoring switch, and the temperature sensor at the cold plate input end uploads the detected temperature data to the computer room monitoring system via the computer room monitoring switch. A temperature sensor is provided at the processing unit, which uploads the detected temperature data to the liquid-cooled server, which then uploads it to the server monitoring system. The computer room monitoring system and the server monitoring system can collect data every 5 minutes or every hour, and then upload the relevant data to the cloud monitoring platform, which executes the server cold plate blockage detection method. The data collection cycle of the computer room monitoring system and the server monitoring system can be set according to actual needs and is not limited here.

[0037] In one embodiment, based on any of the above embodiments, the method also includes: the cooling liquid is transported by the cooling distribution unit to the cold plate of the liquid-cooled server to serve the heat dissipation of the liquid-cooled server, and then returns from the cold plate to the cooling distribution unit; the total amount of liquid supplied by the cooling distribution unit is positively correlated with the total amount of servers served by the cooling distribution unit.

[0038] For example, the total amount of liquid supplied by the cooling distribution unit at the initial stage of operation is basically proportional to the total number of servers, or, considering whether the cold plates of each liquid-cooled server are connected in series or in parallel, the total amount of liquid supplied by the cooling distribution unit is set in combination with the number of servers on each branch.

[0039] Preferably, the current temperature of the processing unit is the current shell temperature of the processing unit. Optionally, the processing unit of the liquid cooling server is at least one of a central processing unit, a field programmable logic gate array, and a graphics processing unit.

[0040] In one embodiment, based on any of the above embodiments, the method further includes: after determining that the cold plate pipeline is blocked, generating and displaying cold plate blockage alarm information to prompt staff to take pipeline clearing operations.

[0041] Furthermore, the alarm level of the cold plate blockage alarm information is positively correlated with the degree of deviation between the current temperature of the processing unit and the target reference temperature.

[0042] Exemplarily, if the current temperature of the processing unit is higher than the target reference temperature, and the difference between the two exceeds the set temperature difference, a cold plate blockage alarm message is generated, prompting staff to handle it in time. For example: if the difference between the current temperature of the processing unit and the target reference temperature is greater than or equal to 5°C, a primary warning is issued; if the difference between the current temperature of the processing unit and the target reference temperature is greater than or equal to 10°C, a blockage alarm is issued; Exemplarily, if the current temperature of the processing unit is higher than the target reference temperature, and the ratio of the two exceeds the set ratio, a cold plate blockage alarm message is generated, prompting staff to handle it in time. For example: if the ratio of the current temperature of the processing unit to the target reference temperature is greater than or equal to 110%, a primary warning is issued; if the ratio of the current temperature of the processing unit to the target reference temperature is greater than or equal to 120%, a blockage alarm is issued.

[0043] It is worth noting that the set temperature difference and the set ratio are not limited to the above specific values, and the number of alarm levels can be set according to actual conditions.

[0044] Compared with the prior art, the method provided by the embodiment of the present invention first obtains the current single-server liquid inlet temperature, the current processing unit load rate, the current single-server liquid flow rate and the current temperature of the processing unit of the liquid-cooled server; then, based on the mapping relationship between the preset liquid-cooled server parameters and the processing unit reference temperature, combined with the current single-server liquid inlet temperature, the current processing unit load rate and the current single-server liquid flow rate, the target reference temperature adapted to the current working condition is determined; when the current temperature of the processing unit is not only greater than the target reference temperature, but also the degree of deviation between the two exceeds the set deviation threshold, it can be determined that the cold plate pipeline is blocked. It can be seen that the embodiment of the present invention determines the target reference temperature adapted to the current working condition based on the current single-server liquid inlet temperature, the current processing unit load rate and the current single-server liquid flow rate of the liquid-cooled server, and then compares it with the current temperature of the processing unit to analyze whether there is a temperature anomaly caused by cold plate blockage, which can timely identify the risk of cold plate blockage, so that the staff can take maintenance measures in advance, effectively reduce the risk of simultaneous failure of batches of servers, greatly reduce serious interference with the business carried by the server, and ensure business continuity and stability.

[0045] See also Figure 2 , an embodiment of the present invention further provides a server cold plate blockage detection device, comprising: The data acquisition module 21 is used to obtain the current liquid inlet temperature of a single server, the current processing unit load rate, the current liquid flow rate of a single server, and the current temperature of the processing unit of the liquid cooling server; wherein the current liquid inlet temperature of a single server is the temperature of the coolant currently entering the corresponding cold plate of the liquid cooling server, and the current liquid flow rate of a single server is the flow rate of the coolant currently entering the corresponding cold plate of the liquid cooling server; A reference temperature determination module 22 is configured to determine a target reference temperature based on a mapping relationship between preset liquid cooling server parameters and processing unit reference temperatures, according to the current single server liquid inlet temperature, the current processing unit load rate, and the current single server liquid flow rate; The blockage detection module 23 is used to detect that the cold plate pipe of the liquid cooling server is blocked when the current temperature of the processing unit is greater than the target reference temperature and the deviation between the current temperature of the processing unit and the target reference temperature is greater than a set deviation threshold.

[0046] In one embodiment, the device further includes a traffic analysis module, configured to: Acquire liquid flow sequence data of a server cluster; wherein the server cluster includes at least two of the liquid-cooled servers; Calculating a cluster flow rate change trend based on the liquid flow rate sequence data; When the cluster traffic change trend is a continuous decrease in traffic, it is determined that a systemic congestion problem exists in the server cluster.

[0047] Optionally, the reference temperature determination module 22 is specifically used to: when the server cluster has the systematic blockage problem, based on the mapping relationship between the preset liquid cooling server parameters and the processing unit reference temperature, determine the target reference temperature according to the current single server liquid inlet temperature, the current processing unit load rate and the current single server liquid flow.

[0048] Optionally, the server cluster is composed of all the liquid-cooled servers served by the cooling distribution unit; or, the server cluster is composed of some of the liquid-cooled servers served by the cooling distribution unit; the cooling liquid is transported by the cooling distribution unit to the cold plate of the liquid-cooled server to serve the heat dissipation of the liquid-cooled server, and then returns to the cooling distribution unit from the cold plate.

[0049] In one embodiment, the mapping relationship between the liquid cooling server parameters and the processing unit reference temperature includes a mapping relationship in a prototype measurement scenario; The mapping relationship in the prototype measurement scenario is determined by the following method: Obtain the mapping relationship between liquid cooling server parameters and processing unit reference temperature under ideal standard scenarios; Constructing a prototype measurement scenario, adjusting a reference single-server liquid inlet temperature, a reference single-server liquid flow rate, and a reference processing unit load rate of the liquid-cooled server in the prototype measurement scenario, and obtaining a reference processing unit temperature of the liquid-cooled server after each adjustment; According to the reference single-server liquid inlet temperature, the reference single-server liquid flow rate, the reference processing unit load rate and the processing unit reference temperature, the mapping relationship under the ideal standard scenario is adjusted to obtain the mapping relationship under the prototype measured scenario.

[0050] In one embodiment, the mapping relationship between the liquid cooling server parameters and the processing unit reference temperature includes a mapping relationship under an ideal standard scenario and a mapping relationship under a prototype measurement scenario; The reference temperature determination module 22 is specifically configured to: Determining a first reference temperature based on the mapping relationship in the prototype measurement scenario and the current single-server liquid inlet temperature, the current processing unit load rate, and the current single-server liquid flow rate; Based on the mapping relationship in the ideal standard scenario, determining a second reference temperature according to the current single server liquid inlet temperature, the current processing unit load rate, and the current single server liquid flow rate; Determining a target initial processing unit temperature based on initial processing unit temperature data of the liquid-cooled server and the current processing unit load rate; wherein the initial processing unit temperature data is a set of initial processing unit temperatures collected by the liquid-cooled server at an initial stage of operation under different parameter conditions; the parameter conditions include an initial single-server liquid inlet temperature, an initial single-server liquid flow rate, and an initial processing unit load rate, and a variable in the parameter conditions includes the initial processing unit load rate; subtracting the second reference temperature from the target processing unit initial temperature to obtain an error adjustment value; The first reference temperature is added to the error adjustment value to obtain a target reference temperature.

[0051] In one embodiment, the device is a cloud monitoring platform; the current single-server liquid flow is detected by a flow meter of a computer room monitoring system and uploaded to the cloud monitoring platform by the computer room monitoring system; the current temperature of the processing unit is monitored by the server monitoring system and uploaded to the cloud monitoring platform by the server monitoring system.

[0052] In one embodiment, the cooling liquid is transported by a cooling distribution unit to the cold plate of the liquid-cooled server to serve the heat dissipation of the liquid-cooled server, and then returns from the cold plate to the cooling distribution unit; the total amount of liquid supplied by the cooling distribution unit is positively correlated with the total number of servers served by the cooling distribution unit.

[0053] In one embodiment, the current temperature of the processing unit is the current shell temperature of the processing unit; the processing unit of the liquid-cooled server is a central processing unit and / or a graphics processing unit.

[0054] In one embodiment, the device further includes an early warning module for generating and displaying a cold plate blockage alarm message after determining that the cold plate pipe is blocked, so as to prompt staff to perform pipe clearing operations.

[0055] Furthermore, the alarm level of the cold plate blockage alarm information is positively correlated with the degree of deviation between the current temperature of the processing unit and the target reference temperature.

[0056] It is worth noting that the working principle of the server cold plate blockage detection device provided in the above embodiment can refer to the working process of the server cold plate blockage detection method provided in any of the above embodiments, which will not be described in detail here.

[0057] Compared with the existing technology, the server cold plate blockage detection device provided by the embodiment of the present invention determines the target reference temperature adapted to the current working conditions based on the current single-server liquid inlet temperature of the liquid-cooled server, the current processing unit load rate, and the current single-server liquid flow rate, and then compares it with the current temperature of the processing unit to analyze whether there is a temperature anomaly caused by cold plate blockage. It can identify the risk of cold plate blockage in a timely manner, allowing staff to take maintenance measures in advance, effectively reducing the risk of simultaneous failure of batches of servers, significantly reducing serious interference with the business carried by the server, and ensuring business continuity and stability.

[0058] See also Figure 3 The embodiment of the present invention further provides a server cold plate blockage detection device, comprising a processor 31, a memory 32, and a computer program stored in the memory 32 and configured to be executed by the processor 31. When the processor 31 executes the computer program, the steps in the above-mentioned server cold plate blockage detection method embodiment are implemented, for example Figure 1 S11~S13 in; or, when the processor 31 executes the computer program, the functions of the modules in the above-mentioned device embodiments are realized.

[0059] Exemplarily, the computer program can be divided into one or more modules, which are stored in the memory 32 and executed by the processor 31 to implement the present invention. The one or more modules can be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program in the server cold plate blockage detection device. For example, the computer program can be divided into multiple modules. The specific operation process of each module can be referred to the operation process of the server cold plate blockage detection device described in the above embodiment, and will not be repeated here.

[0060] The server cold plate blockage detection device can be a computing device such as a desktop computer, laptop, PDA, or cloud server. The server cold plate blockage detection device can include, but is not limited to, a processor 31 and a memory 32. Those skilled in the art will appreciate that the server cold plate blockage detection device can also include input / output devices, network access devices, buses, and the like.

[0061] The processor 31 may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor. The processor 31 serves as the control center of the server cold plate blockage detection device and connects various components of the server cold plate blockage detection device using various interfaces and lines.

[0062] The memory 32 can be used to store the computer programs and / or modules. The processor 31 implements the various functions of the server cold plate blockage detection device by running or executing the computer programs and / or modules stored in the memory 32 and accessing the data stored in the memory 32. The memory 32 may primarily include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function (such as image playback), while the data storage area may store data generated based on the use of the mobile phone. Furthermore, the memory 32 may include high-speed random access memory (RAM) and non-volatile memory, such as a hard disk, internal memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one disk storage device, a flash memory device, or other volatile solid-state storage device.

[0063] If the module integrated into the server cold plate blockage detection device is implemented as a software functional unit and sold or used as a standalone product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention can implement all or part of the process steps in the above-described method embodiments by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When executed by the processor 31, the computer program can implement the steps of each of the above-described method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium can include any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunications signal, and a software distribution medium.

[0064] An embodiment of the present invention further provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the server cold plate blockage detection method as described in any of the above embodiments.

[0065] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.

Claims

1. A server cold plate blockage detection method, characterized in that: include: Obtain the current single-server liquid inlet temperature, current processing unit load rate, current single-server liquid flow rate, and current processing unit temperature of the liquid-cooled server; wherein the current single-server liquid inlet temperature is the temperature of the coolant currently entering the corresponding cold plate of the liquid-cooled server, and the current single-server liquid flow rate is the coolant flow rate currently entering the corresponding cold plate of the liquid-cooled server; Based on a preset mapping relationship between liquid cooling server parameters and processing unit reference temperatures, a target reference temperature is determined according to the current single server liquid inlet temperature, the current processing unit load rate, and the current single server liquid flow rate; When the current temperature of the processing unit is greater than the target reference temperature, and the deviation between the current temperature of the processing unit and the target reference temperature is greater than a set deviation threshold, the cold plate pipe of the liquid cooling server is blocked.

2. The server cold plate blockage detection method according to claim 1, characterized in that: Also includes: Acquire liquid flow sequence data of a server cluster; wherein the server cluster includes at least two of the liquid-cooled servers; Calculating a cluster flow rate change trend based on the liquid flow rate sequence data; When the cluster traffic change trend is a continuous decrease in traffic, it is determined that a systemic congestion problem exists in the server cluster.

3. The server cold plate blockage detection method according to claim 2, characterized in that: The method of determining the target reference temperature based on the mapping relationship between the preset liquid cooling server parameters and the processing unit reference temperature and according to the current single server liquid inlet temperature, the current processing unit load rate, and the current single server liquid flow rate includes: When the server cluster has the systematic blockage problem, the target reference temperature is determined based on the mapping relationship between the preset liquid cooling server parameters and the processing unit reference temperature, according to the current single server liquid inlet temperature, the current processing unit load rate and the current single server liquid flow.

4. The server cold plate blockage detection method according to claim 2, wherein: The server cluster is composed of all the liquid-cooled servers served by the cooling distribution unit; or, the server cluster is composed of some of the liquid-cooled servers served by the cooling distribution unit; the cooling liquid is transported by the cooling distribution unit to the cold plates of the liquid-cooled servers to serve the heat dissipation of the liquid-cooled servers, and then returns to the cooling distribution unit from the cold plates.

5. The server cold plate blockage detection method according to claim 1, characterized in that: The mapping relationship between the liquid cooling server parameters and the processing unit reference temperature includes a mapping relationship under a prototype measurement scenario; The mapping relationship in the prototype measurement scenario is determined by the following method: Obtain the mapping relationship between liquid cooling server parameters and processing unit reference temperature under ideal standard scenarios; Constructing a prototype measurement scenario, adjusting a reference single-server liquid inlet temperature, a reference single-server liquid flow rate, and a reference processing unit load rate of the liquid-cooled server in the prototype measurement scenario, and obtaining a reference processing unit temperature of the liquid-cooled server after each adjustment; According to the reference single-server liquid inlet temperature, the reference single-server liquid flow rate, the reference processing unit load rate and the processing unit reference temperature, the mapping relationship under the ideal standard scenario is adjusted to obtain the mapping relationship under the prototype measured scenario.

6. The server cold plate blockage detection method according to any one of claims 1 to 5, characterized in that: The mapping relationship between the liquid cooling server parameters and the reference temperature of the processing unit includes a mapping relationship under an ideal standard scenario and a mapping relationship under a prototype measurement scenario; The method of determining the target reference temperature based on the mapping relationship between the preset liquid cooling server parameters and the processing unit reference temperature and according to the current single server liquid inlet temperature, the current processing unit load rate, and the current single server liquid flow rate includes: Determining a first reference temperature based on the mapping relationship in the prototype measurement scenario and the current single-server liquid inlet temperature, the current processing unit load rate, and the current single-server liquid flow rate; Based on the mapping relationship in the ideal standard scenario, determining a second reference temperature according to the current single server liquid inlet temperature, the current processing unit load rate, and the current single server liquid flow rate; Determining a target initial processing unit temperature based on initial processing unit temperature data of the liquid-cooled server and the current processing unit load rate; wherein the initial processing unit temperature data is a set of initial processing unit temperatures collected by the liquid-cooled server at an initial stage of operation under different parameter conditions; the parameter conditions include an initial single-server liquid inlet temperature, an initial single-server liquid flow rate, and an initial processing unit load rate, and a variable in the parameter conditions includes the initial processing unit load rate; subtracting the second reference temperature from the initial temperature of the target processing unit to obtain an error adjustment value; The first reference temperature is added to the error adjustment value to obtain a target reference temperature.

7. The server cold plate blockage detection method according to claim 1, wherein: The method is applied to a cloud monitoring platform; the current single-server liquid flow is detected by a flow meter of a computer room monitoring system and uploaded to the cloud monitoring platform by the computer room monitoring system; the current temperature of the processing unit is monitored by a server monitoring system and uploaded to the cloud monitoring platform by the server monitoring system.

8. The server cold plate blockage detection method according to claim 1, wherein: Also includes: The cooling liquid is transported by the cooling distribution unit to the cold plate of the liquid-cooled server to serve the heat dissipation of the liquid-cooled server, and then returns to the cooling distribution unit from the cold plate; The total amount of liquid supplied by the cooling distribution unit is positively correlated with the total amount of servers served by the cooling distribution unit.

9. The server cold plate blockage detection method according to claim 1, wherein: The current temperature of the processing unit is the current shell temperature of the processing unit; the processing unit of the liquid cooling server is a central processing unit and / or a graphics processing unit.

10. The server cold plate blockage detection method according to claim 1, wherein: Also includes: After determining that the cold plate pipeline is blocked, a cold plate blockage alarm message is generated and displayed to prompt the staff to take pipeline unblocking operations.

11. The server cold plate blockage detection method according to claim 10, wherein: The alarm level of the cold plate blockage alarm information is positively correlated to the degree of deviation between the current temperature of the processing unit and the target reference temperature.

12. A server cold plate blockage detection device, characterized in that: include: A data acquisition module is used to obtain the current single-server liquid inlet temperature, the current processing unit load rate, the current single-server liquid flow rate, and the current temperature of the processing unit of the liquid-cooled server; wherein the current single-server liquid inlet temperature is the temperature of the coolant currently entering the corresponding cold plate of the liquid-cooled server, and the current single-server liquid flow rate is the coolant flow rate currently entering the corresponding cold plate of the liquid-cooled server; a reference temperature determination module, configured to determine a target reference temperature based on a mapping relationship between preset liquid cooling server parameters and processing unit reference temperatures, according to the current single server liquid inlet temperature, the current processing unit load rate, and the current single server liquid flow rate; The blockage detection module is used to detect that the cold plate pipe of the liquid cooling server is blocked when the current temperature of the processing unit is greater than the target reference temperature and the deviation between the current temperature of the processing unit and the target reference temperature is greater than a set deviation threshold.

13. A server cold plate blockage detection device, characterized in that: The method comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, the server cold plate blockage detection method according to any one of claims 1 to 11 is implemented.

14. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored computer program, wherein when the computer program is executed, the device where the computer-readable storage medium is located is controlled to execute the server cold plate blockage detection method according to any one of claims 1 to 11.

15. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the server cold plate blockage detection method according to any one of claims 1 to 11 is implemented.

Citation Information

Patent Citations

  • Liquid cooling server intelligent temperature control method based on local software monitoring and liquid cooling server

    CN119536482A

  • Fault monitoring method and device of liquid cooling system and server system

    CN119958889A

  • Air conditioner capable of diagnosing failure of refrigerant shortage and system block and diagnosis method thereof

    CN1737448A

  • IGBT power module water -cooling plate resistance stopper alarm device

    CN206401308U

  • Engine cooling system onboard diagnostic strategy

    US20100095909A1