Server cold plate blockage detection method, device, equipment, storage medium and product
By acquiring temperature and flow data from the liquid-cooled server and combining the mapping relationship to determine cold plate blockage, the problem of not being able to detect cold plate blockage in a timely manner in existing technologies is solved, ensuring the stable operation of the server.
Patent Information
- Application Number
- CN202511262100.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-05
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-09-05
AI Technical Summary
Current technology cannot detect server cold plate blockage in a timely manner, which leads to reduced heat dissipation and may cause batch server failures, affecting business continuity and stability.
By acquiring the current single-server liquid inlet temperature, processing unit load rate, and liquid flow rate of the liquid-cooled server, and combining this with a preset mapping relationship, the target reference temperature is determined. The deviation between the current temperature of the processing unit and the target reference temperature is compared to determine whether the cold plate pipe is blocked.
It enables timely identification of cold plate blockage risks, reduces the risk of batch server failures, and ensures business continuity and stability.
Smart Images

Figure CN120763003B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, in particular, relate to a server cold plate blockage detection method, device, equipment, storage medium and product. BACKGROUND
[0002] In view of the equipment cooling demand of liquid-cooled server, the existing technical solutions mostly adopt cold plate type liquid cooling scheme. After long-term operation of the cold plate, impurities accumulate inside the pipeline, and microorganisms breed in the water, which causes the cold plate to be easily blocked, thereby affecting the heat dissipation effect of the server.
[0003] At present, the discovery of the cold plate blockage problem has obvious hysteresis. Usually, only when the server processing unit shell temperature rises to the set upper limit of temperature, can the staff realize that the cold plate has a blockage fault. More seriously, such cold plate blockage problem is not an individual case, once it occurs, it often causes batch of servers to fail simultaneously, and the cold plate needs to be replaced. This situation will directly cause serious interference to the various businesses carried by the server, which may cause business interruption, response delay and other problems.
[0004] Therefore, it is urgent to provide a way to detect the risk of server cold plate blockage in a timely manner. SUMMARY
[0005] Based on this, the present application provides a server cold plate blockage detection method, device, equipment, storage medium and product to solve the defect that the risk of server cold plate blockage cannot be detected in a timely manner in the prior art.
[0006] To achieve the above-mentioned purpose, the present application provides a server cold plate blockage detection method, comprising:
[0007] Obtaining the current single server liquid inlet temperature, the current processing unit load rate, the current single server liquid flow and the current processing unit temperature of the liquid-cooled server; wherein the current single server liquid inlet temperature is the cooling liquid temperature currently entering the corresponding cold plate of the liquid-cooled server, and the current single server liquid flow is the cooling liquid flow currently entering the corresponding cold plate of the liquid-cooled server;
[0008] Based on the mapping relationship between the preset liquid-cooled server parameters and the processing unit reference temperature, the target reference temperature is determined according to the current single server liquid inlet temperature, the current processing unit load rate and the current single server liquid flow;
[0009] When the current processing unit temperature is greater than the target reference temperature, and the deviation degree between the current processing unit temperature and the target reference temperature is greater than the set deviation threshold, the cold plate pipeline of the liquid-cooled server is blocked.
[0010] To achieve the above object, the embodiment of the present application further provides a server cold plate blockage detection device, comprising:
[0011] The data acquisition module is configured to acquire a current single-server liquid inlet temperature, a current processing unit load rate, a current single-server liquid flow and a current processing unit temperature of the liquid-cooled server, wherein the current single-server liquid inlet temperature is a cooling liquid temperature currently entering a corresponding cold plate of the liquid-cooled server, and the current single-server liquid flow is a cooling liquid flow currently entering the corresponding cold plate of the liquid-cooled server.
[0012] The reference temperature determination module is configured to determine a target reference temperature based on a preset mapping relationship between liquid-cooled server parameters and processing unit reference temperatures, according to the current single-server liquid inlet temperature, the current processing unit load rate and the current single-server liquid flow.
[0013] The blockage detection module is configured to determine that a cold plate pipeline of the liquid-cooled server is blocked when the current processing unit temperature is greater than the target reference temperature, and a deviation degree between the current processing unit temperature and the target reference temperature is greater than a set deviation threshold.
[0014] To achieve the above object, the embodiment of the present application further provides a server cold plate blockage detection device, comprising a processor, a memory and a computer program stored in the memory and configured to be executed by the processor, wherein the processor executes the computer program to realize the server cold plate blockage detection method according to any one of the above embodiments.
[0015] To achieve the above object, the embodiment of the present application further provides a computer readable storage medium, comprising a stored computer program, wherein the computer readable storage medium controls a device where the computer readable storage medium is located to execute the server cold plate blockage detection method according to any one of the above embodiments when the computer program runs.
[0016] To achieve the above object, the embodiment of the present application further provides a computer program product, comprising computer programs / instructions, wherein the computer programs / instructions are executed by a processing unit to realize the server cold plate blockage detection method according to any one of the above embodiments.
[0017] Compared with the prior art, the server cold plate blockage detection method, device, equipment, storage medium and product disclosed by the embodiment of the application first acquire the current single-server liquid inlet temperature, the current processing unit load rate, the current single-server liquid flow and the current processing unit temperature of the liquid cooling server; then, according to the preset mapping relationship between the liquid cooling server parameters and the processing unit reference temperature, the current single-server liquid inlet temperature, the current processing unit load rate and the current single-server liquid flow are combined to determine the target reference temperature adapted to the current working condition; when the current processing unit temperature is not only greater than the target reference temperature, but also the deviation degree between the two exceeds the set deviation threshold, it can be determined that the cold plate pipeline is blocked. As can be seen, the embodiment of the application determines the target reference temperature adapted to the current working condition according to the current single-server liquid inlet temperature, the current processing unit load rate and the current single-server liquid flow of the liquid cooling server, and then compares the current processing unit temperature to analyze whether there is temperature abnormality caused by cold plate blockage, which can timely identify the cold plate blockage risk, so that the staff can take maintenance measures in advance, effectively reduce the risk of simultaneous failure of batch servers, greatly reduce the serious interference to the business carried by the server, and protect the business continuity and stability. BRIEF DESCRIPTION OF DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the present application, the following will briefly introduce the drawings needed to be used in the embodiments. Obviously, the drawings described in the following are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0019] Figure 1 is a flow diagram of a server cold plate blockage detection method provided by an embodiment of the present application;
[0020] Figure 2 is a structural diagram of a server cold plate blockage detection device provided by an embodiment of the present application;
[0021] Figure 3 is a structural diagram of a server cold plate blockage detection device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0022] The technical solutions in the embodiments of the present application will be described clearly and completely in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0023] Referring to Figure 1 , Figure 1is a flowchart of a server cold plate blockage detection method provided by an embodiment of the present application. Specifically, the server cold plate blockage detection method comprises steps S11-S13:
[0024] S11, obtaining a current single-server inlet liquid temperature, a current processing unit load rate, a current single-server liquid flow rate and a current processing unit temperature of a liquid-cooled server; wherein the current single-server inlet liquid temperature is the temperature of the cooling liquid currently entering the corresponding cold plate of the liquid-cooled server, and the current single-server liquid flow rate is the flow rate of the cooling liquid currently entering the corresponding cold plate of the liquid-cooled server;
[0025] S12, determining a target reference temperature based on a preset mapping relationship between the parameters of the liquid-cooled server and the reference temperature of the processing unit, according to the current single-server inlet liquid temperature, the current processing unit load rate and the current single-server liquid flow rate;
[0026] S13, when the current processing unit temperature is greater than the target reference temperature, and the degree of deviation between the current processing unit temperature and the target reference temperature is greater than a set deviation threshold, the cold plate pipeline of the liquid-cooled server is blocked.
[0027] Specifically, the liquid-cooled server is installed in a cabinet, and a cooling distribution unit (CDU) of a liquid-cooled data center delivers cooling liquid to the cold plate of the liquid-cooled server to take away the heat of the processing unit of the liquid-cooled server, and then the cooling liquid returns to the cooling distribution unit from the cold plate. The processing unit can be a central processing unit (CPU), a graphics processing unit (GPU), or other processing units, which are not limited here. The cold plate is arranged near the processing unit to facilitate taking away the heat generated by the operation of the processing unit. The current single-server inlet liquid temperature refers to the temperature of the cooling liquid currently entering the corresponding cold plate of the liquid-cooled server, which can be detected by a temperature sensor arranged at the inlet of the corresponding cold plate of the liquid-cooled server. The current single-server liquid flow rate refers to the flow rate of the cooling liquid currently entering the corresponding cold plate of the liquid-cooled server, which can be detected by a flow meter arranged at the inlet of the corresponding cold plate of the liquid-cooled server. The current processing unit temperature can be detected by a temperature sensor arranged near the housing of the processing unit. It should be noted that the current single-server inlet liquid temperature, the current processing unit load rate, the current single-server liquid flow rate and the current processing unit temperature can be obtained by conventional detection methods, and are not limited to the above specific detection methods.
[0028] Specifically, the current single-server inlet liquid temperature, the current processing unit load rate, the current single-server liquid flow rate and the processing unit current temperature of the liquid cooling server can be acquired at a certain time interval (such as 12 hours, 24 hours or 36 hours, etc.) or in real time, then based on the server model pre-stored or reported by the liquid cooling server, the corresponding mapping relationship between the liquid cooling server parameters and the processing unit reference temperature is selected from the database, the current single-server inlet liquid temperature, the current processing unit load rate and the current single-server liquid flow rate are combined to find the target reference temperature matched with the current working condition, and the processing unit current temperature and the target reference temperature are compared, if the processing unit current temperature is greater than the target reference temperature and the deviation degree of the two is greater than the set deviation threshold, it is considered that there is a risk of blockage of the cold plate pipeline of the liquid cooling server, which causes the processing unit to fail to dissipate heat in time. Optionally, the deviation degree of the two refers to the difference between the processing unit current temperature and the target reference temperature, or the deviation degree of the two refers to the ratio of the processing unit current temperature to the target reference temperature, and the deviation degree of the two can also be represented in other ways, which is not limited here.
[0029] It is worth noting that the specific value of the set deviation threshold can be set according to actual conditions, which is not limited here.
[0030] Optionally, the mapping relationship between the liquid cooling server parameters and the processing unit reference temperature can be the mapping relationship between the liquid cooling server parameters and the processing unit reference temperature under an ideal standard scenario, or the mapping relationship between the liquid cooling server parameters and the processing unit reference temperature under a prototype measurement scenario, which is not limited here. The mapping relationship under the ideal standard scenario can be obtained by simulation, and can also be obtained by experiment in a highly controllable standard laboratory environment; the prototype measurement scenario is closer to the actual state before the liquid cooling server is shipped, and the scene setting relaxes some constraints and actively incorporates non-ideal factors.
[0031] Optionally, the processing unit temperature acquired in real time can also be compared with the processing unit temperature at the initial online stage of the liquid cooling server under the same working condition, if the two are quite different, for example, the processing unit temperature acquired in real time increases by 5-10℃ or 10%-20% compared with the processing unit temperature at the initial online stage of the liquid cooling server under the same working condition, it is considered that there is a risk of blockage of the cold plate pipeline.
[0032] Compared with the prior art, the embodiment of the present application determines the target reference temperature adapted to the current working condition according to the current single-server inlet liquid temperature, the current processing unit load rate and the current single-server liquid flow rate of the liquid cooling server, and then compares it with the processing unit current temperature to analyze whether there is temperature abnormality caused by cold plate blockage, which can identify the risk of cold plate blockage in time, so that the staff can take maintenance measures in advance, effectively reduce the risk of simultaneous failure of batch servers, greatly reduce the serious interference to the business carried by the servers, and protect the business continuity and stability.
[0033] In a preferred embodiment, based on steps S11-S13, the method further comprises:
[0034] obtaining liquid flow sequence data of a server cluster; wherein the server cluster comprises at least two liquid cooling servers;
[0035] calculating a cluster flow change trend according to the liquid flow sequence data;
[0036] when the cluster flow change trend is a continuous decrease in flow, determining that the server cluster has a systematic clogging problem.
[0037] For example, assuming that the server cluster comprises all liquid cooling servers serviced by a cold energy distribution unit, the total flow can be detected by a flow meter arranged at the liquid supply port or the liquid return port of the cold energy distribution unit, and the liquid flow sequence data is formed in chronological order. Under constant pressure setting, if the overall flow continuously decreases within a certain period of time (e.g., within 1 month), it can be preliminarily determined that there is a systematic clogging problem. Alternatively, the total flow can also be the sum of the current single-server liquid flow of all liquid cooling servers in the server cluster, and the total flow can also be determined by other means, which is not limited herein.
[0038] Preferably, step S12 specifically comprises: when the server cluster has the systematic clogging problem, determining a target reference temperature based on a preset mapping relationship between the liquid cooling server parameters and the processing unit reference temperature, according to the current single-server liquid inlet temperature, the current processing unit load rate, and the current single-server liquid flow.
[0039] It can be understood that, compared with directly analyzing the processing unit temperature, analyzing the processing unit temperature after determining that there is a systematic clogging problem can reduce the data operation amount and reduce the operation pressure.
[0040] Alternatively, the server cluster is composed of all liquid cooling servers serviced by a cold energy distribution unit; or the server cluster is composed of part of the liquid cooling servers serviced by the cold energy distribution unit; the cooling liquid is delivered by the cold energy distribution unit to the cold plate of the liquid cooling server, serves to dissipate heat of the liquid cooling server, and then returns to the cold energy distribution unit from the cold plate.
[0041] It can be understood that if the server scale served by the cold quantity distribution unit is small, for example, the total number of servers served by the cold quantity distribution unit is less than a set upper limit value (which can be set according to actual conditions), the server cluster refers to all liquid-cooled servers served by the cold quantity distribution unit. If the server scale served by the cold quantity distribution unit is large, the plurality of liquid-cooled servers can be divided to obtain a plurality of server clusters, and flow analysis is performed on each server cluster. According to the flow analysis result, it is determined whether to further perform processing unit temperature analysis on each liquid-cooled server in the server cluster, so as to reduce the data operation amount.
[0042] In an embodiment, on the basis of any of the above embodiments, the mapping relationship between the liquid-cooled server parameter and the processing unit reference temperature comprises a mapping relationship in a prototype measurement scenario.
[0043] The mapping relationship in the prototype measurement scenario is determined by the following method:
[0044] Obtain the mapping relationship between the liquid-cooled server parameter and the processing unit reference temperature in an ideal standard scenario.
[0045] Construct a prototype measurement scenario, in which the reference single-server liquid inlet temperature, the reference single-server liquid flow rate, and the reference processing unit load rate of the liquid-cooled server are adjusted, and after each adjustment, the processing unit reference temperature of the liquid-cooled server is obtained.
[0046] According to the reference single-server liquid inlet temperature, the reference single-server liquid flow rate, the reference processing unit load rate, and the processing unit reference temperature, the mapping relationship in the ideal standard scenario is adjusted to obtain the mapping relationship in the prototype measurement scenario.
[0047] For example, assume that the mapping relationship between the liquid cooling server parameters and the processing unit reference temperature is the mapping relationship between the liquid cooling server parameters and the processing unit reference temperature in a prototype test scenario. The processing unit temperature of the liquid cooling server is mainly determined by the processing unit load rate and the single-server liquid flow. In an ideal standard scenario, for example, a certain liquid cooling server under normal conditions: when the CPU load rate is 100% and the single-server liquid flow is 0.5 L / min, if the single-server inlet liquid temperature is 40°C, then the corresponding CPU temperature is 68°C. The mapping relationship between the liquid cooling server parameters and the processing unit reference temperature in the ideal standard scenario can be constructed in the following way: under different single-server liquid flows and different single-server inlet liquid temperatures, such as 0.5-3 L / min, adjust the CPU load rate, and detect the corresponding CPU case temperature, which is regarded as the processing unit reference temperature. Then, the test results are fitted to form an algorithm function: Tce = f(CPUload, L) + (Tin-T); wherein Tce is the processing unit reference temperature, the unit of which is Celsius; f is the function, CPUload is the CPU load rate, L is the single-server liquid flow, the unit of L is L / min; Tin is the single-server inlet liquid temperature, the unit of Tin is Celsius, T is an empirical parameter, the actual value of T can be set according to the actual situation, such as 35, 40 or 45, etc., which is not limited herein. After obtaining the mapping relationship between the liquid cooling server parameters and the processing unit reference temperature in the ideal standard scenario, the prototype test scenario is set, and the corresponding CPU reference temperature under different reference single-server inlet liquid temperatures, different reference single-server liquid flows and different reference processing unit load rates is tested in the scenario; then, the mapping relationship in the ideal standard scenario is adjusted according to these test data to obtain the mapping relationship in the prototype test scenario: processing unit reference temperature = f(CPUload, a x L + b) + (Tin-T), wherein a and b are coefficients.
[0048] In the embodiment, the mapping relationship in the ideal standard scene is established first, and then the mapping relationship close to the actual prototype measurement scene is obtained by adjustment, rather than directly testing and fitting in the prototype measurement scene, so that the influence of factors such as processing unit load rate, single service liquid flow and the influence of non-core interference variables such as equipment assembly error can be avoided. The ideal standard scene can completely strip non-core interference by strictly controlling variables (such as constant environmental temperature, eliminating assembly error, and stable flow), so that the temperature change of the processing unit is determined only by the three core variables of CPU load, L and Tin. For example, in the ideal standard scene, when the preset single service liquid temperature is CPU load = 50%, L = 2L / min, the CPU shell temperature is 58℃, and when L is adjusted to 1.9L / min, the CPU shell temperature rises to 59.5℃, so that the core relationship of “the flow decreases by 0.1L / min, and the shell temperature rises by 1.5℃” can be determined. This relationship is a “baseline ruler” for subsequent analysis of actual scene deviation, and if direct fitting is performed in the prototype measurement scene, such clear core logic cannot be obtained. In addition, the mapping relationship in the ideal standard scene is a general framework for different prototype measurement scenes, which is suitable for individual differences of different batches of prototypes and environmental differences of different machine rooms. The mapping relationship in the ideal standard scene only needs to adjust the coefficients a and b according to the measured data of each batch of prototypes, without the need to retest the core variable relationship.
[0049] In one embodiment, on the basis of any of the above embodiments, the mapping relationship between the liquid-cooled server parameters and the processing unit reference temperature includes a mapping relationship in an ideal standard scene and a mapping relationship in a prototype measurement scene.
[0050] The mapping relationship between the liquid-cooled server parameters and the processing unit reference temperature based on the preset mapping relationship is used to determine a target reference temperature according to the current single service liquid temperature, the current processing unit load rate and the current single service liquid flow, and includes:
[0051] The mapping relationship in the prototype measurement scene is used to determine a first reference temperature according to the current single service liquid temperature, the current processing unit load rate and the current single service liquid flow.
[0052] The mapping relationship in the ideal standard scene is used to determine a second reference temperature according to the current single service liquid temperature, the current processing unit load rate and the current single service liquid flow.
[0053] determining a target processing unit initial temperature according to the current processing unit load rate based on processing unit initial temperature data of the liquid-cooled server; wherein the processing unit initial temperature data is a set of processing unit initial temperatures collected by the liquid-cooled server at an initial stage of operation under different parameter conditions; the parameter conditions include an initial single-server liquid inlet temperature, an initial single-server liquid flow rate, and an initial processing unit load rate, and the variable in the parameter conditions includes the initial processing unit load rate;
[0054] subtracting the second reference temperature from the target processing unit initial temperature to obtain an error adjustment value;
[0055] adding the first reference temperature to the error adjustment value to obtain a target reference temperature.
[0056] For example, the target parameter temperature is determined by the following formula: f(CPUload, a x L + b) + (Tin - T) + ε; wherein f is a function, CPUload is the CPU load rate, L is the single-server liquid flow rate, Tin is the single-server liquid inlet temperature, T is an empirical parameter, the actual value of T can be set according to the actual situation, and ε is the error adjustment value, which is the difference between the target processing unit initial temperature and the second reference temperature.
[0057] For example, the processing unit initial temperature data comes from the initial test of the liquid-cooled server when it is put into operation. The processing unit initial temperature can be obtained by setting two variables during the test: 1. fixing the initial single-server liquid inlet temperature and the initial single-server liquid flow rate, and adjusting only the processing unit load rate as a variable to detect the corresponding processing unit initial temperature; 2. adjusting the processing unit load rate and at least one of the initial single-server liquid inlet temperature and the initial single-server liquid flow rate as a variable to detect the corresponding processing unit initial temperature.
[0058] It is worth noting that in this embodiment, the core function of the error adjustment value is to capture the difference between the actual state and the theoretical benchmark of the liquid-cooled server when it is first operated in the computer room, and to reduce interference.
[0059] In one embodiment, on the basis of any of the above embodiments, the method is applied to a cloud monitoring platform; the current single-server liquid flow rate is detected by a flow meter of a computer room monitoring system and uploaded to the cloud monitoring platform by the computer room monitoring system; and the current processing unit temperature is monitored by a server monitoring system and uploaded to the cloud monitoring platform by the server monitoring system.
[0060] Exemplarily, flow meters are arranged at the liquid supply port and the liquid return port of the cold quantity distribution unit, a temperature sensor is arranged at the input end of the cold plate of the liquid-cooled server, and flow meters are arranged at the cold plate of the liquid-cooled server, which are in communication connection with the machine room monitoring switch, the machine room monitoring switch is in communication connection with the machine room monitoring system, each flow meter uploads the detected flow data to the machine room monitoring system through the machine room monitoring switch, and the temperature sensor at the input end of the cold plate uploads the detected temperature data to the machine room monitoring system through the machine room monitoring switch. A temperature sensor is arranged at the processing unit, which uploads the detected temperature data to the liquid-cooled server and then to the server monitoring system. The machine room monitoring system and the server monitoring system can collect data every 5 minutes or every hour, and then upload the relevant data to the cloud monitoring platform, which executes the server cold plate blockage detection method. The data collection period of the machine room monitoring system and the server monitoring system can be set according to actual needs, which is not limited herein.
[0061] In an embodiment, on the basis of any of the above-mentioned embodiments, the method further comprises: the cooling liquid is delivered by the cold quantity distribution unit to the cold plate of the liquid-cooled server, serves to dissipate heat of the liquid-cooled server, and is returned from the cold plate to the cold quantity distribution unit; and the total liquid supply amount of the cold quantity distribution unit is in a positive correlation with the total amount of servers served by the cold quantity distribution unit.
[0062] Exemplarily, the total liquid supply amount of the cold quantity distribution unit at the initial stage of operation is substantially in a positive correlation with the total amount of servers, or the total liquid supply amount of the cold quantity distribution unit is set in combination with the number of servers on each branch, considering whether the cold plates of the liquid-cooled servers are in series or in parallel.
[0063] Preferably, the current temperature of the processing unit is the current shell temperature of the processing unit. Optionally, the processing unit of the liquid-cooled server is at least one of a central processing unit, a field programmable gate array, and a graphics processing unit.
[0064] In an embodiment, on the basis of any of the above-mentioned embodiments, the method further comprises: after determining that the cold plate pipeline is blocked, generating and displaying cold plate blockage warning information to prompt staff to take pipeline dredging operation.
[0065] Further, the alarm level of the cold plate blockage warning information is positively correlated with the deviation degree between the current temperature of the processing unit and the target reference temperature.
[0066] For example, if the current temperature of the processing unit is higher than the target reference temperature, and the difference between the two exceeds the set temperature difference, a cold plate blockage warning information is generated to prompt the staff to handle it in time. For example, if the difference between the current temperature of the processing unit and the target reference temperature is greater than or equal to 5℃, a primary early warning is issued; if the difference between the current temperature of the processing unit and the target reference temperature is greater than or equal to 10℃, a blockage warning is issued. For example, if the ratio of the current temperature of the processing unit to the target reference temperature exceeds the set ratio, a cold plate blockage warning information is generated to prompt the staff to handle it in time. For example, if the ratio of the current temperature of the processing unit to the target reference temperature is greater than or equal to 110%, a primary early warning is issued; if the ratio of the current temperature of the processing unit to the target reference temperature is greater than or equal to 120%, a blockage warning is issued.
[0067] It should be noted that the set temperature difference and the set ratio are not limited to the above specific values, and the number of warning levels can be set according to actual conditions.
[0068] Compared with the prior art, the method provided by the embodiment of the present application first acquires the current single-server liquid inlet temperature, the current processing unit load rate, the current single-server liquid flow rate, and the current processing unit temperature of the liquid cooling server; then, according to the preset mapping relationship between the liquid cooling server parameters and the processing unit reference temperature, the current single-server liquid inlet temperature, the current processing unit load rate, and the current single-server liquid flow rate are combined to determine the target reference temperature adapted to the current working condition; when the current processing unit temperature is not only greater than the target reference temperature, but also the deviation between the two exceeds the set deviation threshold, it can be determined that the cold plate pipeline is blocked. As can be seen, by determining the target reference temperature adapted to the current working condition according to the current single-server liquid inlet temperature, the current processing unit load rate, and the current single-server liquid flow rate of the liquid cooling server, and then comparing it with the current processing unit temperature, whether there is a temperature anomaly caused by cold plate blockage can be analyzed, the cold plate blockage risk can be identified in time, the staff can take maintenance measures in advance, the risk of simultaneous failure of batches of servers can be effectively reduced, the serious interference on the business carried by the servers can be greatly reduced, and the business continuity and stability can be ensured.
[0069] Referring to Figure 2 The embodiment of the present application also provides a server cold plate blockage detection device, which comprises:
[0070] The data acquisition module 21 is configured to acquire the current single-server liquid inlet temperature, the current processing unit load rate, the current single-server liquid flow rate, and the current processing unit temperature of the liquid cooling server; wherein the current single-server liquid inlet temperature is the cooling liquid temperature currently entering the corresponding cold plate of the liquid cooling server, and the current single-server liquid flow rate is the cooling liquid flow rate currently entering the corresponding cold plate of the liquid cooling server;
[0071] The reference temperature determination module 22 is configured to determine a target reference temperature based on a preset mapping relationship between liquid-cooled server parameters and processing unit reference temperatures, according to the current single-server liquid inlet temperature, the current processing unit load rate, and the current single-server liquid flow rate.
[0072] The blockage detection module 23 is configured to determine that the cold plate pipeline of the liquid-cooled server is blocked when the current processing unit temperature is greater than the target reference temperature, and the deviation between the current processing unit temperature and the target reference temperature is greater than a set deviation threshold.
[0073] In an embodiment, the device further comprises a flow analysis module configured to:
[0074] Obtain liquid flow sequence data of a server cluster, wherein the server cluster comprises at least two liquid-cooled servers.
[0075] Calculate a cluster flow change trend based on the liquid flow sequence data.
[0076] When the cluster flow change trend is a continuous decrease in flow, determine that the server cluster has a systematic blockage problem.
[0077] Optionally, the reference temperature determination module 22 is specifically configured to determine a target reference temperature based on a preset mapping relationship between liquid-cooled server parameters and processing unit reference temperatures, according to the current single-server liquid inlet temperature, the current processing unit load rate, and the current single-server liquid flow rate, when the server cluster has the systematic blockage problem.
[0078] Optionally, the server cluster is composed of all liquid-cooled servers served by a cold quantity distribution unit, or the server cluster is composed of part of the liquid-cooled servers served by the cold quantity distribution unit. The cooling liquid is delivered by the cold quantity distribution unit to the cold plate of the liquid-cooled server, serves heat dissipation of the liquid-cooled server, and is returned to the cold quantity distribution unit from the cold plate.
[0079] In an embodiment, the mapping relationship between the liquid-cooled server parameters and the processing unit reference temperatures comprises a mapping relationship in a prototype measurement scenario.
[0080] The mapping relationship in the prototype measurement scenario is determined by the following method:
[0081] Obtain a mapping relationship between liquid-cooled server parameters and processing unit reference temperatures in an ideal standard scenario.
[0082] constructing a prototype measurement scene, under the prototype measurement scene, adjusting the reference single-server liquid inlet temperature, the reference single-server liquid flow rate and the reference processing unit load rate of the liquid-cooled server, and after each adjustment, obtaining the processing unit reference temperature of the liquid-cooled server;
[0083] According to the reference single-server liquid inlet temperature, the reference single-server liquid flow rate, the reference processing unit load rate and the processing unit reference temperature, the mapping relationship under the ideal standard scene is adjusted to obtain the mapping relationship under the prototype measurement scene.
[0084] In an embodiment, the mapping relationship between the liquid-cooled server parameters and the processing unit reference temperature includes a mapping relationship under an ideal standard scene and a mapping relationship under a prototype measurement scene.
[0085] The reference temperature determination module 22 is specifically configured to:
[0086] Based on the mapping relationship under the prototype measurement scene, the first reference temperature is determined according to the current single-server liquid inlet temperature, the current processing unit load rate and the current single-server liquid flow rate.
[0087] Based on the mapping relationship under the ideal standard scene, the second reference temperature is determined according to the current single-server liquid inlet temperature, the current processing unit load rate and the current single-server liquid flow rate.
[0088] Based on the processing unit initial temperature data of the liquid-cooled server, the target processing unit initial temperature is determined according to the current processing unit load rate; wherein the processing unit initial temperature data is a set of processing unit initial temperatures collected by the liquid-cooled server at the initial stage of operation under different parameter conditions; the parameter conditions include initial single-server liquid inlet temperature, initial single-server liquid flow rate and initial processing unit load rate, and the variables in the parameter conditions include the initial processing unit load rate.
[0089] The target processing unit initial temperature is subtracted from the second reference temperature to obtain an error adjustment value.
[0090] The first reference temperature is added to the error adjustment value to obtain a target reference temperature.
[0091] In an embodiment, the device is a cloud monitoring platform; the current single-server liquid flow rate is detected by a flow meter of a machine room monitoring system and uploaded to the cloud monitoring platform by the machine room monitoring system; and the current processing unit temperature is monitored by a server monitoring system and uploaded to the cloud monitoring platform by the server monitoring system.
[0092] In an embodiment, the cooling liquid is delivered by a cooling capacity distribution unit to the cold plate of the liquid-cooled server, serves heat dissipation of the liquid-cooled server, and is returned to the cooling capacity distribution unit from the cold plate; the total amount of the cooling liquid supplied by the cooling capacity distribution unit is in a positive correlation with the total amount of servers served by the cooling capacity distribution unit.
[0093] In an embodiment, the current temperature of the processing unit is a current housing temperature of the processing unit; and the processing unit of the liquid-cooled server is a central processing unit and / or a graphics processing unit.
[0094] In an embodiment, the device further comprises a pre-warning module configured to generate and display cold plate blockage warning information to prompt a staff to take a pipe dredging operation after determining that the cold plate pipe is blocked.
[0095] Further, the warning level of the cold plate blockage warning information is positively correlated with the deviation degree between the current temperature of the processing unit and the target reference temperature.
[0096] It is worth noting that the working principle of the server cold plate blockage detection device provided in the above embodiments can refer to the working process of the server cold plate blockage detection method provided in any of the above embodiments, which will not be repeated here.
[0097] Compared with the prior art, the server cold plate blockage detection device provided in the embodiments of the present application can determine the target reference temperature suitable for the current working condition according to the current single-server liquid inlet temperature, the current processing unit load rate, and the current single-server liquid flow rate of the liquid-cooled server, and then compare the current temperature of the processing unit to analyze whether there is temperature abnormality caused by cold plate blockage, so as to identify the cold plate blockage risk in time, so that the staff can take maintenance measures in advance, effectively reduce the risk of simultaneous failure of batch servers, greatly reduce the serious interference to the business carried by the servers, and ensure the business continuity and stability.
[0098] Referring to Figure 3 , the embodiments of the present application also provide a server cold plate blockage detection device, which comprises a processor 31, a memory 32, and a computer program stored in the memory 32 and configured to be executed by the processor 31, and the processor 31 implements the steps in the above server cold plate blockage detection method embodiments when executing the computer program, such as S11-S13 in Figure 1 ; or the processor 31 implements the functions of each module in the above device embodiments when executing the computer program.
[0099] For example, the computer program can be divided into one or more modules, which are stored in the memory 32 and executed by the processor 31 to complete the present application. The one or more modules can be a series of computer program instruction segments capable of completing a specific function, which are used to describe the execution process of the computer program in the server cold plate blockage detection device. For example, the computer program can be divided into multiple modules, and the specific working process of each module can refer to the working process of the server cold plate blockage detection device described in the above embodiments, which will not be described here.
[0100] The server cold plate blockage detection device can be a desktop computer, a notebook computer, a palm computer, a cloud server and the like. The server cold plate blockage detection device can include, but is not limited to, a processor 31 and a memory 32. Those skilled in the art can understand that the server cold plate blockage detection device can also include input / output devices, network access devices, buses and the like.
[0101] The processor 31 can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor and the like. The processor 31 is the control center of the server cold plate blockage detection device, which connects all parts of the server cold plate blockage detection device through various interfaces and lines.
[0102] The memory 32 can be used to store the computer programs and / or modules, and the processor 31 realizes various functions of the server cold plate blockage detection device by running or executing the computer programs and / or modules stored in the memory 32, and calling the data stored in the memory 32. The memory 32 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application program required by a function (such as an image playing function, etc.), and the like; and the data storage area can store data created according to the use of the mobile phone, etc. In addition, the memory 32 can include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0103] The modules integrated in the server cold plate blockage detection device can be stored in a computer readable storage medium if they are realized in the form of software function units and sold or used as independent products. Based on this understanding, all or part of the processes in the above-mentioned embodiment methods can also be completed by a computer program instructing related hardware, and the computer program can be stored in a computer readable storage medium. The computer program can realize the steps of each method embodiment when executed by the processor 31. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or some intermediate forms, etc. The computer readable medium can include any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0104] The embodiment of the present application also provides a computer program product, which includes computer programs / instructions, and the computer programs / instructions are executed by a processor to realize the server cold plate blockage detection method according to any one of the above-mentioned embodiments.
[0105] The above-mentioned is the preferred embodiment of the present application, and it should be pointed out that, for ordinary skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, and these improvements and refinements are also considered to be within the protection scope of the present application.
Claims
1. A method of server cold plate blockage detection, the method comprising: The method comprises the following steps: obtaining the current single-server inlet liquid temperature, the current processing unit load rate, the current single-server liquid flow rate and the current processing unit temperature of the liquid-cooled server; wherein the current single-server inlet liquid temperature is the current cooling liquid temperature entering the corresponding cold plate of the liquid-cooled server, and the current single-server liquid flow rate is the current cooling liquid flow rate entering the corresponding cold plate of the liquid-cooled server; determining a target reference temperature according to the current single-server inlet liquid temperature, the current processing unit load rate and the current single-server liquid flow rate based on a preset mapping relationship between the liquid-cooled server parameters and the processing unit reference temperature; the mapping relationship between the liquid-cooled server parameters and the processing unit reference temperature comprises a mapping relationship between the liquid-cooled server parameters and the processing unit reference temperature in a prototype measurement scene or a mapping relationship between the liquid-cooled server parameters and the processing unit reference temperature in an ideal standard scene; the ideal standard scene refers to a highly controllable standardized experimental scene, and the prototype measurement scene refers to a scene close to the actual state of the liquid-cooled server before leaving the factory; when the current processing unit temperature is greater than the target reference temperature, and the deviation between the current processing unit temperature and the target reference temperature is greater than a set deviation threshold, the cold plate pipeline of the liquid-cooled server is blocked.
2. The server cold plate blockage detection method of claim 1, wherein, The method further comprises the following steps: obtaining liquid flow sequence data of a server cluster; wherein the server cluster comprises at least two liquid-cooled servers; calculating a cluster flow change trend according to the liquid flow sequence data; when the cluster flow change trend is a continuous decrease in flow, determining that the server cluster has a systematic blockage problem.
3. The server cold plate blockage detection method of claim 2, wherein, The method of determining a target reference temperature according to the current single-server inlet liquid temperature, the current processing unit load rate and the current single-server liquid flow rate based on a preset mapping relationship between the liquid-cooled server parameters and the processing unit reference temperature comprises the following steps: when the server cluster has the systematic blockage problem, determining a target reference temperature according to the current single-server inlet liquid temperature, the current processing unit load rate and the current single-server liquid flow rate based on a preset mapping relationship between the liquid-cooled server parameters and the processing unit reference temperature.
4. The server cold plate blockage detection method of claim 2, wherein, The server cluster is composed of all liquid-cooled servers served by a cold energy distribution unit; or the server cluster is composed of part of the liquid-cooled servers served by the cold energy distribution unit; cooling liquid is delivered by the cold energy distribution unit to the cold plate of the liquid-cooled server, serves heat dissipation of the liquid-cooled server, and is returned to the cold energy distribution unit from the cold plate.
5. The server cold plate blockage detection method of claim 1, wherein, The mapping relationship between the liquid-cooled server parameters and the processing unit reference temperature comprises a mapping relationship in a prototype measurement scene; The mapping relationship in the prototype measurement scene is determined by the following method: obtaining a mapping relationship between the liquid-cooled server parameters and the processing unit reference temperature in an ideal standard scene; constructing a prototype measurement scene, adjusting the reference single-server inlet liquid temperature, the reference single-server liquid flow rate and the reference processing unit load rate of the liquid-cooled server in the prototype measurement scene, and obtaining the processing unit reference temperature of the liquid-cooled server after each adjustment; Adjust the mapping relationship under the ideal standard scene according to the reference single-serve liquid inlet temperature, the reference single-serve liquid flow, the reference processing unit load rate and the processing unit reference temperature, to obtain the mapping relationship under the prototype measured scene.
6. The server cold plate blockage detection method of any one of claims 1-5, wherein, The mapping relationship between the liquid-cooled server parameter and the processing unit reference temperature comprises a mapping relationship under an ideal standard scene and a mapping relationship under a prototype measured scene. The method for determining the target reference temperature according to the current single-serve liquid inlet temperature, the current processing unit load rate and the current single-serve liquid flow based on the preset mapping relationship between the liquid-cooled server parameter and the processing unit reference temperature comprises: Determine a first reference temperature according to the current single-serve liquid inlet temperature, the current processing unit load rate and the current single-serve liquid flow based on the mapping relationship under the prototype measured scene. Determine a second reference temperature according to the current single-serve liquid inlet temperature, the current processing unit load rate and the current single-serve liquid flow based on the mapping relationship under the ideal standard scene. Determine a target processing unit initial temperature according to the current processing unit load rate based on processing unit initial temperature data of the liquid-cooled server; wherein the processing unit initial temperature data is a set of processing unit initial temperatures collected by the liquid-cooled server at an initial stage of operation under different parameter conditions; the parameter conditions comprise an initial single-serve liquid inlet temperature, an initial single-serve liquid flow and an initial processing unit load rate, and the variables in the parameter conditions comprise the initial processing unit load rate. Subtract the second reference temperature from the target processing unit initial temperature to obtain an error adjustment value. Add the error adjustment value to the first reference temperature to obtain the target reference temperature.
7. The server cold plate blockage detection method of claim 1, wherein, The method is applied to a cloud monitoring platform; the current single-serve liquid flow is detected by a flow meter of a machine room monitoring system and uploaded to the cloud monitoring platform by the machine room monitoring system; and the processing unit current temperature is monitored by a server monitoring system and uploaded to the cloud monitoring platform by the server monitoring system.
8. The server cold plate blockage detection method of claim 1, wherein, Further comprising: The cooling liquid is delivered by a cold quantity distribution unit to a cold plate of the liquid-cooled server, serves heat dissipation of the liquid-cooled server, and is returned to the cold quantity distribution unit from the cold plate; The total liquid supply amount of the cold quantity distribution unit is in a positive correlation with the total amount of servers served by the cold quantity distribution unit.
9. The server cold plate blockage detection method of claim 1, wherein, The processing unit current temperature is a current processing unit shell temperature; and the processing unit of the liquid-cooled server is a central processing unit and / or a graphics processing unit.
10. The server cold plate blockage detection method of claim 1, wherein, Further comprising: After determining that the cold plate pipeline is blocked, generate and display cold plate blockage alarm information to prompt staff to take pipeline dredging operations.
11. The server cold plate blockage detection method of claim 10, wherein, The alarm level of the cold plate blockage alarm information is positively correlated with the degree of deviation between the processing unit current temperature and the target reference temperature.
12. A server cold plate blockage detection apparatus, comprising: Further comprising: The data acquisition module is configured to acquire a current single-server liquid inlet temperature, a current processing unit load rate, a current single-server liquid flow rate, and a current processing unit temperature of the liquid-cooled server. The current single-server liquid inlet temperature is a current cooling liquid temperature entering a corresponding cold plate of the liquid-cooled server, and the current single-server liquid flow rate is a current cooling liquid flow rate entering the corresponding cold plate of the liquid-cooled server. The mapping relationship between the liquid-cooled server parameters and the processing unit reference temperature includes a mapping relationship between the liquid-cooled server parameters and the processing unit reference temperature in a prototype measurement scene or a mapping relationship between the liquid-cooled server parameters and the processing unit reference temperature in an ideal standard scene. The ideal standard scene refers to a highly controllable standardized experimental scene, and the prototype measurement scene refers to a scene close to an actual state of the liquid-cooled server before being shipped. The reference temperature determination module is configured to determine a target reference temperature based on a preset mapping relationship between liquid-cooled server parameters and processing unit reference temperatures, according to the current single-server liquid inlet temperature, the current processing unit load rate, and the current single-server liquid flow rate. The clogging detection module is configured to determine that a cold plate pipeline of the liquid-cooled server is clogged when the current processing unit temperature is greater than the target reference temperature, and a deviation degree between the current processing unit temperature and the target reference temperature is greater than a set deviation threshold.
13. A server cold plate blockage detection apparatus, comprising: The computer readable storage medium includes a stored computer program, wherein the computer readable storage medium controls a device in which the computer readable storage medium is located to perform the server cold plate clogging detection method according to any one of claims 1 to 11 when the computer program runs.
14. A computer readable storage medium characterized by, The computer program / instruction is executed by the processor to implement the server cold plate clogging detection method according to any one of claims 1 to 11.
15. A computer program product comprising computer programs / instructions, characterized in that,
Citation Information
Patent Citations
Liquid cooling server intelligent temperature control method based on local software monitoring and liquid cooling server
CN119536482A
Fault monitoring method and device of liquid cooling system and server system
CN119958889A