Fault early warning method and system for cooling unit of server cabinet

By setting up measurement points in server cabinets, monitoring operating data and images, using temporal convolutional networks to analyze abnormal features, and combining data comparison and judgment, the problems of low accuracy and high false alarm rate in existing technologies are solved, and efficient fault warning and improved operation and maintenance quality are achieved.

CN120673115AInactive Publication Date: 2025-09-19SHENZHEN JINGXINLONG HARDWARE PRODUCTS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510601324.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-09-19
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing server cabinet early warning technology has low accuracy, poor precision, is prone to false alarms, and cannot effectively identify cabinet cooling unit failures.

Method used

By setting up measurement points in the cabinet to monitor operating data and images, a temporal convolutional network is used to analyze abnormal image features, and combined with operating data comparison and judgment, the corresponding early warning results are output.

Benefits of technology

It achieves more accurate and efficient fault diagnosis and early warning, reduces the probability of misjudgment, improves operation and maintenance quality, and reduces operation and maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673115A_ABST
    Figure CN120673115A_ABST
Patent Text Reader

Abstract

The invention discloses a server cabinet cooling unit fault early warning method and system, and the method comprises the steps: setting a measurement point at a cabinet, and monitoring and obtaining the operation data and operation images of the cabinet; acquiring historical operation data, calculating a normal operation data range, acquiring historical operation images, and analyzing abnormal image features by adopting a time convolution network; and when the operation data exceeds a threshold, comparing and judging the operation data, when the operation data exceeds the threshold, comparing and judging the operation image in the same time period, and outputting a corresponding early warning result based on a comparison result. According to the method, more accurate and efficient fault judgment and early warning are realized by performing composite analysis on the operation parameters and the images of the cabinet, the misjudgment probability is effectively reduced, the operation and maintenance quality is further improved, and the operation and maintenance cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of cabinet early warning technology, and in particular to a method and system for early warning of a server cabinet cooling unit failure. Background Art

[0002] In recent years, with the rapid development of internet technology, whole-cabinet servers have become increasingly popular due to their low cost, high integration, centralized management, centralized cooling, and easy operation and maintenance. Centralized cooling and management particularly reduce management and maintenance costs. However, while this high level of integration improves convenience, it also concentrates risks and hidden dangers. If problems such as overheating or leakage occur in a part of the cabinet, without timely warning and response, the damage can easily spread to other servers in the cabinet, leading to further losses. Therefore, comprehensive monitoring and management of server cabinets is necessary to reduce the risk of accidents spreading.

[0003] Existing server cabinet early warning technology typically involves installing measurement points to directly monitor relevant operating parameters and issuing warnings when these parameters exceed warning values. This approach is relatively simple, direct, and easy to implement, but it suffers from low accuracy, poor precision, and the susceptibility to false alarms. Summary of the Invention

[0004] This application mainly provides a server cabinet cooling unit failure warning method and system to solve the above problems.

[0005] To solve the above technical problems, a technical solution adopted by the present application is to provide a server cabinet cooling unit failure early warning method, comprising the steps of:

[0006] S10: Setting measurement points in the cabinet to monitor and obtain operation data and operation images of the cabinet;

[0007] S20: Acquire historical operation data, calculate the normal operation data range, and simultaneously acquire historical operation images, and use a temporal convolutional network to analyze abnormal image features;

[0008] S30: performing a comparison and determination on the operation data. When the operation data exceeds a threshold, performing a comparison and determination on the operation images of the same period, and outputting a corresponding warning result based on the comparison result.

[0009] In a possible implementation, the step of setting a measurement point in the cabinet to monitor and obtain the operating data and operating image of the cabinet includes:

[0010] S11: performing unified timing on the measuring points, and establishing a mapping relationship between the operation image, the operation data and the actual position of the cabinet.

[0011] In a possible implementation, the step of setting measurement points in the cabinet to monitor and obtain operating data and operating images of the cabinet further includes:

[0012] S12: Acquire an infrared image and a visible light image of the cabinet, and enhance features of the infrared image and the visible light image by weighted fusion;

[0013] S13: Analyze the motion direction of the feature based on its texture and dynamics.

[0014] In one possible implementation, the step of comparing and determining the operating data, and when the operating data exceeds a threshold, comparing and determining the operating images of the same period, and outputting a corresponding warning result based on the comparison result, includes:

[0015] S31: When the operation data exceeds a threshold, the operation image of the same period is retrieved for analysis. If the analysis result of the operation image indicates that there is a visual leak, a level 1 warning is output;

[0016] S32: Continuously monitor the operating data, and if the difference exceeding the threshold value continues to increase, output a second-level warning.

[0017] In a possible implementation, the step of when the operating data exceeds a threshold includes:

[0018] S33: Plotting a normal value range based on the historical operating data, and obtaining its similarity with the operating data through dynamic time warping;

[0019] S34: Calculate the comprehensive leakage assessment value and compare it with the threshold value to determine whether there is a leakage.

[0020] In a possible implementation, after the step of continuously monitoring the operating data and outputting a level 2 warning if the difference exceeding the threshold value continues to increase, the step further includes:

[0021] S35: Continuously monitor the operating data, and if the increasing rate of the comprehensive leakage assessment value is positive within a predetermined time, output a third-level warning.

[0022] In a possible implementation, the step of when the operating data exceeds a threshold further includes:

[0023] S36: If the change speed of the operating data exceeds a threshold, a third-level warning is output.

[0024] In a possible implementation, after the steps of obtaining historical operation data, calculating the normal operation data range, and simultaneously obtaining historical operation images, and analyzing abnormal image features using a temporal convolutional network, the method further includes:

[0025] S21: storing the currently acquired real-time operation data and operation image, and using them to update the historical operation data and the historical operation image.

[0026] In a possible implementation, the operation data includes node pressure, cold source pressure, cold source flow, cold source flow velocity, cold source temperature, and humidity outside the pipe; and the operation image includes an infrared image and a visible light image.

[0027] To solve the above technical problems, another technical solution adopted by the present application is to provide a server cabinet cooling unit failure warning system, which is applicable to the server cabinet cooling unit failure warning method described above, including:

[0028] The monitoring module includes measuring points and cameras, which are used to monitor the operating data of the cabinet and obtain operating images respectively;

[0029] A storage module, configured to store the operation data and the operation image;

[0030] A preprocessing module, configured to process the running image to analyze features;

[0031] An analysis module, configured to compare and analyze the operation data and the operation image, and output a determination result;

[0032] The early warning module outputs a warning of the corresponding situation based on the determination result.

[0033] The beneficial effect of the present application is that, different from the prior art, the present application discloses a server cabinet cooling unit fault warning method and system, which realizes more accurate and efficient fault judgment and warning by performing a composite analysis of the operating parameters and images of the cabinet, effectively reduces the probability of misjudgment, further improves the operation and maintenance quality, and reduces the operation and maintenance costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. Those skilled in the art can also derive other drawings based on these drawings without inventive work, among which:

[0035] Figure 1 This is a flow chart of a server cabinet cooling unit failure warning method according to an embodiment of the present application. DETAILED DESCRIPTION

[0036] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0037] The terms "first", "second" and "third" in the embodiments of the present application are only used for descriptive purposes and are not to be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first", "second" and "third" may explicitly or implicitly include at least one of the features. In the description of the present application, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device comprising a series of steps or units is not limited to the listed steps or units, but may optionally also include steps or units that are not listed, or may optionally also include other steps or units inherent to these processes, methods, products or devices.

[0038] References to "embodiments" herein mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0039] See also Figure 1 , the embodiment of the present application includes: a server cabinet cooling unit failure warning method, comprising the steps of:

[0040] S10: Setting measurement points in the cabinet to monitor and obtain operating data and images of the cabinet;

[0041] S20: Acquire historical operation data, calculate the normal operation data range, and simultaneously acquire historical operation images, and use a temporal convolutional network to analyze abnormal image features;

[0042] S30: Compare and judge the operation data. When the operation data exceeds the threshold, compare and judge the operation images of the same period, and output a corresponding warning result based on the comparison result.

[0043] Specifically, in step S10, various measurement points are set up at key locations within the cabinet. For example, pressure measurement points are set up at each pipeline node and within the cooling source; flow measurement points, flow velocity measurement points, and temperature measurement points are set up within the cooling source; and humidity measurement points are set up outside the pipeline and on the outer casing. These measurement points monitor operational data such as pressure, flow, flow velocity, temperature, or humidity. The operational data obtained by these measurement points can directly or indirectly determine whether there are any anomalies or leaks in the cabinet. Therefore, a comprehensive assessment is subsequently conducted to improve the accuracy of the judgment.

[0044] The measuring point also includes a camera. The camera in this embodiment includes an infrared camera and a visible light camera, which are used to capture infrared images and visible light images respectively for further judgment of leakage.

[0045] In step S20, historical data and images are acquired. Based on the different operating states in the historical operating data and images, the data characteristics of normal and abnormal operation are summarized and analyzed to set the abnormality threshold for the operating data. This analysis process can be performed using a temporal convolutional network (TCN). This method has the advantages of high computational efficiency and greater stability, and it uses a multi-step analysis process to obtain features from the image.

[0046] In step S30, the operating data is first evaluated. If multiple elements of the operating data exceed thresholds, an image review is initiated, analyzing images from the same time period. If the images from the same time period are judged to be abnormal, an abnormality is considered to have occurred, and a Level 1 warning is issued. Further monitoring of the operating data is also performed. If the operating data trend continues to deviate from the threshold, the abnormality is considered to have worsened, and a Level 2 warning is issued. This phased evaluation and warning process allows for accurate abnormality assessment and reduces the probability of misjudgment.

[0047] In one embodiment, the steps of setting measurement points in a cabinet to monitor and obtain operating data and operating images of the cabinet include:

[0048] S11: Perform unified timing for the measurement points and establish a mapping relationship between the operation image, operation data and the actual position of the cabinet.

[0049] Specifically, unified timing is implemented for all measurement points, sensors, cameras, and other devices to synchronize timestamps and ensure that data from different sources for the same event is aligned on the timeline. A mapping relationship is established to correspond the pixel coordinates of the collected operational images to their physical locations, ensuring that the detected features in the images can be located at the actual locations on the cabinet. This step preprocesses the acquired data and images to facilitate subsequent data organization and analysis.

[0050] In one embodiment, the step of setting a measurement point in the cabinet to monitor and obtain the operating data and operating image of the cabinet further includes:

[0051] S12: Obtain an infrared image and a visible light image of the cabinet, and enhance features of the infrared image and the visible light image by weighted fusion;

[0052] S13: Analyze the motion direction of the feature based on its texture and dynamics.

[0053] Specifically, the images processed in step S12 include infrared and visible light images. The infrared image focuses on detecting and identifying areas with abnormal temperatures, while the visible light image is used to detect drips and traces of liquid leaks. Weighted fusion enhances features. In this embodiment, the two images are aligned and fused based on OpenCV, with a weight of 0.6 for the visible light image and 0.4 for the infrared image. This setting adjusts the importance of different images. It is generally believed that visible light images have more features, while infrared images focus more on temperature detection and involve fewer feature elements, so they are assigned lower weights.

[0054] In step S13, the running image is trained to extract texture and motion characteristics of features, and based on these characteristics, objects such as droplets, water vapor, and water stains are detected. This embodiment can use the YOLO model to detect objects in the image, extract deep image features, and combine optical flow to analyze the flow direction of the liquid, further determining whether there is leakage or dripping, and determining the specific characteristics of the leak.

[0055] In one embodiment, the steps of comparing and determining the operating data, comparing and determining the operating images of the same period when the operating data exceeds a threshold, and outputting a corresponding warning result based on the comparison result include:

[0056] S31: When the operating data exceeds the threshold, the operating image of the same period is retrieved for analysis. If the analysis result of the operating image indicates that there is a visual leak, a level 1 warning is output;

[0057] S32: Continuously monitor the operating data. If the difference exceeding the threshold value continues to increase, a level 2 warning is output.

[0058] Specifically, in step S31, a preliminary judgment is made on the operation data, and a threshold comparison can be performed on several items in the operation data. In one embodiment, when more than three parameters exceed the threshold, an image review can be triggered, and the operation images of the same period are analyzed. If the operation images of this period are also determined to be leaks, a first-level warning is output. If the operation images of this period are not determined to be leaks, no warning is issued.

[0059] In step S32, after issuing a first-level warning, the operating data is continuously monitored. If the parameters exceeding the threshold continue to deviate from the normal value range, or more parameters exceed the threshold, it can be considered that the leakage anomaly is more serious, and a second-level warning is output.

[0060] In one embodiment, when the operating data exceeds a threshold, the steps include:

[0061] S33: Based on the historical operating data, plot the normal value range and obtain its similarity with the operating data through dynamic time warping (DTW);

[0062] S34: Calculate the comprehensive leakage assessment value and compare it with the threshold value to determine whether there is a leakage.

[0063] Specifically, in step S33, parameters of a normal state may be obtained based on historical operating data and compared with the current operating parameters to obtain similarity.

[0064] In another embodiment, the operating data and the image may be comprehensively judged, and the formula for calculating the comprehensive leakage assessment value P is as follows:

[0065] P=α·S d +β·S i +γ·S t

[0066] Among them, α, β, and γ are the first coefficient, the second coefficient, and the third coefficient respectively; S d S is the comprehensive score of the running data. i S is the confidence score of the running image. t is the DTW similarity (time series matching). The comprehensive operational data score is the total score obtained after weighted evaluation of all measurement points; the confidence score is the confidence of the YOLO model, which specifically includes category confidence and existence confidence. The former is the probability of determining whether it is a specific target (such as a droplet or water vapor), and the latter is the probability of determining whether any target (not limited to droplets, water vapor, or liquid traces) exists. The first coefficient, second coefficient, and third coefficient can be adjusted based on the actual environment and application scenario to adjust the weight of the corresponding parameters.

[0067] In one embodiment, the operation data is continuously monitored. If the difference exceeding the threshold value continues to increase, after the step of outputting a second-level warning, the method further includes:

[0068] S35: Continuously monitor the operating data. If the rate of increase of the comprehensive leakage assessment value is positive within a predetermined time, output a level 3 warning.

[0069] Specifically, during the continuous monitoring of the operating data, if it is found that the comprehensive leakage assessment value continues to increase and the rate of increase is positive, it means that there is an abnormality and the abnormal situation continues to worsen, such as pipe burst, breakage, pump stoppage and other abnormalities. At this time, a more advanced warning should be output to remind relevant personnel.

[0070] In one embodiment, when the operating data exceeds a threshold, the step further includes:

[0071] S36: If the change rate of the operating data exceeds the threshold, a level 3 warning is output.

[0072] This step is similar to step S35 and is a supplementary determination of serious abnormal situations. If the detected initial operating data has obvious serious deviations, a level 3 warning is directly issued to further improve the response speed.

[0073] The warning forms in this embodiment include multiple types, and can correspond to different degrees and forms of prompt information, sound and light warnings, etc. from level 1 to level 3 warnings. The severity of the abnormality represented by level 1 to level 3 warnings increases step by step, and level 3 warnings include at least multiple forms of warning reminders.

[0074] In one embodiment, after the steps of obtaining historical operation data, calculating the normal operation data range, and simultaneously obtaining historical operation images, and analyzing abnormal image features using a temporal convolutional network, the method further includes:

[0075] S21: storing the currently acquired real-time operation data and operation image, and using them to update the historical operation data and historical operation image.

[0076] The operation data is continuously stored and used as historical operation data to continuously correct the historical operation data to compensate for the data changes caused by the aging of the equipment itself with use, the data changes caused by seasonal environment, etc., and further improve the accuracy of abnormality judgment.

[0077] In one embodiment, the operation data includes node pressure, cold source pressure, cold source flow, cold source flow velocity, cold source temperature, and humidity outside the pipe, and the operation image includes an infrared image and a visible light image.

[0078] A server cabinet cooling unit fault warning system, applicable to the server cabinet cooling unit fault warning method described above, comprises:

[0079] The monitoring module includes measuring points and cameras, which are used to monitor the operating data of the cabinet and obtain operating images respectively;

[0080] A storage module, used for storing operation data and operation images;

[0081] A preprocessing module is used to process the running image to analyze the features;

[0082] The analysis module is used to compare and analyze the operation data and operation images and output the judgment results;

[0083] The early warning module outputs warnings of the corresponding situation based on the judgment results.

[0084] The above description is merely an embodiment of the present application and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A server cabinet cooling unit failure early warning method, characterized in that: Including steps: S10: Setting measurement points in the cabinet to monitor and obtain operation data and operation images of the cabinet; S20: Acquire historical operation data, calculate the normal operation data range, and simultaneously acquire historical operation images, and use a temporal convolutional network to analyze abnormal image features; S30: performing a comparison and determination on the operation data. When the operation data exceeds a threshold, performing a comparison and determination on the operation images of the same period, and outputting a corresponding warning result based on the comparison result.

2. The server cabinet cooling unit failure early warning method according to claim 1, characterized in that: The step of setting a measuring point on the cabinet to monitor and obtain the operating data and operating image of the cabinet includes: S11: performing unified timing on the measuring points, and establishing a mapping relationship between the operation image, the operation data and the actual position of the cabinet.

3. The server cabinet cooling unit failure early warning method according to claim 1, characterized in that: The step of setting a measuring point in the cabinet to monitor and obtain the operating data and operating image of the cabinet also includes: S12: Acquire an infrared image and a visible light image of the cabinet, and enhance features of the infrared image and the visible light image by weighted fusion; S13: Analyze the motion direction of the feature based on its texture and dynamics.

4. The server cabinet cooling unit failure early warning method according to claim 1, characterized in that: The step of comparing and judging the operation data, and when the operation data exceeds a threshold, comparing and judging the operation images of the same period, and outputting a corresponding warning result based on the comparison result, includes: S31: When the operation data exceeds a threshold, the operation image of the same period is retrieved for analysis. If the analysis result of the operation image indicates that there is a visual leak, a level 1 warning is output; S32: Continuously monitor the operating data, and if the difference exceeding the threshold value continues to increase, output a second-level warning.

5. The server cabinet cooling unit failure early warning method according to claim 4, characterized in that: The step of when the operating data exceeds a threshold comprises: S33: Plotting a normal value range based on the historical operating data, and obtaining its similarity with the operating data through dynamic time warping; S34: Calculate the comprehensive leakage assessment value to determine whether there is a leakage.

6. The server cabinet cooling unit failure early warning method according to claim 5, characterized in that: After the step of continuously monitoring the operating data and outputting a second-level warning if the difference exceeding the threshold value continues to increase, the method further includes: S35: Continuously monitor the operating data, and if the increasing rate of the comprehensive leakage assessment value is positive within a predetermined time, output a third-level warning.

7. The server cabinet cooling unit failure early warning method according to claim 4, characterized in that: The step of when the operating data exceeds a threshold value further includes: S36: If the change speed of the operating data exceeds a threshold, a third-level warning is output.

8. The server cabinet cooling unit failure early warning method according to claim 1, characterized in that: After the steps of obtaining historical operation data, calculating the normal operation data range, obtaining historical operation images, and analyzing abnormal image features using a temporal convolutional network, the method further includes: S21: storing the currently acquired real-time operation data and operation image, and using them to update the historical operation data and the historical operation image.

9. The server cabinet cooling unit failure early warning method according to claim 1, characterized in that: The operation data includes node pressure, cold source pressure, cold source flow, cold source flow velocity, cold source temperature, and humidity outside the pipe; the operation images include infrared images and visible light images.

10. A server cabinet cooling unit failure warning system, applicable to the server cabinet cooling unit failure warning method according to any one of claims 1 to 9, characterized in that: include: The monitoring module includes measuring points and cameras, which are used to monitor the operating data of the cabinet and obtain operating images respectively; A storage module, configured to store the operation data and the operation image; A preprocessing module, configured to process the running image to analyze features; An analysis module, configured to compare and analyze the operation data and the operation image, and output a determination result; The early warning module outputs a warning of the corresponding situation based on the determination result.