A data center CDU system fault identification method

By collecting multi-parameter data in the data center CDU system, establishing a multivariate linear regression model and dynamic thresholds, and identifying system anomalies, the problem of the inability to accurately identify system failures in existing technologies is solved, early warning and false alarm reduction are achieved, and maintenance costs are reduced.

CN120560898BActive Publication Date: 2025-10-10SICHUAN CRUN CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511062185.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-10-10
Estimated Expiration
2045-07-31

AI Technical Summary

Technical Problem

In the existing technology, fault identification of the data center CDU system mainly relies on the output of a single component, which cannot accurately identify faults of the entire system, resulting in false alarms and unplanned downtime.

Method used

By collecting historical data on motor frequency, valve opening, input power, current, cooling water temperature, and ambient temperature and humidity, a multivariate linear regression model is established, dynamic thresholds are calculated, and system anomalies are identified based on seasonal and load changes. Multi-parameter collaborative analysis and dynamic threshold adjustment are used to improve fault identification accuracy.

Benefits of technology

It can issue early warnings hours to days before a fault occurs, reduce false alarms, improve fault detection rates, lower maintenance costs, and reduce unnecessary downtime.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120560898B_ABST
    Figure CN120560898B_ABST
Patent Text Reader

Abstract

The application discloses a data center CDU system fault identification method, and belongs to the technical field of data center CDU, which comprises the following steps: S1, data acquisition and feature extraction; S2, establishing a basic working condition model, establishing a reference relationship model of power and current, and respectively establishing quantitative relationships between power / current and motor frequency and valve opening degree by using a multiple linear regression model; S3, selecting a dynamic threshold, obtaining a working condition model based on S2 and obtaining historical statistical data based on S1, calculating residual errors epsilon under different working conditions, and calculating dynamic threshold data based on corresponding working condition residual errors epsilon; S4, establishing a fault feature engineering, calculating dynamic thresholds in different seasons, and identifying features indicating faults, including: the absolute value of the residual error exceeding the threshold value, and the residual error variance significantly increasing; and S5, establishing a fault detection model, and determining that the system is abnormal when the residual error exceeds the threshold value. The application solves the problem in the prior art that a single element itself output fault cannot identify the whole system fault.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of data center CDU technology, and in particular relates to a method for identifying system faults in a data center CDU. Background Art

[0002] With the advancement of information technology, artificial intelligence has become an indispensable part of social development. Data centers, as their physical carrier, play a key role in this process. The CDU, as the cooling distribution unit in a data center, plays a vital role in its normal operation.

[0003] Currently, abnormal diagnosis of CDU system operation in data centers plays a vital role in the normal operation of the entire data center. Traditional system fault diagnosis mainly relies on feedback from component operation data to determine the fault.

[0004] The method proposed in this paper establishes a fault identification model based on multiple factors, including motor frequency, valve opening, input power, and current, and adjusts sensitivity to accurately and promptly identify faults, reducing false alarms and unplanned downtime. Furthermore, the model can be optimized based on environmental conditions and different loads, improving the adaptability of the fault identification method. Summary of the Invention

[0005] The purpose of this application is to overcome the problems of the prior art and disclose a method for identifying faults in a data center CDU system. The method of this application solves the technical problem in the prior art of relying on the output fault of a single component itself and being unable to identify the fault of the entire system.

[0006] The purpose of this application is achieved through the following technical solutions:

[0007] A data center CDU system fault identification method, the data center CDU system fault identification method comprising:

[0008] S1: Data acquisition and feature extraction: Collect historical data on motor frequency, valve opening, input power, current, cooling water inlet and outlet temperatures, and ambient temperature and humidity during normal operation of the data center CDU system, and extract key features: the mean and variance of power and current.

[0009] S2: Establish a basic operating condition model. When the motor frequency and valve opening are stable, establish a benchmark relationship model between power and current. Use a multiple linear regression model to establish the quantitative relationship between power P / current I and motor frequency F and valve opening V respectively.

[0010] S3: Select a dynamic threshold, calculate the residual ε under different working conditions based on the working condition model obtained in S2 and the historical statistical data obtained in S1, and obtain the dynamic threshold data based on the residual ε under the corresponding working conditions;

[0011] S4: Establish fault feature engineering, calculate dynamic thresholds for different seasons, and identify features that indicate faults, including: determining an abnormality when the absolute value of the residual exceeds the threshold, and determining an abnormality when the residual variance exceeds 3 times the baseline value during normal operation;

[0012] S5: Establish a fault detection model and determine that the system is abnormal when the residual exceeds the threshold.

[0013] According to a preferred embodiment, the basic operating condition model expression established in step S2 includes:

[0014]

[0015]

[0016] Among them, β0 and α0 are intercept terms, which are the reference values ​​of power P / current I when the independent variables frequency F and opening V are zero; β1, β2, α1, α2 are regression coefficients, and ε is the residual, which is the difference between the actual observation value and the model prediction value.

[0017] According to a preferred embodiment, in step S2, the regression coefficients β1, β2, α1, α2, and the intercept terms β0, α0 are calculated using the collected n groups of observation data.

[0018] According to a preferred embodiment, the dynamic threshold obtained in step S3 is:

[0019]

[0020] in, is the residual obtained based on the historical data statistics in summer The corresponding standard deviation, is the residual obtained based on the historical data statistics of the transition season The corresponding standard deviation, is the residual obtained based on winter historical data statistics The corresponding standard deviation, is the humidity influence coefficient, The humidity deviation is that the monthly average temperature is greater than 25℃ in summer, the monthly average temperature is less than 10℃ in winter, 10℃≤monthly average temperature≤25℃ in transition season, and the average daily relative humidity in summer is greater than 60% in rainy season.

[0021] According to a preferred embodiment, the humidity deviation Indicates the difference between the current humidity and the reference humidity, where

[0022]

[0023] is the current real-time measured relative humidity, is the baseline humidity, which represents the historical average humidity of the current season.

[0024] According to a preferred embodiment, in step S5, the alarm strategy includes: level one alarm: a single point exceeds the threshold range; level two alarm: three consecutive points exceed the threshold range.

[0025] According to a preferred embodiment, the data center CDU system fault identification method further includes: S6: model updating and optimization, establishing a comparison curve between predicted and actual power / current values, and regularly retraining the model with new data.

[0026] The aforementioned main solution of this application and its further options can be freely combined to form multiple solutions, all of which can be adopted and protected by this application. After understanding the solution of this application, those skilled in the art will understand that there are many combinations based on existing technology and common knowledge, all of which are technical solutions to be protected by this application, and these are not exhaustive here.

[0027] Beneficial effects of this application:

[0028] The data center CDU system fault identification method of this application adopts multi-parameter collaborative analysis and simultaneously monitors parameters such as motor frequency, valve opening, power, current, temperature, etc. to improve the accuracy of fault identification. Based on dynamic thresholds and trend analysis (such as the slope change of power residual), it can issue an early warning several hours to several days before the fault occurs completely; the alarm threshold is automatically adjusted according to load and seasonal changes to avoid false alarms caused by fixed thresholds, improve the fault detection rate, greatly reduce maintenance costs, and reduce unnecessary downtime. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 This is a schematic diagram of the application process;

[0030] Figure 2 This is a schematic diagram of the data center CDU system structure. DETAILED DESCRIPTION

[0031] The following describes the embodiments of the present application through specific examples. Those skilled in the art can easily understand the other advantages and effects of the present application from the content disclosed in this specification. The present application can also be implemented or applied through other different specific embodiments. The details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that the following embodiments and features in the embodiments can be combined with each other unless they conflict.

[0032] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings.

[0033] In the description of this application, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," "outer," etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings, or are the orientations or positional relationships in which the product of this application is typically placed when in use. These terms are intended only to facilitate the description of this application and simplify the description, and are not intended to indicate or imply that the device or element referred to must have a specific orientation, be constructed, or operate in a specific orientation. Therefore, they should not be construed as limitations on this application. Furthermore, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0034] Furthermore, terms such as "horizontal," "vertical," and "overhanging" do not necessarily imply that a component must be absolutely horizontal or overhanging, but rather that it can be slightly tilted. For example, "horizontal" simply means that its direction is more horizontal than "vertical," and does not mean that the structure must be completely horizontal, but rather that it can be slightly tilted.

[0035] It should also be noted that, in the description of this application, unless otherwise expressly specified or limited, the terms "disposed," "installed," "connected," and "connected" should be understood in a broad sense. For example, they can refer to fixed connections, detachable connections, or integral connections; they can refer to mechanical connections or electrical connections; they can refer to direct connections or indirect connections through an intermediate medium; and they can refer to internal connections between two components. Those skilled in the art will understand the specific meanings of the above terms in this application based on the specific circumstances.

[0036] In addition, the present application would like to point out that, in the present application, unless the specific structures, connection relationships, positional relationships, power source relationships, etc. are specifically written out, the structures, connection relationships, positional relationships, power source relationships, etc. involved in the present application are all known to those skilled in the art based on the existing technology without creative work.

[0037] Example 1

[0038] The present application discloses a method for identifying faults in a data center CDU system, wherein the data center CDU system includes: a plurality of motors, a plurality of three-way valve actuators, a plurality of temperature sensors, and a plurality of temperature and humidity sensors.

[0039] Specifically, if Figure 2As shown in the figure, the right side is the refrigeration circuit, which is used to complete load cooling; the left side is the cold source circuit, which is used to provide cold source. The data center CDU system includes: at least one motor for circulating the cooling water in the system refrigeration circuit, P01 in the figure; at least one three-way valve actuator for switching the cooling water circulation between internal and external, V601 in the figure; at least two temperature sensors installed at the cooling water inlet and outlet, for the system outlet temperature and inlet temperature, TT01 and TT05 in the figure. At least one temperature and humidity sensor installed around the system to detect the ambient temperature and humidity, TRT01 in the figure, also includes four valves: V009, V010, V011, and V012.

[0040] like Figure 1 As shown, the data center CDU system fault identification method of the present application includes the following steps.

[0041] Step S1, data collection and feature extraction: collect historical data of motor frequency, valve opening, input power, current value, cooling water inlet and outlet temperatures, and ambient temperature and humidity under normal operating conditions.

[0042] Extract key features: mean μ and variance σ² of power and current;

[0043]

[0044]

[0045] The mean μ represents the average value in the data set and reflects the central tendency of the data. n The variance σ² represents the average of the squared differences between the data points and the mean, reflecting the degree of data dispersion. This is used to monitor system stability: a sudden increase in variance may indicate a fault.

[0046] Step S2, establishing a basic operating condition model: When the motor frequency and valve opening are stable, a baseline relationship model of power and current is established, and a multiple linear regression model is used to establish a quantitative relationship between power P / current I and motor frequency F and valve opening V.

[0047] Model expression:

[0048]

[0049]

[0050] Among them, β0 and α0 are intercept terms, which are the reference values ​​of power P / current I when the independent variables frequency F and opening V are zero; β1, β2, α1, α2 are regression coefficients, and ε is the residual, which is the difference between the actual observation value and the model prediction value.

[0051] And through the collected n groups of observation data, the regression coefficients β1, β2, α1, α2, and the intercept terms β0, α0 are calculated.

[0052] Step S3: Obtain the basic operating condition model from step S2. Based on historical data statistics, calculate the residual ε under different operating conditions (different load rates). After calculating a single residual ε value, it is usually necessary to analyze its overall distribution:

[0053] 1. The mean of the residual ε should theoretically be 0 (if the model is unbiased), but in practice it can be close to 0.

[0054] 2. The standard deviation of the residual ε is used to measure the fluctuation range of the error and to set the fault detection threshold.

[0055] 3. Select a dynamic threshold, which is established based on the statistical characteristics of the power / current model residual.

[0056] 4. The fault model establishes a unit to determine the season and apply different thresholds to automatically compensate for the impact of seasonal temperature and humidity changes on cooling efficiency. Seasonal determination data can come from: ① ambient temperature and humidity sensors; ② historical meteorological data; and ③ timestamps.

[0057] Specifically: the monthly average temperature > 25℃ is summer, the monthly average temperature < 10℃ is winter, the monthly average temperature 10℃ ≤ ≤ 25℃ is the transition season, and the relative humidity range in summer > 60% is the rainy season; historical data is divided into seasonal labels: summer / winter / rainy season / transition season.

[0058] Calculate the seasonal thresholds separately. In the rainy season, σ needs to be calculated based on the humidity deviation. Dynamic adjustment (α is the humidity influence coefficient).

[0059] Specifically, the dynamic threshold obtained in step S3 is:

[0060]

[0061] in, is the residual obtained based on the historical data statistics in summer The corresponding standard deviation, is the residual obtained based on the historical data statistics of the transition season The corresponding standard deviation, is the residual obtained based on winter historical data statistics The corresponding standard deviation.

[0062] Indicates the difference between the current humidity and the baseline humidity, used to quantify the additional impact of ambient humidity on the cooling system. The calculation formula is:

[0063]

[0064] : Current real-time measured relative humidity (%), : Baseline humidity, the historical average humidity of the current season.

[0065] α is calibrated by linear regression:

[0066] ①Data preparation:

[0067] Collecting historical data: actual residuals at different humidity levels (Actual power - model predicted power).

[0068] Ensure that the data covers the typical humidity range (e.g. 40% to 85%).

[0069] Build a regression model:

[0070] ∣ ∣=α +β

[0071] Dependent variable: Absolute value of residual | ∣

[0072] Independent variables:

[0073] Regression coefficient: β

[0074] The slope α is the humidity influence coefficient.

[0075] During the rainy season, > 0 → Relax the threshold to reduce false positives.

[0076] During the dry season, <0 → Tighten the threshold and increase sensitivity.

[0077] S4. Establish fault feature engineering. Based on the temperature and humidity information obtained by measurement in different seasons, calculate the dynamic threshold value (also known as the current threshold value) of different seasons. For example, if the real-time temperature measurement statistics show that the monthly temperature is less than 10 degrees Celsius, the current threshold value is .

[0078] The system also identifies features that may indicate a fault. The fault judgment rules are as follows: ① The absolute value of the residual exceeds the threshold; ② The residual variance increases significantly, meaning that an abnormality is determined when the variance exceeds three times the baseline value during normal operation. ③ If the standard deviation of the residual ε suddenly increases but the data does not exceed the threshold, this may indicate a potential fault.

[0079] S5. Establish a fault detection model

[0080] When the residual exceeds the threshold, it is judged as abnormal. The alarm strategies include: Level 1 alarm: a single point exceeds the threshold range; Level 2 alarm: three consecutive points exceed the threshold range.

[0081] S6. Model update and optimization.

[0082] Visual monitoring: Create a curve comparing predicted power / current values ​​to actual values. Regularly (e.g., monthly) retrain the model with new data. Add fault logging: Record all abnormal events and the system status at the time.

[0083] Through the multi-parameter collaborative analysis of the method of this application, parameters such as motor frequency, valve opening, power, current, temperature, etc. are monitored simultaneously to improve the accuracy of fault identification. Based on dynamic thresholds and trend analysis (such as the slope change of power residual), an early warning can be issued several hours to several days before the fault occurs completely; the alarm threshold is automatically adjusted according to load and seasonal changes to avoid false alarms caused by fixed thresholds, improve the fault detection rate, greatly reduce maintenance costs, and reduce unnecessary downtime.

[0084] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements and improvements made within the spirit and principles of the present application should be included in the scope of protection of the present application.

Claims

1. A data center CDU system fault identification method, characterized in that: The data center CDU system fault identification method includes: S1: Data acquisition and feature extraction: Collect historical data on motor frequency, valve opening, input power, current, cooling water inlet and outlet temperatures, and ambient temperature and humidity during normal operation of the data center CDU system, and extract key features: the mean and variance of power and current. S2: Establish a basic operating condition model. When the motor frequency and valve opening are stable, establish a benchmark relationship model between power and current. Use a multiple linear regression model to establish the quantitative relationship between power P / current I and motor frequency F and valve opening V respectively. The basic operating condition model expression established in step S2 includes: Where β0 and α0 are intercept terms, which are the reference values ​​of power P / current I when the independent variables frequency F and opening V are zero; β1, β2, α1, α2 are regression coefficients, and ε is the residual, which is the difference between the actual observation value and the model prediction value; In step S2, the regression coefficients β1, β2, α1, α2, and intercept terms β0, α0 are calculated using the collected n sets of observation data; S3: Select a dynamic threshold, calculate the residual ε under different working conditions based on the working condition model obtained in S2 and the historical statistical data obtained in S1, and obtain the dynamic threshold data based on the residual ε under the corresponding working conditions; The dynamic threshold obtained in step S3 is: in, is the residual obtained based on the historical data statistics in summer The corresponding standard deviation, is the residual obtained based on the historical data statistics of the transition season The corresponding standard deviation, is the residual obtained based on winter historical data statistics The corresponding standard deviation, is the humidity influence coefficient, The humidity deviation is defined as summer when the monthly average temperature is greater than 25°C, winter when the monthly average temperature is less than 10°C, transition season when the monthly average temperature is 10°C ≤ ≤ 25°C, and rainy season when the daily average relative humidity in summer is greater than 60%. S4: Establish fault feature engineering, calculate dynamic thresholds for different seasons, and identify features that indicate faults, including: determining an abnormality when the absolute value of the residual exceeds the threshold, and determining an abnormality when the residual variance exceeds 3 times the baseline value during normal operation; S5: Establish a fault detection model and determine that the system is abnormal when the residual exceeds the threshold.

2. The data center CDU system fault identification method according to claim 1, characterized in that: Humidity deviation Indicates the difference between the current humidity and the reference humidity, where is the current real-time measured relative humidity, is the baseline humidity, which represents the historical average humidity of the current season.

3. The data center CDU system fault identification method according to claim 1, characterized in that: In step S5, the alarm strategy includes: level one alarm: a single point exceeds the threshold range; level two alarm: three consecutive points exceed the threshold range.

4. The data center CDU system fault identification method according to claim 1, characterized in that: The data center CDU system fault identification method further includes: S6: model updating and optimization, Create a comparison curve between the predicted and actual power / current values, and regularly retrain the model with new data.

Citation Information

Patent Citations

  • Wind turbine generator fault early warning and abnormal parameter inspection method and system

    CN116644343A

  • Machine-learning based optimization of data center designs and risks

    US20200348993A1