A machine learning-based data center equipment failure prediction method and system

By collecting multi-phase flow data and training a generative network, combined with a discriminator network, the problem of obtaining the scale thickness in the liquid cooling system is solved, enabling multi-dimensional perception and accurate fault prediction of the liquid cooling system, and adapting to precise operation and maintenance under different working conditions.

CN121029544BActive Publication Date: 2026-02-06BEIJING AVIC XINBERUN TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511564726.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-30
Publication Date
2026-02-06
Estimated Expiration
2045-10-30

AI Technical Summary

Technical Problem

In existing technologies, liquid cooling system monitoring schemes based on flow rate and temperature cannot directly obtain the thickness of the scale layer, resulting in a high rate of false alarms and missed alarms. Furthermore, the fixed warning thresholds cannot be adapted to different operating conditions, and the severity of the fault cannot be quantified or the remaining operating time cannot be predicted.

Method used

By collecting multi-phase flow data, the system health status indicators and the thickness of the scale layer on the inner wall of the pipeline are calculated. Adversarial training is carried out using generative networks and discriminator networks to generate reconstructed data. In the prediction stage, the reconstruction error and anomaly score are calculated, and the threshold is dynamically adjusted to determine the fault evolution and predict the remaining running time.

Benefits of technology

It enables multi-dimensional perception of liquid cooling systems, accurately detects scaling conditions, reduces the risk of false alarms and missed alarms, provides quantitative fault prediction and operation and maintenance support, and adapts to precise operation and maintenance decisions under different operating conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121029544B_ABST
    Figure CN121029544B_ABST
Patent Text Reader

Abstract

The application provides a machine learning-based data center equipment failure prediction method and system, relating to the technical field of data center intelligent operation and maintenance, and the application calculates system overall health state indexes according to data correlation by collecting multi-phase flow data of a cooling system; obtains a scale layer thickness estimation value according to feedback signal attenuation changes by generating micro-vibration through a piezoelectric actuator; takes the multi-phase flow data, the health state indexes and the thickness estimation value as joint inputs, introduces a discriminator network in a generative network for adversarial training, and generates reconstructed data that is difficult to distinguish from normal state; calculates errors of real and reconstructed data, obtains discriminator abnormal scores, and when the two exceed historical data thresholds, determines failure evolution and predicts residual effective time, so that precise prediction of data center cooling system scale failure and residual effective operation time estimation can be realized, and an effective technical solution is provided for early warning of equipment failure.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent operation and maintenance technology for data centers, and in particular to a method and system for predicting data center equipment failures based on machine learning. Background Technology

[0002] In liquid cooling scenarios with high water hardness or long-term operation, impurities in the cooling medium easily form scale on the inner walls of pipes, leading to a gradual decrease in heat exchange efficiency and causing overheating of high-power chips and equipment shutdown. With the development of AI computing servers, their power consumption has increased significantly, and liquid cooling systems have more stringent requirements for heat dissipation stability. There is an urgent need for technology to detect the gradual failure signals caused by scaling in advance, accurately predict the failure trend and remaining operating time, and avoid unplanned downtime losses.

[0003] Currently, a typical solution for this requirement is a flow-temperature correlation monitoring scheme based on ultrasonic flow meters. This scheme uses an external clamp-on ultrasonic flow meter to collect real-time coolant flow rate, and a temperature sensor to obtain the inlet and outlet temperature difference, establishing a baseline correlation model between the two. When flow fluctuations or temperature differences deviate from the baseline to a certain extent, a scaling warning is triggered, achieving preliminary monitoring of abnormal heat exchange efficiency.

[0004] However, this solution has obvious limitations: it relies solely on two parameters, flow rate and temperature, and does not directly obtain core data such as scale layer thickness, making it impossible to distinguish the causes of flow fluctuations and prone to false alarms and missed alarms; the warning thresholds are mostly fixed and do not take into account variables such as water hardness and running time, making it difficult to adapt to different operating conditions and fault patterns; it can only provide simple warnings, cannot quantify the severity of faults, and cannot predict the remaining running time, making it difficult to support precise operation and maintenance. Summary of the Invention

[0005] The purpose of this application is to provide a data center equipment fault prediction method and system based on machine learning, so as to solve the problems of fixed thresholds being difficult to adapt to multiple operating conditions and the inability to quantify the degree of fault in the prior art.

[0006] To address the aforementioned technical problems, in a first aspect, this application provides a data center equipment fault prediction method based on machine learning, comprising:

[0007] The monitoring data collected from the data center cooling system constitutes multi-phase flow data, and the overall health status index of the system is calculated based on the correlation between the multi-phase flow data.

[0008] Micro-vibrations are generated by a piezoelectric actuator, and the thickness of the scale layer on the inner wall of the pipe is estimated based on the attenuation change of the feedback signal of the micro-vibrations.

[0009] The multi-phase flow data, the system health state indicator, and the thickness estimation value are taken as joint inputs, and a discriminator network is introduced in a generative network structure for adversarial training, a probability distribution of the joint inputs under normal working conditions is learned, and reconstructed data that is difficult to distinguish from normal states is generated;

[0010] In the prediction phase, a reconstruction error between real observation data and the reconstructed data is calculated, and an anomaly score output by the discriminator network is obtained, and when the values of the reconstruction error and the anomaly score exceed adaptive thresholds dynamically adjusted based on historical running data, it is determined that the liquid cooling system is evolving towards a fouling failure state, and the remaining effective running time is predicted.

[0011] Optionally, the multi-phase flow data, the system health state indicator, and the thickness estimation value are taken as joint inputs, and a discriminator network is introduced in a generative network structure for adversarial training, a probability distribution of the joint inputs under normal working conditions is learned, and reconstructed data that is difficult to distinguish from normal states is generated, including:

[0012] The multi-phase flow data, the health state indicator, and the thickness estimation value are combined into a multi-dimensional joint input sample;

[0013] A generative network and a discriminator network are respectively constructed, the generative network is used to receive the joint input sample, and a reconstructed sample is generated by learning a probability distribution of the joint input sample under normal working conditions;

[0014] The discriminator network is used to distinguish whether the input sample is a real normal working condition sample or the reconstructed sample;

[0015] A large number of normal working condition joint input samples are used to alternately train the generative network and the discriminator network until the discriminator network cannot distinguish between real normal samples and the reconstructed samples;

[0016] After training, the reconstructed sample generated by the generative network on the normal working condition sample is taken as the reconstructed data corresponding to the normal state probability distribution.

[0017] Optionally, a large number of normal working condition joint input samples are used to alternately train the generative network and the discriminator network until the discriminator network cannot distinguish between real normal samples and the reconstructed samples, including:

[0018] A large number of normal working condition joint input samples are used to alternately update internal parameters of the discriminator network and the generative network;

[0019] After each alternating update, calculate the average difference of the output values of the discriminator network for a batch of real samples and corresponding reconstructed samples, and stop training when the average difference is lower than a preset threshold.

[0020] Optionally, the thickness estimation of the fouling layer on the inner wall of the pipeline is calculated according to the attenuation change of the feedback signal of the micro-vibration generated by the piezoelectric actuator, comprising:

[0021] The micro-vibration of a specific frequency is generated by the piezoelectric actuator, and the attenuation characteristics of the feedback signal of the micro-vibration are extracted;

[0022] The attenuation characteristics are compared with the pre-calibrated reference attenuation characteristics corresponding to different fouling thicknesses to estimate the thickness estimation.

[0023] Optionally, the attenuation characteristics are compared with the pre-calibrated reference attenuation characteristics corresponding to different fouling thicknesses to estimate the thickness estimation, comprising:

[0024] A set of corresponding relationships established in advance through experiments is obtained, each entry in the set of corresponding relationships contains a known fouling layer thickness value and a corresponding reference attenuation characteristic value;

[0025] The current extracted attenuation characteristics are compared one by one with each reference attenuation characteristic value in the set of corresponding relationships to determine the two reference attenuation characteristic values with the highest similarity to the current attenuation characteristics;

[0026] According to the similarity of the current attenuation characteristics and the two reference attenuation characteristic values, interpolation calculation is performed between the fouling layer thickness values corresponding to the two reference attenuation characteristic values respectively, and the value obtained by interpolation calculation is taken as the thickness estimation of the current fouling layer on the inner wall of the pipeline.

[0027] Optionally, in the prediction stage, the reconstruction error between the real observation data and the reconstructed data is calculated, and the anomaly score output by the discriminator network is obtained, and when the values of the reconstruction error and the anomaly score exceed the adaptive threshold dynamically adjusted based on historical operation data, it is determined that the liquid cooling system is evolving towards a fouling failure state, and the remaining effective operation time is predicted, comprising:

[0028] In the prediction stage, the newly obtained real observation data is input into the trained generative network, and the corresponding reconstructed data is output by the generative network;

[0029] The differences in each corresponding dimension between the real observation data and the reconstructed data are calculated, and the differences in each corresponding dimension are integrated into a reconstruction error;

[0030] The real observation data is input into the trained discriminator network, and a corresponding anomaly score is obtained through the discriminator network, the anomaly score representing a likelihood value of being judged as normal data.

[0031] Optionally, the reconstruction error and the anomaly score of the current data are compared with system historical adaptive thresholds respectively, and when at least one of the reconstruction error and the anomaly score continuously exceeds the corresponding adaptive threshold, it is determined that the liquid cooling system is evolving towards a fouling failure state.

[0032] According to the degree and duration that the reconstruction error and the anomaly score exceed the threshold, the state deterioration trend of the system is inferred, and the remaining effective operation time is predicted according to the state deterioration trend.

[0033] In a second aspect, the present application provides a machine learning-based data center equipment failure prediction system, comprising:

[0034] The acquisition module is configured to acquire multi-phase flow data of the data center cooling system, and calculate a health state indicator of the system as a whole based on an association relationship between the multi-phase flow data.

[0035] The calculation module is configured to generate micro-vibration through a piezoelectric actuator, and calculate an estimated value of the thickness of the fouling layer on the inner wall of the pipeline according to the attenuation change of a feedback signal of the micro-vibration.

[0036] The reconstruction module is configured to take the multi-phase flow data, the health state indicator of the system, and the estimated value of the thickness as joint inputs, perform adversarial training by introducing a discriminator network in a generative network structure, learn the probability distribution of the joint inputs under normal working conditions, and generate reconstructed data that is difficult to distinguish from the normal state.

[0037] The prediction module is configured to calculate a reconstruction error between real observation data and the reconstructed data in a prediction phase, and obtain an anomaly score output by the discriminator network, and when the values of the reconstruction error and the anomaly score exceed adaptive thresholds dynamically adjusted based on historical operation data, it is determined that the liquid cooling system is evolving towards a fouling failure state, and the remaining effective operation time is predicted.

[0038] In a third aspect, the present application provides an electronic device, comprising:

[0039] The memory is configured to store a computer program.

[0040] The processor is configured to execute the computer program to implement the steps of the machine learning-based data center equipment failure prediction method of the first aspect.

[0041] Fourthly, this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, can implement the steps of the machine learning-based data center equipment fault prediction method described in the first aspect above.

[0042] The data center equipment fault prediction method provided in this application, based on machine learning, collects multi-phase flow data from the data center cooling system and calculates the overall health status index of the system based on the correlation between the multi-phase flow data. This enables multi-dimensional perception of the liquid cooling system's operating status, breaking through the information blind spots of single data and providing a comprehensive basis for subsequent fault diagnosis. Furthermore, by generating micro-vibrations through piezoelectric actuators and calculating the estimated thickness of the scale layer on the inner wall of the pipe based on the attenuation changes of the feedback signal of these micro-vibrations, the method utilizes the electromechanical conversion characteristics of piezoelectric materials to accurately capture the vibration propagation changes caused by scale buildup, achieving direct and efficient detection of core scale status parameters. By combining the multi-phase flow data and the system... Using the health status index and the thickness estimate as joint inputs, a discriminator network is introduced into the generative network structure for adversarial training. This learns the probability distribution of the joint inputs under normal operating conditions and generates reconstructed data, effectively learning the characteristic patterns of normal operating modes. Even in scenarios where fault samples are scarce, a reliable benchmark model can be built. By calculating the reconstruction error between the actual observed data and the reconstructed data during the prediction phase, the anomaly score output by the discriminator network is obtained. When both exceed an adaptive threshold dynamically adjusted based on historical data, the fault evolution is determined and the remaining operating time is predicted. This enables accurate fault identification and trend prediction, reduces the risk of false alarms and missed alarms caused by fixed thresholds, and provides quantitative decision support for precise operation and maintenance. Attached Figure Description

[0043] To more clearly illustrate the technical solutions of the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0044] Figure 1 A flowchart illustrating a machine learning-based data center equipment fault prediction method provided in this application embodiment;

[0045] Figure 2 A flowchart illustrating a specific implementation of a machine learning-based data center equipment fault prediction method provided in this application embodiment;

[0046] Figure 3A schematic diagram illustrating a specific implementation of a machine learning-based data center equipment fault prediction method provided in this application embodiment;

[0047] Figure 4 This is a schematic diagram of the structure of a data center equipment fault prediction system based on machine learning, provided in an embodiment of this application. Detailed Implementation

[0048] In the monitoring of scaling in data center liquid cooling systems, existing solutions based on ultrasonic flow meters have significant shortcomings: relying solely on flow rate and temperature data for judgment does not directly obtain crucial information such as the thickness of scaling inside the pipe. For example, a decrease in flow rate could be caused by scaling or a problem with the pump, easily leading to misjudgment or missed detection. Moreover, the judgment criteria for warnings are fixed, failing to consider the differences in water hardness and equipment operating time in different scenarios. For instance, scaling occurs rapidly in areas with hard water, yet the same standard is used for areas with soft water, resulting in poor adaptability. More importantly, it can only indicate a possible problem, but cannot specify the severity of scaling or predict how long the equipment can continue to operate normally, causing difficulties in subsequent maintenance.

[0049] To address these issues, this application proposes a machine learning-based method for predicting data center equipment failures. This method first collects multi-dimensional operational data of the liquid cooling system and calculates the system's health status. Then, a specialized device is used to detect the thickness of scale buildup in the pipes. Next, an intelligent model learns normal operating patterns. Finally, dynamically adjusted judgment criteria are used to monitor for failures. Specifically, multi-dimensional data and scale thickness analysis can accurately identify the cause of problems, avoiding false alarms and missed alarms; dynamic criteria can adapt to different water qualities and operating durations; and it can also determine the severity of scale buildup and estimate remaining operating time. This effectively solves the shortcomings of existing solutions in terms of accuracy, scenario adaptability, and maintenance guidance, making the maintenance of liquid cooling systems more precise and reliable.

[0050] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are merely some embodiments of the present application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0051] The core of this application is to provide a data center equipment fault prediction method based on machine learning, and a flowchart of one specific implementation is shown below. Figure 1 As shown, the method includes:

[0052] S101. Collect monitoring data of the data center cooling system to form multi-phase flow data, and calculate the overall health status index of the system based on the correlation between the multi-phase flow data.

[0053] The monitoring data includes temperature data, water conductivity data, and fluid pressure and flow rate data. Temperature data is acquired by deploying temperature sensors such as thermocouples and platinum resistance sensors near the inlet and outlet of cooling pipes, key heat exchange nodes, and equipment heat dissipation areas to monitor real-time temperature changes of the coolant at different locations. Water conductivity data is typically collected by online conductivity sensors installed in the cooling circuit, indirectly reflecting the content of ionic impurities in the water by detecting the coolant's conductivity, thus helping to determine the potential sources of scaling substances. Fluid pressure data can be measured using pressure sensors installed on the pipes, such as strain gauge pressure sensors, to sense the pressure state of the coolant within the pipes in real time. Flow rate data can be acquired using flow meters such as electromagnetic flow meters and ultrasonic flow meters to monitor the flow rate and volume of the coolant within the pipes in real time. Next, the temperature data, water conductivity data, and fluid pressure and flow rate data are time-series aligned to form multi-phase flow data.

[0054] Specifically, the process begins by collecting monitoring data from the cooling system, including temperature data, water conductivity data, and fluid pressure and flow rate data. Next, the temperature, conductivity, pressure, and flow rate data are time-series aligned to form multi-phase flow data. Then, based on this multi-phase flow data under normal operating conditions, a baseline model is established. By calculating the overall deviation of the multi-phase flow data from the baseline model across all dimensions, a numerical health status index is generated. Finally, through statistical analysis or physical principles, normal correlation patterns between the data points are uncovered, such as the implicit correlation between flow rate and temperature difference, and conductivity and pressure / flow rate, establishing a correlation baseline range under normal operating conditions. This results in a complete data stream containing basic data, correlation patterns, and the baseline range.

[0055] It should be noted that when constructing the correlation benchmark, it is necessary to base it on the historical data of stable operation of the system without scaling failure, and the benchmark parameters need to be adjusted according to different liquid cooling scenarios (such as water hardness and equipment running time differences) to ensure benchmark adaptability.

[0056] The system health status index is calculated based on the data stream. This involves comparing the deviations of multi-phase stream data with the associated benchmark in real time, weighting and fusing each deviation according to its impact weight, and quantifying the health value that reflects the overall status of the system.

[0057] S102. Micro-vibration is generated by a piezoelectric actuator, and the thickness of the scale layer on the inner wall of the pipe is estimated based on the attenuation change of the feedback signal of the micro-vibration.

[0058] Optionally, step S102 may specifically include the following steps:

[0059] S1021. A micro-vibration of a specific frequency is generated by a piezoelectric actuator, and the attenuation characteristics of the feedback signal of the micro-vibration are extracted.

[0060] S1022. The attenuation characteristics are compared with the pre-calibrated benchmark attenuation characteristics corresponding to different scale thicknesses to estimate the thickness estimate.

[0061] Specifically, step S1022 may include the following process: obtaining a pre-established set of correspondences through experiments, wherein each entry in the set contains a known scale thickness value and a corresponding baseline attenuation characteristic value; comparing the currently extracted attenuation characteristic with each baseline attenuation characteristic value in the set of correspondences one by one, and determining the two baseline attenuation characteristic values ​​with the highest similarity to the current attenuation characteristic; performing interpolation calculation between the scale thickness values ​​corresponding to the two baseline attenuation characteristic values ​​based on the similarity between the current attenuation characteristic and the two baseline attenuation characteristic values, and using the interpolated value as the estimated thickness of the scale layer on the inner wall of the pipe.

[0062] In the above steps, a piezoelectric actuator is a device that converts electrical energy into mechanical vibration, generating micro-vibrations with controllable frequency and amplitude to transmit vibration signals to the pipeline. A specific frequency micro-vibration refers to a pre-set vibration frequency (usually in the kHz range) that can propagate stably within the pipeline and is sensitive to scaling, based on the material (e.g., stainless steel) and diameter of the liquid-cooled pipeline. The attenuation characteristic of the feedback signal refers to the degree to which the signal strength weakens due to the absorption of vibration energy by scaling on the inner wall of the pipeline during micro-vibration propagation; it is commonly expressed by indicators such as vibration amplitude attenuation rate and signal propagation time delay. The pre-calibrated baseline attenuation characteristic is a set of "scale thickness - attenuation characteristic" data established by experimentally testing the attenuation law of vibration signals under different known scale thicknesses. The correspondence set is a database storing multiple "known scale thickness values ​​- baseline attenuation characteristic values" entries. Interpolation calculation is a mathematical calculation method that estimates the current thickness by using the thickness values ​​corresponding to the attenuation characteristics of two adjacent baselines when the current attenuation characteristic does not perfectly match any baseline in the set.

[0063] As one possible approach, firstly, the installation location of the piezoelectric actuator is determined in step S1021. Typically, a straight section of the liquid-cooled pipe is chosen to avoid interference from bends in vibration propagation. A specific vibration frequency is set based on the pipe material and diameter. For example, for a 304 stainless steel pipe with an 80mm diameter, a frequency of 1.5kHz is set. This frequency has been tested and found to effectively penetrate the pipe and is not significantly affected by scaling. The control module drives the piezoelectric actuator to generate stable micro-vibrations, for example, the actuator outputs a vibration amplitude of 6V, continuously transmitting vibration signals to the pipe. Secondly, a piezoelectric sensor is installed on the other side of the pipe, 2 meters away from the actuator, ensuring sufficient propagation distance for the vibration signal to exhibit attenuation. The feedback vibration signal is collected in real time, and the collected analog signal is converted into a digital signal via a data acquisition card and transmitted to a signal processing program in the computer. The program calculates the attenuation characteristics by comparing the amplitude difference and phase change between the initial vibration signal and the feedback signal. For example, with an initial signal amplitude of 6V and a feedback signal amplitude of 5.2V, the calculated attenuation rate is approximately 13.3%, thus completing the extraction of attenuation characteristics.

[0064] Secondly, a pre-calibrated set of correspondences is obtained through step 1022. This set is constructed through previous experiments: First, pipe samples with the same material and diameter as the actual tested pipes are selected. Standard scale layers with thicknesses of 0.2mm, 0.4mm, 0.6mm, 0.8mm, and 1.0mm are artificially created on the inner walls of the samples. Calcium carbonate can be used to simulate the actual scale composition. Then, a 1.5kHz micro-vibration is applied to each sample using the same method, and the attenuation rate corresponding to each thickness is recorded, forming a set of correspondences such as "0.2mm-5%", "0.4mm-9%", "0.6mm-14%", "0.8mm-18%", and "1.0mm-23%", which is stored in the system database. Next, the extracted current attenuation characteristic, such as the actual detected attenuation rate of 11%, is compared one by one with the benchmark attenuation characteristic values ​​in the set. It is found that 11% is between 9% (corresponding to 0.4mm) and 14% (corresponding to 0.6mm), and has the highest similarity to both. Finally, linear interpolation is used to calculate the current thickness. According to the interpolation formula "current thickness = smaller thickness value + (current attenuation rate - smaller attenuation rate) / (larger attenuation rate - smaller attenuation rate) × (larger thickness value - smaller thickness value)", substituting the data, we get "current thickness = 0.4mm + (11% - 9%) / (14% - 9%) × (0.6mm - 0.4mm) = 0.48mm". 0.48mm is used as the estimated thickness of the scale layer on the inner wall of the pipe, thus completing the estimation.

[0065] Specifically, for example, in a scenario involving scale detection of liquid-cooled pipes in a type A data center, for a stainless steel liquid-cooled pipe of specification B with a diameter adapted to the heat dissipation requirements of a type A data center, a C-brand piezoelectric actuator installed on the outer wall of the pipe first generates a micro-vibration at a specific frequency of 1.8kHz according to preset parameters. Simultaneously, a signal acquisition device located at the same position receives the vibration signal fed back from the other side of the pipe, and the amplitude attenuation of the feedback signal is extracted using dedicated data processing software. Next, a pre-established set of correspondences, based on the same material and phase of the pipe as specification B, is retrieved from the system database. For sample pipes of the same diameter, scale layers of different thicknesses were artificially created for testing. Each entry included a known scale thickness value and a corresponding baseline attenuation characteristic value. The extracted current attenuation characteristic was compared one by one with each baseline attenuation characteristic value in the set. It was found that the current attenuation characteristic had the highest similarity to the baseline attenuation characteristic corresponding to "0.3mm scale thickness" and "0.5mm scale thickness". Subsequently, a linear interpolation algorithm was used to calculate the estimated thickness of the scale layer on the inner wall of the pipe between 0.3mm and 0.5mm based on the closeness of the current attenuation characteristic to these two baseline attenuation characteristics.

[0066] In the overall scheme of step S102 above, by controlling the specificity of the vibration frequency, the vibration signal is ensured to propagate stably in the pipeline and can effectively reflect the scaling state, thereby improving the pertinence and effectiveness of signal extraction. By using a pre-calibrated set of correspondences, combined with one-to-one comparison and interpolation calculation, the accuracy of the previous experimental data is utilized, and the gaps between the reference data are filled by interpolation, making the thickness estimate more consistent with the actual scaling situation and effectively avoiding estimation errors caused by the dispersion of the reference data. Overall, this scheme can achieve accurate and efficient estimation of the thickness of the scaling layer on the inner wall of the liquid-cooled pipeline, providing key core data support for subsequent judgment on whether the liquid-cooling system is evolving into a scaling failure state, helping to detect scaling problems in advance, and ensuring the heat dissipation efficiency and stable operation of the liquid-cooling system.

[0067] S103. The multi-phase flow data, system health status indicators and thickness estimate are used as joint inputs. Adversarial training is performed by introducing a discriminator network into the generative network structure to learn the probability distribution of the joint inputs under normal operating conditions and generate reconstructed data that is difficult to distinguish from the normal state.

[0068] Optionally, such as Figure 2 As shown, step S103 may specifically include the following steps:

[0069] S1031. Combine the multi-phase flow data, the health status index, and the thickness estimate into a multi-dimensional joint input sample;

[0070] S1032. Construct a generative network and a discriminator network respectively. The generative network is used to receive the joint input samples and generate reconstructed samples by learning the probability distribution of the joint input samples under normal working conditions.

[0071] S1033, The discriminator network is used to distinguish whether the input sample is a real normal working condition sample or the reconstructed sample;

[0072] S1034. Using a large number of joint input samples from normal operating conditions, the generative network and the discriminator network are subjected to adversarial training in an alternating manner until the discriminator network can no longer distinguish between real normal samples and the reconstructed samples.

[0073] Specifically, step S1034 may include the following process: using a large number of joint input samples under normal operating conditions, alternately update the internal parameters of the discriminator network and the generative network; after each alternate update, calculate the average difference of the output values ​​of the discriminator network to a batch of real samples and corresponding reconstructed samples, and stop training when the average difference is lower than a preset threshold.

[0074] S1035. After training is completed, the reconstructed samples generated by the generative network on normal working condition samples are used as the corresponding reconstructed data that conforms to the normal state probability distribution.

[0075] In the above steps, joint input refers to a multi-dimensional data set formed by integrating multi-phase flow data, system health status indicators, and thickness estimates, which provides comprehensive input for model training; generative network is a deep learning network whose core function is to learn the probability distribution of data under normal operating conditions, and then generate reconstructed samples similar to normal data features; discriminator network, also a deep learning network, is responsible for classifying input samples and determining whether they are samples from real normal operating conditions or reconstructed samples generated by the generative network; adversarial training refers to the process of alternating training and mutual competition between the generative network and the discriminator network. Through this process, the samples generated by the generative network become closer and closer to real normal samples, and the discriminator network's discrimination ability is continuously optimized, eventually reaching a dynamic equilibrium; joint input sample is structured data that can be directly input into the model after the joint input has been standardized in data format; reconstructed sample is a data sample simulating normal operating conditions generated by the generative network after learning the normal data distribution; preset threshold is a pre-set indicator threshold used to determine whether training has stopped. When the difference between the discriminator network's output on real samples and reconstructed samples is lower than this value, it indicates that the model training has achieved the expected results.

[0076] In this embodiment, firstly, in step S1031, the previously collected multi-phase flow data, the calculated system health status indicators, and the estimated thickness values ​​are combined in a fixed data dimension order to avoid the impact of dimension confusion on subsequent model input. To eliminate differences in data volume, each data point is normalized to ultimately form a multi-dimensional joint input sample.

[0077] For example, after normalizing the temperature (25℃) to 0.6, the conductivity (200μS / cm) to 0.3, the pressure (0.3MPa) to 0.5, the flow rate (5L / min) to 0.7, the health index (85) to 0.85, and the thickness (0.4mm) to 0.2, the above data are integrated into a structured joint input sample of [0.6, 0.3, 0.5, 0.7, 0.85, 0.2] to ensure that it can be directly read by the model.

[0078] Secondly, using S1032, a deep neural network architecture is employed to construct a generative network and a discriminator network, respectively, clarifying their functional roles. The generative network adopts an encoder-decoder structure, where the encoder compresses high-dimensional joint input samples into low-dimensional feature vectors through convolutional and pooling layers, and the decoder restores the low-dimensional features to reconstructed samples with the same dimensions as the original samples through deconvolutional layers and activation functions. The core objective is to learn the probability distribution of joint input samples under normal operating conditions. The discriminator network adopts a fully connected network structure, where the input layer receives sample data, the intermediate layers extract sample features through the ReLU activation function, and the output layer is set to binary classification, with 0 representing reconstructed samples and 1 representing real normal samples. The core objective is to accurately distinguish sample types.

[0079] For example, the generative network first compresses the 6-dimensional sample into 3-dimensional features. Then, after receiving the joint input sample [0.6, 0.3, 0.5, 0.7, 0.85, 0.2], it is processed by the encoder-decoder to generate the reconstructed sample [0.58, 0.32, 0.49, 0.71, 0.84, 0.21]. The discriminator network can then receive the reconstructed sample and the real normal sample, and initially learn the feature differences between the two to complete the differentiation.

[0080] Next, through S1033, the joint input samples of real normal operating conditions and the reconstructed samples generated by the generative network are mixed in a 1:1 ratio and uniformly input into the discriminator network. The discriminator network calculates the matching degree between the sample features and the preset real sample features, and outputs the probability value that the sample is a real normal sample. It should be noted that the closer the probability is to 1, the more likely it is to be a real sample; the closer it is to 0, the more likely it is to be a reconstructed sample.

[0081] For example, 500 samples are taken from the joint input samples of real normal working conditions and 500 samples are taken from the reconstructed samples generated by the generative network. When the real sample [0.6,0.3,0.5,0.7,0.85,0.2] is input, the output is 0.95, and when the reconstructed sample [0.58,0.32,0.49,0.71,0.84,0.21] is input, the output is 0.4, thereby realizing the differentiation of sample types.

[0082] Then, adversarial training is conducted using a large number (e.g., 10,000) of joint input samples from normal operating conditions via S1034: First, the parameters of the generative network are fixed, and real samples and reconstructed samples are input into the discriminator network. The discrimination error is calculated using the cross-entropy loss function, such as the loss when a real sample is misclassified as a reconstructed sample and the loss when a reconstructed sample is misclassified as a real sample. The gradient descent algorithm (e.g., Adam optimizer) is used to update the weight parameters of the discriminator network to improve its ability to distinguish between real samples and reconstructed samples. Next, the parameters of the discriminator network are fixed, and the joint input samples are input into the generative network to generate new reconstructed samples. Based on the output error of the discriminator network for the new reconstructed samples, if the output value is too low, it indicates that the sample is easily judged as a reconstructed sample. Similarly, the encoder and decoder parameters of the generative network are updated using the gradient descent algorithm to optimize the authenticity of the generated samples. This alternating training process is then performed. After each alternating update, a batch of samples, such as 100 real samples and corresponding reconstructed samples, is randomly selected. The average difference between the output values ​​of the discriminator network for the two types of samples is calculated. For example, the average output of real samples is 0.9, and the average output of reconstructed samples is 0.6, with an average difference of 0.3. When this average difference is lower than a preset threshold, such as 0.1, it indicates that the discriminator can no longer distinguish between the two types of samples, the generative network has fully learned the normal data distribution, and training stops.

[0083] Finally, through S1035, after training is completed, the reconstructed samples generated by the generative network on all normal operating condition samples, including the training set and the validation set, are summarized. Since these reconstructed samples have been aligned with the data probability distribution under normal operating conditions, they can be used as reconstructed data that conforms to the probability distribution of normal states. This data can be used for comparative analysis with real observation data in the subsequent fault prediction stage, providing a normal data benchmark for fault judgment.

[0084] In practical applications, in the fault prediction project of a type A data center liquid cooling system, multi-phase flow data of the data center liquid cooling system, including temperature, water conductivity, fluid pressure, and flow rate, are first collected. The temperature range is 24-26℃, water conductivity range is 180-220 μS / cm, fluid pressure range is 0.28-0.32 MPa, and flow rate range is 4.8-5.2 L / min. This data is then combined with 82-88 calculated system health status indicators and estimated pipe scale thickness of 0.3-0.5 mm, and after normalization, a 6-dimensional joint input sample is generated. Subsequently, a generative network and a discriminator network are constructed using the TensorFlow framework. The generative network's encoder contains two convolutional layers, and the decoder contains two deconvolutional layers. This generative network receives the joint input samples and learns their distribution. The discriminator network contains three fully connected layers to distinguish sample types. Next, the joint input samples of real normal operating conditions, such as [0.55, 0.28, 0.48, 0.68, 0.83, 0.18], are input together with the reconstructed samples generated by the generative network, such as [0.56, 0.29, 0.47, 0.69, 0.82, 0.19], into the discriminator network. The discriminator outputs a probability of 0.93 for the real sample and a probability of 0.52 for the reconstructed sample, thus distinguishing the sample types. Then, the two networks are trained alternately using 10,000 normal operating condition samples from the data center. Each training iteration first updates the discriminator parameters to improve discrimination, then updates the generative network parameters to optimize sample quality. After each training round, the average difference between the outputs of the validation set samples is calculated, and training stops when the difference is below 0.1. Finally, all the reconstructed samples generated by the generative network on the normal samples are used as reconstructed data conforming to the normal state for subsequent comparison with the real observation data.

[0085] In the overall scheme of step S103 above, this application breaks through the limitations of traditional single-data modeling. By fusing multi-dimensional data to form a joint input, it achieves a comprehensive characterization of the liquid cooling system's state. Simultaneously, it introduces adversarial training between generative and discriminator networks, enabling the model to accurately learn the distribution of normal operating condition data and generate highly realistic reconstructed data. Compared to traditional monitoring schemes, its progress is reflected in: multi-data fusion avoids misjudgment based on a single parameter, providing a more comprehensive basis for fault prediction; adversarial training solves the problem of difficulty in capturing the characteristics of normal operating condition samples, enabling the reconstructed data to accurately match the real normal state; and the final normal state benchmark data provides a reliable comparison standard for subsequent fault identification, fundamentally improving the accuracy and reliability of fault prediction and supporting precise operation and maintenance decisions.

[0086] S104. In the prediction phase, the reconstruction error between the actual observation data and the reconstructed data is calculated, and the anomaly score output by the discriminator network is obtained. When the values ​​of the reconstruction error and the anomaly score exceed the adaptive threshold dynamically adjusted based on historical operating data, it is determined that the liquid cooling system is evolving towards a scaling failure state, and the remaining effective operating time is predicted.

[0087] Optionally, step S104 may specifically include the following steps:

[0088] S1041. In the prediction stage, the newly obtained real observation data is input into the trained generative network, and the corresponding reconstructed data is output through the generative network.

[0089] S1042. Calculate the differences between the real observation data and the reconstructed data in each corresponding dimension, and integrate the differences in each corresponding dimension into the reconstruction error;

[0090] S1043. Input the real observation data into the trained discriminator network, and obtain the corresponding anomaly score through the discriminator network. The anomaly score represents the probability value of being judged as normal data.

[0091] S1044. The reconstruction error and the anomaly score of the current data are compared with the historical adaptive threshold of the system. When at least one of the reconstruction error and the anomaly score continuously exceeds the corresponding adaptive threshold, it is determined that the liquid cooling system is evolving towards a scaling failure state.

[0092] S1045. Based on the degree and duration of the reconstruction error and anomaly score exceeding the threshold, infer the deterioration trend of the system's state, and predict the remaining effective running time based on the deterioration trend.

[0093] In the above steps, the prediction phase refers to the stage after the model has been trained, during which it is used to monitor the real-time operating status of the liquid cooling system and identify faults; the real observation data refers to the multi-phase flow data of the liquid cooling system collected in real time during the prediction phase, including temperature, water conductivity, fluid pressure, flow rate, as well as system health indicators and estimated pipe scale thickness, consistent with the joint input sample format of the training phase; the reconstructed data refers to the data generated by the trained generative network after receiving the real observation data, conforming to the probability distribution of normal operating conditions; the reconstruction error refers to the difference between the real observation data and the reconstructed data, reflecting the degree to which the real data deviates from the normal state; reconstruction Error refers to the comprehensive error value after integrating the differences between the real observation data and the reconstructed data in various dimensions such as temperature and pressure; Anomaly score refers to the output value of the trained discriminator network on the real observation data, which is used to indicate the probability that the data is judged as normal data. The lower the score, the more likely the data is to be abnormal; Adaptive threshold is a critical value that is dynamically adjusted based on the system's historical operating data. It is used to determine whether the reconstruction error and anomaly score exceed the normal range and will adapt to changes in scenarios such as water hardness and operating time; Remaining effective operating time refers to the estimated time from the current moment until the liquid cooling system may shut down due to scaling failure.

[0094] Specifically, firstly, through S1041, in the prediction phase of the data center liquid cooling system, sensors collect real-time data on temperature, water conductivity, fluid pressure, and flow rate at a given moment. Using a pre-established correlation model, system health indicators are calculated. A piezoelectric actuator detects and calculates the estimated thickness of pipe scale. These data are combined in the order of "temperature-conductivity-pressure-flow rate-health indicators-thickness estimate," and after normalization, form real observation data. This data is then input into a pre-trained encoder-decoder generative network. The encoder compresses features from the real observation data, and the decoder restores the compressed features based on normal operating conditions, outputting reconstructed data. For example, if the temperature dimension values ​​in the real observation data deviate from the normal range, the temperature dimension values ​​in the reconstructed data will conform to the normal range.

[0095] Secondly, through S1042, the difference between the real observation data and the reconstructed data is calculated for each dimension. For example, the difference in temperature is calculated using absolute error, and the difference in pressure is calculated using relative error. Then, the weights are set according to the influence of each dimension on heat dissipation efficiency. Among them, the weights of temperature and pressure dimensions are higher. The difference of each dimension is multiplied by the corresponding weight and then summed to obtain the reconstruction error. This error can comprehensively reflect the degree of deviation between the real data and the normal state.

[0096] Next, in step S1043, real observation data is input into the trained discriminator network with a fully connected layer architecture. The input layer of the discriminator network receives the data, the intermediate layers extract features using the ReLU activation function and compare them with the features of normal samples stored during training, and the output layer outputs anomaly scores using the Sigmoid activation function. For example, if real observation data deviates from the normal state due to scaling, the discriminator network will output a low anomaly score, indicating that it is unlikely to classify the data as normal.

[0097] Then, through S1044, the adaptive threshold dynamically updated by the system based on the normal operation data of the past 6 months is retrieved. The current reconstruction error is compared with the error threshold and the abnormal score is compared with the score threshold. If a certain indicator exceeds the threshold and continues for 2 monitoring cycles, each of which is 10 minutes, the system is determined to evolve towards a fault, thus avoiding misjudgment caused by instantaneous fluctuations.

[0098] Finally, through S1045, the magnitude of the reconstruction error exceeding the threshold, the magnitude of the abnormal score below the threshold, and the duration are statistically analyzed. Referring to the time it takes for the system to go from abnormal to failure under the same magnitude and duration in history, and combined with the current system's heat dissipation margin, the trend of state deterioration is inferred, and the remaining effective running time is predicted. For example, in similar historical cases, it takes a certain amount of time for the system to go from a similar state to failure. Combined with the current system's better heat dissipation foundation, the estimated duration is appropriately adjusted.

[0099] In practical applications, for example, in a fault prediction project for a liquid cooling system in a type A data center, after the system enters the prediction phase, it collects key data every 15 minutes via sensors: coolant temperature is 29℃ and the value is high; water conductivity is 260μS / cm and exceeds the stable range; pressure is 0.38MPa and slightly higher than the normal range; and flow rate is 4.1L / min and shows a downward trend. Based on this data, the system calculates a health status index of 68 points, which is below the health threshold of 80 points. Simultaneously, the piezoelectric detection module measures a pipe scale thickness of 0.9mm, which is close to the warning value of 0.1mm. This data is integrated and normalized to form real observation data. After inputting this data into a trained generative network, reconstructed data corresponding to normal operating conditions is obtained. In the reconstructed data, the temperature is 25℃ and the conductivity is 200μS / cm, which are within the parameter range of normal system operation. The system calculates the differences between the real observation data and the reconstructed data in each dimension, and then performs a weighted summation according to preset weights, where temperature and pressure each account for 0.3, and conductivity and flow rate each account for 0.2, finally obtaining a reconstruction error of 0.28. At the same time, the real observation data is input into the trained discriminator network, and the network outputs an anomaly score of 0.35, which indicates that the real observation data is unlikely to be judged as normal data.

[0100] The system then retrieved adaptive thresholds dynamically adjusted based on 12 months of historical operating data, with a reconstruction error threshold of 0.25 and an anomaly score threshold of 0.5. The comparison revealed that the current reconstruction error exceeded the error threshold, while the anomaly score fell below the score threshold. Furthermore, this anomaly had occurred for four consecutive monitoring cycles, totaling one hour, indicating that the liquid cooling system was gradually evolving towards a scaling failure state. The system further retrieved historical fault databases, comparing similar cases from the past three years. These cases exceeded the thresholds by more than 12% in error and 30% in score, lasting 1-2 hours. Combined with the current system's heat dissipation margin, which showed a 18% decrease in heat dissipation efficiency compared to the design value, the system's deterioration rate was estimated at approximately 2% per day, ultimately predicting a remaining effective operating time of 9 days. The operations and maintenance platform synchronized this prediction to the operations and maintenance management system of the Type A data center and automatically generated maintenance recommendations: schedule pipe descaling within 7 days, and prepare suitable descaling equipment and coolant in advance to prevent further reduction in heat exchange efficiency due to scaling, which could lead to overheating or shutdown of high-power chips.

[0101] In the overall scheme of step S104 above, the system status is judged from two dimensions: data difference and normal probability, by using the generative network to output reconstructed data and calculate reconstruction error, combined with the anomaly score output by the discriminator network, thus avoiding the one-sidedness of judgment based on a single indicator. The introduction of adaptive thresholds overcomes the limitation of fixed thresholds being difficult to adapt to different operating conditions, and can dynamically adjust according to the historical operating status of the system, reducing the risk of misjudgment or omission. The design of continuously monitoring the state of indicators exceeding the threshold effectively eliminates the interference of instantaneous fluctuations on fault judgment, thereby improving the accuracy of judgment. The prediction of the remaining effective running time provides maintenance personnel with a clear time reference, facilitating advance planning of maintenance work and avoiding unplanned downtime losses. Overall, this step achieves accurate prediction and maintenance guidance for scaling faults in liquid cooling systems, significantly improving the stability of system operation and the initiative of maintenance.

[0102] The following is a complete embodiment for steps S101 to S104:

[0103] like Figure 3As shown, in a cold plate liquid cooling system of a type A data center, when using this method for fault prediction, sensors deployed at the inlet and outlet of the liquid cooling pipes and chip heat dissipation nodes are used to collect real-time data on the temperature, conductivity, pressure, and flow rate of the coolant. The temperature monitoring range is 22-28℃, the conductivity range is 180-250 μS / cm, the pressure range is 0.25-0.35 MPa, and the flow rate range is 4.5-5.5 L / min. Multi-phase flow data is constructed using this data. Then, based on a preset physical correlation model, the relationships between the data are analyzed, such as the correlation between flow rate and inlet / outlet temperature difference, and conductivity and pressure decay. A system health status index on a 0-100 scale is calculated. Initially, this index stabilizes at 85-88 points, with a score above 80 indicating a healthy state. Subsequently, a B-brand piezoelectric actuator was installed on the outer wall of the straight section of the liquid-cooled pipeline. A frequency of 1.8kHz was set to generate micro-vibration. The feedback signal was received through the matching module and the amplitude attenuation characteristics were extracted. The "scale thickness-attenuation characteristics" correspondence set established in the early stage based on sample pipes of the same material and diameter and artificially made scale layers of 0.1-1.0mm was called up. The estimated value of scale thickness on the inner wall of the pipeline was calculated by comparison and linear interpolation. Initially, the thickness was about 0.2-0.3mm.

[0104] Next, the multi-phase flow data, health indicators, and thickness estimates are integrated and normalized to form a joint input sample. 10,000 samples from 3 months of normal system operation are used as training data. The TensorFlow framework is used to build a generative network and a discriminator network. The generative network contains an encoder with 2 convolutional layers and a decoder with 2 deconvolutional layers. The discriminator network contains 3 fully connected layers. Adversarial training is carried out by alternately updating parameters until the difference between the discriminator's output for real samples and reconstructed samples is lower than a preset threshold, so that the generative network can output reconstructed data that is highly similar to normal samples.

[0105] After entering the prediction phase, new data is collected in real time, and health indicators and scale thickness are calculated. At this point, the health indicator score drops to 72-75, and the scale thickness increases to 0.6-0.7 mm. These data together constitute the real observation data. The real observation data is input into the generative network to obtain reconstructed data, and the weighted summation of the reconstruction error is calculated. Simultaneously, it is input into the discriminator network to obtain the anomaly score. At this point, the anomaly score drops to 0.4-0.45, and the lower the score, the more abnormal the data. An adaptive threshold is invoked, which is dynamically adjusted based on historical data. The current reconstruction error threshold is 0.3, and the anomaly score threshold is 0.5. The comparison shows that the reconstruction error exceeds the threshold, and the anomaly score is below the threshold, and this state has been maintained for 5 consecutive 15-minute monitoring cycles. Based on similar historical cases, the deterioration trend of the system status is inferred, and the remaining effective operating time is predicted to be 8-10 days. A suggestion is pushed to the operation and maintenance system, prompting that pipeline descaling be completed within 7 days.

[0106] The data center equipment fault prediction method based on machine learning provided in this application breaks through the limitations of single-parameter monitoring by deeply integrating key data from multiple sources. It uses scale thickness as a fundamental parameter to fundamentally solve the problems of false alarms and missed alarms in traditional solutions. It introduces an adversarial training model, which can accurately learn the distribution of normal operating conditions and generate highly realistic reconstructed data without a large amount of fault data. This overcomes the modeling difficulties in scenarios with scarce fault samples and greatly expands the scope of application. By adapting dynamic thresholds to different scenario changes and predicting the remaining time, it upgrades the binary judgment of "whether there is a fault" to a quantitative guidance of "when the fault will occur". This completely changes the passive situation of "early warning equals emergency repair" in traditional solutions, leaving sufficient time for operation and maintenance, significantly reducing the risk of chip overheating and equipment downtime, and ensuring the long-term stable operation of the liquid cooling system. Compared with traditional scale monitoring solutions for liquid cooling systems, it achieves multi-dimensional technological breakthroughs.

[0107] Figure 4 This is a schematic diagram illustrating a specific implementation of a machine learning-based data center equipment fault prediction system provided in this application. (Refer to...) Figure 4 The system may include:

[0108] The acquisition module 41 is used to acquire monitoring data of the data center cooling system to form multi-phase flow data, and calculate the overall health status index of the system based on the correlation between the multi-phase flow data.

[0109] The calculation module 42 is used to generate micro-vibrations through a piezoelectric actuator and calculate the estimated thickness of the scale layer on the inner wall of the pipe based on the attenuation change of the feedback signal of the micro-vibrations.

[0110] The reconstruction module 43 is used to take the multi-phase flow data, system health status indicators and thickness estimate as joint inputs, and learn the probability distribution of the joint inputs under normal operating conditions by introducing a discriminator network into the generative network structure for adversarial training, and generate reconstructed data that is difficult to distinguish from the normal state.

[0111] The prediction module 44 is used to calculate the reconstruction error between the actual observation data and the reconstructed data during the prediction phase, and to obtain the anomaly score output by the discriminator network. When the values ​​of the reconstruction error and the anomaly score exceed the adaptive threshold dynamically adjusted based on historical operating data, it is determined that the liquid cooling system is evolving towards a scaling failure state, and the remaining effective operating time is predicted.

[0112] The machine learning-based data center equipment fault prediction system of this application is used to implement the aforementioned machine learning-based data center equipment fault prediction method. Therefore, the specific implementation of the machine learning-based data center equipment fault prediction system can be found in the embodiment section of the machine learning-based data center equipment fault prediction method above. The specific implementation can be referred to the description of the corresponding embodiment, and will not be repeated here.

[0113] This application also provides an electronic device, including: a memory for storing a computer program; and a processor for executing the computer program to implement the steps of any of the above-described machine learning-based data center equipment fault prediction methods.

[0114] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of any of the above-described machine learning-based data center equipment fault prediction methods.

[0115] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as USB flash drives, read-only memory, random access memory, portable hard drives, magnetic disks, or optical disks.

[0116] Embodiments of the present invention also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the embodiments of the machine learning-based data center equipment fault prediction method described above.

[0117] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0118] The foregoing has provided a detailed description of a data center equipment fault prediction method, system, electronic device, and storage medium based on machine learning, as provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and its core ideas. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of this application.

Claims

1. A machine learning based data center equipment failure prediction method, characterized by, The method comprises the following steps: Collecting monitoring data of a data center cooling system, the monitoring data comprising temperature data, water conductivity data, and fluid pressure and flow data, time-aligning the temperature data, water conductivity data, and fluid pressure and flow data, constructing multi-phase flow data, and calculating a system overall health state index based on the correlation between the multi-phase flow data; Generating micro-vibration through a piezoelectric actuator, and calculating an estimated thickness of a fouling layer on the inner wall of a pipeline according to the attenuation change of a feedback signal of the micro-vibration; Taking the multi-phase flow data, system health state index, and estimated thickness as joint inputs, learning the probability distribution of the joint inputs under normal working conditions through adversarial training by introducing a discriminator network in a generative network structure, and generating reconstructed data that is difficult to distinguish from the normal state; In the prediction stage, the reconstruction error between the real observation data and the reconstructed data is calculated, and the abnormal score output by the discriminator network is obtained, and when the values of the reconstruction error and the abnormal score exceed the adaptive threshold dynamically adjusted based on historical operation data, it is determined that the liquid cooling system is evolving towards a fouling failure state, and the remaining effective operation time is predicted.

2. The method of claim 1, wherein, Taking the multi-phase flow data, system health state index, and estimated thickness as joint inputs, learning the probability distribution of the joint inputs under normal working conditions through adversarial training by introducing a discriminator network in a generative network structure, and generating reconstructed data that is difficult to distinguish from the normal state, comprising: Combining the multi-phase flow data, the health state index, and the estimated thickness into multi-dimensional joint input samples; Respectively constructing a generative network and a discriminator network, the generative network being used to receive the joint input samples and generate reconstructed samples by learning the probability distribution of the joint input samples under normal working conditions; The discriminator network is used to distinguish whether the input sample is a real normal working condition sample or the reconstructed sample; Using a large number of normal working condition joint input samples, the generative network and the discriminator network are alternately trained adversarially until the discriminator network cannot distinguish between real normal samples and reconstructed samples; After training, the reconstructed samples generated by the generative network on normal working condition samples are used as reconstructed data corresponding to the normal state probability distribution.

3. The method of claim 2, wherein, Using a large number of normal working condition joint input samples, the generative network and the discriminator network are alternately trained adversarially until the discriminator network cannot distinguish between real normal samples and reconstructed samples, comprising: Using a large number of normal working condition joint input samples, alternately updating the internal parameters of the discriminator network and the generative network; After each alternately updating, calculating the average difference of the output values of the discriminator network for a batch of real samples and corresponding reconstructed samples, and stopping training when the average difference is lower than a preset threshold.

4. The method of claim 1, wherein, Generating micro-vibration through a piezoelectric actuator, and calculating an estimated thickness of a fouling layer on the inner wall of a pipeline according to the attenuation change of a feedback signal of the micro-vibration, comprising: extracting an attenuation characteristic of a feedback signal of the micro-vibration generated by the piezoelectric actuator; comparing the attenuation characteristic with pre-calibrated reference attenuation characteristics corresponding to different fouling thicknesses to estimate a thickness estimate.

5. The method of claim 4, wherein, comparing the attenuation characteristic with pre-calibrated reference attenuation characteristics corresponding to different fouling thicknesses to estimate a thickness estimate, comprising: obtaining a set of pre-established correspondence relations, each entry of the set of correspondence relations containing a known fouling layer thickness value and a corresponding reference attenuation characteristic value; comparing the attenuation characteristic extracted at present with each reference attenuation characteristic value in the set of correspondence relations one by one to determine two reference attenuation characteristic values with the highest similarity to the current attenuation characteristic; performing interpolation calculation between the fouling layer thickness values corresponding to the two reference attenuation characteristic values according to the similarity of the current attenuation characteristic to the two reference attenuation characteristic values, and taking the value obtained by interpolation calculation as the thickness estimate of the fouling layer on the inner wall of the pipeline.

6. The method of claim 1, wherein, In the prediction stage, the reconstruction error between the real observation data and the reconstructed data is calculated, and the anomaly score output by the discriminator network is obtained, comprising: In the prediction stage, the newly obtained real observation data is input into the trained generative network, and the corresponding reconstructed data is output through the generative network; calculate the difference between the real observation data and the reconstructed data in each corresponding dimension, and integrate the difference in each corresponding dimension into a reconstruction error; input the real observation data into the trained discriminator network, and obtain the corresponding anomaly score through the discriminator network, which represents the possibility value of being judged as normal data.

7. The method of claim 1, wherein, When the values of the reconstruction error and the anomaly score exceed the adaptive threshold dynamically adjusted based on historical operation data, it is determined that the liquid cooling system is evolving towards a fouling failure state, and the remaining effective operation time is predicted, comprising: respectively compare the reconstruction error and the anomaly score of the current data with the system historical adaptive threshold, and when at least one of the reconstruction error and the anomaly score continuously exceeds the corresponding adaptive threshold, it is determined that the liquid cooling system is evolving towards a fouling failure state; According to the degree and duration that the reconstruction error and anomaly score exceed the threshold, the state deterioration trend of the system is inferred, and the remaining effective operation time is predicted according to the state deterioration trend.

8. A machine learning based data center equipment failure prediction system, characterized by, comprising: a collection module for collecting monitoring data of the data center cooling system, the monitoring data including temperature data, water conductivity data, and fluid pressure and flow data, time-aligning the temperature data, water conductivity data, and fluid pressure and flow data to form multi-phase flow data, and calculating the health status index of the system as a whole based on the correlation between the multi-phase flow data; a calculation module for generating micro-vibration through a piezoelectric actuator and calculating a thickness estimate of the fouling layer on the inner wall of the pipeline according to the attenuation change of the feedback signal of the micro-vibration; The reconstruction module is configured to take the multi-phase flow data, the system health state indicator, and the thickness estimation value as joint inputs, perform adversarial training by introducing a discriminator network in a generative network structure, learn a probability distribution of the joint inputs under a normal working condition, and generate reconstructed data that is difficult to distinguish from a normal state; The prediction module is configured to calculate a reconstruction error between real observation data and the reconstructed data in a prediction stage, obtain an anomaly score output by the discriminator network, and determine that the liquid cooling system is evolving towards a fouling failure state and predict a remaining effective operation time when values of the reconstruction error and the anomaly score exceed adaptive thresholds dynamically adjusted based on historical operation data.

9. An electronic device, comprising: The method comprises the following steps: a memory configured to store a computer program; a processor configured to implement steps of the machine learning-based data center equipment failure prediction method according to any one of claims 1 to 7 when the computer program is executed.

10. A computer-readable storage medium, characterized in that, The computer program stored in the computer readable storage medium can implement the machine learning-based data center equipment failure prediction method according to any one of claims 1 to 7 when executed by the processor.

Citation Information

Patent Citations

  • AI intelligent diagnosis method and system based on intelligent system

    CN120233758A

  • Switch cabinet partial discharge fault identification method based on multi-fault information fusion

    CN120724277A