Fault prediction method and device, equipment and storage medium

By introducing a multi-level model architecture and adaptive weighted model into the fault prediction model, and combining transfer learning and incremental learning, the problem of poor adaptability of the existing fault prediction model in a diverse environment is solved, achieving higher prediction accuracy and adaptability.

CN120045370APending Publication Date: 2025-05-27LIUZHOU DADI TELECOMM EQUIP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510063587.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing fault prediction models perform poorly and have poor adaptability when facing different equipment and diverse and complex computer room environments. They need to readjust the model parameters or training, and the deployment cost is high.

Method used

A multi-level model architecture is adopted, including global models, device type models and device individual models, combining adaptive weighting models and transfer learning and incremental learning mechanisms, dynamically adjusting model weights and updating parameters to adapt to changes in different devices and environments.

Benefits of technology

It improves the accuracy of fault prediction in the face of different equipment and diverse and complex computer room environments, enhances adaptability, reduces deployment costs, and reduces false alarm rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045370A_ABST
    Figure CN120045370A_ABST
Patent Text Reader

Abstract

The invention discloses a fault prediction method and device, equipment and a storage medium. The fault prediction method comprises the steps of inputting first feature data and second feature data corresponding to environment data of a target machine room and operation data of each device at the current moment into a fault prediction model of a multi-stage model architecture with a global model, a device type model and a device individual model, and outputting a fault prediction result, the fault prediction result comprises a fault prediction result corresponding to the macroscopic environmental data prediction result of the target machine room, a fault prediction result corresponding to the target equipment type and a fault prediction result corresponding to the target equipment, so that the fault prediction method can pay attention to the macroscopic trend and microscopic characteristics of the target machine room equipment at the same time; the overall safety of the machine room system is pre-warned, and meanwhile, the operation state of the specific equipment type and the characteristic equipment is concerned, so that the fault prediction accuracy in the face of different equipment and various and complex machine room environments is improved, and the adaptability is relatively high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of fault prediction and diagnosis, and particularly to a fault prediction method, device, equipment and storage medium. Background Art

[0002] The fault prediction model is the core module in the fault prediction and diagnosis system for the infrastructure of communication computer rooms. At present, common fault prediction models mainly include an anomaly detection module, a time series analysis module, an equipment fault prediction module, and a model optimization and adaptive update mechanism. Among them, the task of the anomaly detection module is to timely identify anomalies in the device operation data before obvious faults occur in the device, usually including key technologies such as Isolation Forest, probability density-based anomaly detection (PCA + KDE), and rule engine combined with machine learning; the task of the time series analysis module is to predict the future state of the device by modeling the historical data of the device, and the time series analysis methods it adopts include Long Short-Term Memory Network (LSTM), Autoregressive Integrated Moving Average Model (ARIMA), Holt-Winters method, etc.; the equipment fault prediction module is trained by machine learning algorithms based on fault sample data and health status indicators extracted by feature extraction to predict the future health status and possible fault risks of the device, and common algorithms include Random Forest, Gradient Boosting Decision Trees (GBDT), Support Vector Machine (SVM), Health Score Model, etc.; the model optimization and adaptive update mechanism is used to optimize the fault prediction model and perform adaptive updates in response to changes.

[0003] Although the above-mentioned fault prediction model is comprehensively designed, there are still some disadvantages and limitations. This fault prediction model generally optimizes for a specific computer room or the target device type in the computer room, but a computer room usually contains multiple devices and the environment is complex and diverse. Therefore, this fault prediction model performs poorly and has poor adaptability when facing other devices or in different environments, and it is necessary to re-adjust the model parameters or train, resulting in a high deployment cost. Therefore, there are still technical problems to be solved in the related technologies. Summary of the Invention

[0004] Embodiments of the present invention provide a fault prediction method, device, equipment and storage medium, aiming to improve the fault prediction accuracy when facing different devices and diverse and complex computer room environments, and improve the adaptability of the fault prediction method.

[0005] In a first aspect, an embodiment of the present invention provides a fault prediction method, which includes:

[0006] Obtain the environmental data of the target computer room and the operation data of each device in the target computer room at the current moment, and preprocess the environmental data and the operation data to obtain the first feature data corresponding to the environmental data and the second feature data corresponding to the operation data of each device;

[0007] Input the first feature data and the second feature data into a preset fault prediction model to obtain a fault prediction result. The fault prediction result includes a first prediction result, a second prediction result, and a third prediction result. The fault prediction model includes a global model, a device type model, and a device individual model. The global model is configured to output an environmental data prediction result of the target computer room according to the first feature data, and perform fault prediction based on the environmental data prediction result to obtain the first prediction result. The device type model is configured to output a second prediction result corresponding to the target device type according to the second feature data corresponding to the target device type, where the target device type is preset. The device individual model is configured to output a third prediction result corresponding to the target device according to the second feature data corresponding to the target device, where the target device is preset.

[0008] In a second aspect, an embodiment of the present invention provides a fault prediction device, where the fault prediction device includes:

[0009] An acquisition unit, configured to obtain the environmental data of the target computer room and the operation data of each device in the target computer room at the current moment, and preprocess the environmental data and the operation data to obtain the first feature data corresponding to the environmental data and the second feature data corresponding to the operation data of each device;

[0010] A prediction unit, configured to input the first feature data and the second feature data into a preset fault prediction model to obtain a fault prediction result. The fault prediction result includes a first prediction result, a second prediction result, and a third prediction result. The fault prediction model includes a global model, a device type model, and a device individual model. The global model is configured to output an environmental data prediction result of the target computer room according to the first feature data, and perform fault prediction based on the environmental data prediction result to obtain the first prediction result. The device type model is configured to output a second prediction result corresponding to the target device type according to the second feature data corresponding to the target device type, where the target device type is preset. The device individual model is configured to output a third prediction result corresponding to the target device according to the second feature data corresponding to the target device, where the target device is preset.

[0011] In a third aspect, an embodiment of the present invention further provides a fault prediction device, including a processor and a memory. The memory stores a computer program, and when the computer program is executed by the processor, the processor is caused to execute the steps of any one of the fault prediction methods provided by the embodiments of the present invention.

[0012] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium. The computer-readable storage medium includes a computer program, and when the computer program runs on an electronic device, the computer program is used to cause the electronic device to execute the steps of any one of the fault prediction methods provided by the embodiments of the present invention.

[0013] The beneficial effects of the present invention are as follows:

[0014] By obtaining the environmental data of the target computer room and the operation data of each device in the target computer room at the current moment, and preprocessing the environmental data and operation data to obtain corresponding first feature data and second feature data, then inputting the first feature data and second feature data into a fault prediction model with a multi-level model architecture including a global model, a device type model, and a device individual model, and outputting a fault prediction result, the fault prediction result includes the fault prediction result corresponding to the predicted result of the environmental data of the target computer room macroscopically, the fault prediction result corresponding to the target device type in the target computer room, and the fault prediction result corresponding to the target device in the target computer room. Therefore, the fault prediction method of the present invention can simultaneously focus on the macroscopic trend and microscopic characteristics of the target computer room devices, pay attention to the operation status of specific device types and characteristic devices while warning the overall safety of the computer room system, improve the fault prediction accuracy when facing different devices and diverse and complex computer room environments, and has high adaptability. Description of the Drawings

[0015] Figure 1 is a schematic flowchart of an embodiment of the fault prediction method provided by the embodiment of the present invention;

[0016] Figure 2 is a schematic structural diagram of the fault prediction device provided by the embodiment of the present invention;

[0017] Figure 3 is a schematic structural diagram of the fault prediction device provided by the embodiment of the present invention. Detailed Embodiments

[0018] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention. At the same time, in the description of the embodiments of the present invention, terms such as "first" and "second" are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more features. In the description of the embodiments of the present invention, the meaning of "a plurality" is two or more, unless otherwise specifically defined.

[0019] The embodiments of the present invention provide a fault prediction method, device, equipment and storage medium.

[0020] Specifically, this embodiment will be described from the perspective of the fault prediction device. The fault prediction device can be specifically integrated in the fault prediction equipment. The fault prediction equipment can be an image sensor chip, an image sensor, a camera, or an electronic device such as a mobile phone or a computer. That is, the fault prediction method in the embodiments of the present invention can be executed by the fault prediction equipment.

[0021] The following will be described in detail with reference to the accompanying drawings. In this embodiment, the execution subject is the fault prediction equipment as an example. It should be noted that the description order of the following embodiments does not limit the preferred order of the embodiments. Although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from that shown in the drawings.

[0022] The fault prediction model is the core module in the fault prediction and diagnosis system for the infrastructure of the communication machine room. Currently, the common fault prediction models mainly include an anomaly detection module, a time series analysis module, an equipment fault prediction module, and a model optimization and adaptive update mechanism. Among them:

[0023] The task of the anomaly detection module is to timely identify anomalies in the device operation data before obvious faults occur in the device. The design of the anomaly detection module includes several key technologies such as Isolation Forest, probability density-based anomaly detection (PCA+KDE), and rule engine combined with machine learning. Among them, Isolation Forest is an unsupervised learning algorithm for detecting abnormal data and is applicable to multi-dimensional data. This algorithm constructs multiple randomly split trees and calculates the isolation degree of each data point, featuring high efficiency and low storage requirements. Isolation Forest is suitable for detecting abnormal points in features such as server CPU, memory, and temperature; Probability density-based anomaly detection uses principal component analysis (PCA) to reduce the dimension of high-dimensional data, extracts the main features, and then uses kernel density estimation (KDE) to judge the anomaly degree of the data distribution. This method is suitable for processing real-time changing data streams such as network device traffic load and latency; The rule engine combined with machine learning, for some key devices (such as UPS and air conditioners) and environmental data, combines simple rules (such as too high temperature or abnormal humidity) and is used jointly with machine learning detection methods to achieve rapid detection. For example, when the temperature sensor detects that the data exceeds the predetermined threshold, the alarm mechanism can be directly triggered.

[0024] The data of the power and environment system (dynamic environment system) and IT devices often present time series characteristics. The task of the time series analysis module is to predict the future state of the device by modeling the historical data of the device. The time series analysis methods adopted by the time series analysis module include Long Short-Term Memory Network (LSTM), Autoregressive Integrated Moving Average Model (ARIMA), and Holt-Winters method. Among them, LSTM is a recurrent neural network structure suitable for time series data and is good at learning the long-term dependence of data. The LSTM module is trained based on continuous time data such as temperature, humidity, and server performance metrics, and can predict the future parameter change trend, such as the change trend of temperature or power consumption; ARIMA is a traditional time series analysis method applicable to the periodic data of the device and can perform short-term trend prediction. For seasonal or periodic data such as power consumption and network traffic, ARIMA performs well in the short term and can be used to predict future load changes; The Holt-Winters method is suitable for data with seasonal characteristics, such as the fluctuations of environmental variables such as computer room temperature and humidity. By decomposing the trend and seasonal components, this method can effectively cope with the fluctuations of seasonal changes and provide accurate predictions.

[0025] The device fault prediction module is trained through machine learning algorithms based on fault sample data and health status indicators extracted from features to predict the possible fault risks of the future health status of the device. Commonly used algorithms include Random Forest, Gradient Boosting Decision Trees (GBDT), Support Vector Machine (SVM), Health Score Model, etc. Among them, Random Forest is an ensemble learning algorithm based on decision trees, with high accuracy and interpretability. Random Forest analyzes features (such as temperature, power consumption, CPU occupancy rate), generates a health status score, and outputs the fault probability. Random Forest is suitable for fault prediction of IT devices and network devices; GBDT continuously iterates and optimizes decision trees and is suitable for scenarios with relatively complex fault types, such as UPS battery performance degradation, air conditioning system failures, etc. GBDT combines the non-linear relationships of different features and can more precisely predict the attenuation trend or fault risk of the device; SVM performs excellently when the data dimension is high or there are obvious boundaries between features. For some key devices, such as precision air conditioners, SVM can determine whether the device is within the normal range through the boundary; The Health Score Model combines multi-dimensional features of IT devices, network devices, and power and environmental devices, and obtains the health score of the device through comprehensive calculation. The health score is standardized as the comprehensive health status score of the device, used to evaluate the overall operation of the device and predict the possible fault time window.

[0026] Although the above-mentioned fault prediction model is designed relatively comprehensively, there are still some drawbacks and limitations. This fault prediction model generally optimizes for a specific computer room or the target device type in the computer room. However, a computer room usually contains multiple devices and the environment is complex and diverse. Therefore, the performance of this fault prediction model is poor and the adaptability is low when facing other devices or in different environments, and it is necessary to re-adjust the model parameters or retrain, resulting in a high deployment cost. It should be noted that in the above-mentioned fault prediction model, although multi-model fusion improves the prediction accuracy, it also increases the complexity and debugging difficulty of the system. At the same time, the computational overhead brought by multi-model fusion is greater, and the update and optimization of the model become more complex, resulting in an increase in the operation and maintenance cost. In addition, the operating environment, load, and device type of the devices in the computer room may change, which will cause the data distribution to drift, thereby reducing the accuracy of the prediction results output by the model.

[0027] To solve the above problems, the present invention discloses a fault prediction method. Please refer to Figure 1 , and the specific process of this fault prediction method can be as follows in steps S101 to S102, where:

[0028] Step S101, obtain the environmental data of the target computer room and the operation data of each device in the target computer room at the current moment, and preprocess the environmental data and the operation data to obtain the first feature data corresponding to the environmental data and the second feature data corresponding to the operation data of each device.

[0029] Among them, the target computer room is the object for which the fault prediction method of the embodiment of the present invention performs fault prediction, and can be any communication computer room. In practical applications, a communication computer room includes many different types of devices, such as IT devices, network devices, and power and environmental devices. The operation data and environmental data of these devices can be collected and transmitted in real time by sensors or monitoring systems.

[0030] Exemplarily, consider a typical computer room environment, which includes several servers and network switches. To monitor the operation status of these devices, the following data of these devices needs to be obtained:

[0031] IT device data: For example, the CPU usage rate, memory occupancy rate, disk read and write speed of the server, etc. These metrics can reflect whether the device is overloaded and whether there are resource bottlenecks, thus helping to predict the decline of system performance or potential hardware failures.

[0032] Network device data: For example, the port traffic, network latency, packet loss rate of the switch, etc. These data help to evaluate the health status of the network. If packet loss occurs on the switch port or the network traffic fluctuates abnormally, it may indicate a network device failure or configuration problem.

[0033] Power and environment data: For example, the temperature, humidity, power supply status, UPS battery power in the computer room, etc. These environmental and power factors are crucial for the long-term stable operation of the device. If the temperature is too high or the UPS battery power is insufficient, it may cause device overheating or power interruption, resulting in system failures.

[0034] In the embodiment of the present invention, sensors and intelligent monitoring devices arranged at various positions in the target computer room are used to monitor these metrics in real time, and the collected data is preprocessed for further analysis.

[0035] Specifically, the embodiment of the present invention preprocesses the environmental data and the operation data to obtain the first feature data corresponding to the environmental data and the second feature data corresponding to the operation data of each device, which can convert the massive operation data and environmental data of the target computer room collected into information meaningful for the health status of the target computer room, facilitating subsequent fault prediction based on the first feature data and the second feature data obtained after preprocessing.

[0036] Further, in some embodiments, preprocessing the environmental data and the operation data to obtain first feature data corresponding to the environmental data and second feature data corresponding to the operation data of each device may include:

[0037] Perform data cleaning, data standardization processing, and feature extraction on the environmental data and the operation data in sequence to obtain the first feature data and the second feature data.

[0038] Among them, data cleaning is the first step of data preprocessing. Due to reasons such as sensor or network fluctuations, the environmental data and operation data collected in step S101 may have noise or missing values. Exemplarily, when monitoring the CPU usage rate of a server, abnormal data points may occur due to hardware failures or network interruptions. These abnormal data points will mislead the subsequent data analysis process, thereby reducing the reliability and accuracy of the prediction results. In the embodiments of the present invention, through data cleaning, abnormal data points in the environmental data and operation data are identified and removed to ensure that the data finally input into the fault prediction model is clean and reliable.

[0039] Data standardization is the second step of data preprocessing. Since the data sources of various devices in the target computer room are different, their dimensions and units may vary greatly. For example, the unit of temperature is degrees Celsius, while the server CPU usage rate is a percentage value. If these data are directly analyzed, unnecessary errors may occur. In the embodiments of the present invention, through data standardization, different types of data can be converted into a unified scale. Exemplarily, some embodiments of the present invention can adopt normalization or standardization methods to make each environmental data and each operation data within a common range, facilitating comprehensive analysis.

[0040] Finally, feature extraction is to extract important information related to faults from the cleaned and standardized data. For example, by statistically analyzing the historical data such as the CPU usage rate and memory occupancy rate of a server, it can be determined which characteristic parameters (such as the fluctuation range of CPU load, the growth rate of memory occupancy, etc.) have a strong correlation with the occurrence of faults. These features can be used as inputs for subsequent prediction models to help identify potential device faults or performance bottlenecks.

[0041] It can be understood that through the above data preprocessing operations, the embodiments of the present invention can convert the collected environmental data and operation data into features that are practically meaningful for fault prediction, namely the first feature data and the second feature data, so as to facilitate subsequent fault prediction and diagnosis based on these feature data. The above data preprocessing process not only improves the quality of the data, but also provides strong support for fault prediction. Based on the finally obtained first feature data and second feature data, the embodiments of the present invention can more accurately identify which devices may fail in the future, thereby realizing timely early warning and repair.

[0042] Step S102: Input the first feature data and the second feature data into a preset fault prediction model to obtain a fault prediction result.

[0043] Among them, the fault prediction result includes a first prediction result, a second prediction result, and a third prediction result, and the fault prediction model includes a global model, a device type model, and a device individual model; the global model is configured to output an environmental data prediction result of the target computer room according to the first feature data, and perform fault prediction based on the environmental data prediction result to obtain the first prediction result; the device type model is configured to output a second prediction result corresponding to the target device type according to the second feature data corresponding to the target device type, and the target device type is preset; the device individual model is configured to output a third prediction result corresponding to the target device according to the second feature data corresponding to the target device, and the target device is preset.

[0044] It should be noted that the fault prediction model of the embodiments of the present invention has a multi-level model architecture composed of the above global model, device type model, and device individual model, aiming at the fault prediction requirements of different types of devices in the target computer room, and performing data analysis and prediction at different levels through a hierarchical structure. Among them, the global model, device type model, and device individual model respectively process macroscopic fault prediction, device category fault prediction, and refined prediction of key devices, so as to improve the fault prediction accuracy and adaptability when facing multiple devices and diverse complex environments.

[0045] It can be understood that, as described above, existing fault prediction models generally optimize for specific computer rooms or target device types within a computer room. However, a computer room usually contains multiple types of devices and has a complex and diverse environment. Therefore, the performance of existing fault prediction models is poor and their adaptability is low when facing other devices or in different environments. The multi-level model architecture adopted in the embodiments of the present invention is a hierarchical fault prediction method, aiming to cope with the diversity and complexity of device types and environments in the target computer room, so as to improve the adaptability and prediction accuracy of the fault prediction model. The multi-level model architecture is divided into three main levels: the global model, the device type model, and the device individual model. The models at each level respectively perform prediction and analysis on the overall device status, device category characteristics, and unique operating conditions of individual devices in the computer room, ensuring that the system can respond in a timely manner to abnormal situations at different levels and provide effective fault prediction information. Therefore, the embodiments of the present invention improve the ability of the fault prediction model to adapt to different devices and environments through hierarchical prediction at three levels: the global model, the device type model, and the device individual model.

[0046] Specifically, the global model is configured to output a prediction result of the environmental data of the target computer room based on the first feature data, and perform fault prediction based on the prediction result of the environmental data to obtain a first prediction result. The global model is the first level in the multi-level model architecture, and is used to capture the macroscopic operation modes and states of all devices in the target computer room. It focuses on the detection of abnormalities in overall trends and can quickly identify possible widespread abnormal events at the system level, such as an increase in the overall temperature or humidity of the target computer room, an increase in power load, etc.

[0047] Optionally, in some embodiments, the global model can model the overall distribution of data such as temperature, humidity, and current in the computer room based on anomaly detection algorithms such as Isolation Forest and PCA. When it detects that these parameters deviate from the normal range, it triggers a preliminary alarm signal. The main advantage of the global model is its high operating efficiency, which is suitable for real-time monitoring of the macroscopic state of the entire system and can quickly screen out potential risks for further detailed analysis at subsequent levels.

[0048] The device type model is configured to output a second prediction result corresponding to the target device type according to the second feature data corresponding to the target device type. The device type model is the middle layer of the multi-level model architecture and focuses on the detailed prediction of a specific device type (i.e., the pre-set target device type). Different devices in the target computer room (such as servers, switches, UPSs, air conditioners, etc.) vary in operating characteristics, failure modes, and data characteristics, so separate modeling is required. Each device type model is optimized for the key features of the device type and can more accurately capture the failure modes of this type of device. For example, the server model may predict abnormal fluctuations in CPU occupancy based on LSTM, the air conditioner device model may use ARIMA to predict the temperature and humidity change trend, and the UPS model may adopt an algorithm based on health scores to monitor the decay of battery performance. The device type model can utilize the characteristics of different device types to provide refined fault warning information, making the prediction results more targeted and accurate.

[0049] The device individual model is configured to output a third prediction result corresponding to the target device according to the second feature data corresponding to the target device. The device individual model, as the bottom layer of the multi-level model architecture, focuses on modeling the micro state of a single key device (i.e., the pre-set target device). The device individual model can independently learn the specific operating states, environmental variables, and other operating data of each device, ensuring that the individual differences of key devices are fully considered. Exemplarily, for a core server, an embodiment of the present invention can train a dedicated model that records and analyzes the specific operating mode, historical load, and temperature fluctuation law of the server. If the CPU occupancy and temperature of the server suddenly increase during a certain period, but the overall computer room and the operation of similar servers remain normal, the device individual model will be able to identify this unique anomaly and give an alarm in a timely manner, thus effectively reducing the false alarm risk. The device individual model is particularly important for critical devices because it can deeply analyze the unique behavior of the device and discover hidden fault problems that are difficult to detect by the global and type models.

[0050] It can be understood that, through the multi-level model architecture in the embodiment of the present invention, the global model, the device type model, and the device individual model progress step by step, enabling the fault prediction module to simultaneously focus on the macro trends and micro characteristics of the device. Among them, the global model ensures overall safety, the device type model improves the prediction accuracy, and the device individual model ensures the reliable operation of key devices. This multi-level model architecture not only improves the adaptability of the fault prediction method in the embodiment of the present invention but also can significantly reduce the false alarm rate, making the entire fault prediction method more intelligent and efficient in dealing with complex and diverse computer room environments.

[0051] Furthermore, as can be seen from the foregoing, the multi-model combination adopted by the existing fault prediction model will bring greater computational overhead, and at the same time, the update and optimization of the model become more complex, resulting in an increase in operation and maintenance costs. To avoid this problem, the multi-level model architecture adopted in the embodiments of the present invention introduces an adaptive weighted model. Before step S102, the fault prediction method in the embodiments of the present invention may further include:

[0052] Input the first feature data and the second feature data into a pre-trained adaptive weighted model to obtain the weights of the machine learning models of the fault prediction model.

[0053] It should be noted that the technology adopted in the embodiments of the present invention is based on an Adaptive Weighted Ensemble Model, which is a technology that integrates multiple machine learning models and assigns dynamic weights to the prediction results of each model, thereby improving the overall prediction accuracy and adaptability. This method is particularly suitable for multi-source data prediction scenarios, and can adjust the weights according to the real-time performance of each model to cope with the dynamic changes of device types, environmental changes, and data characteristics. Compared with the model fusion method with fixed weights, the adaptive weighted ensemble model is more flexible and adaptable, effectively reducing the complexity brought by model fusion.

[0054] It can be understood that with the help of a pre-trained adaptive weighted model, the fault prediction model in the embodiments of the present invention can automatically adjust the weights according to the real-time performance of each machine learning model, avoiding manual adjustment and complex fusion strategies. Through this dynamic optimization, the embodiments of the present invention simplify the fusion process of multiple models, can effectively improve the prediction accuracy, and reduce the complexity and computational overhead that may occur in traditional model fusion.

[0055] Specifically, the adaptive weighted ensemble model is composed of multiple different types of machine learning models, and each model analyzes data features from different perspectives and provides independent prediction results. For example, the random forest model can capture complex non-linear relationships and is suitable for identifying discrete features of device operating states; the LSTM model is good at processing time series data and can be used to capture the time trends of temperature, humidity, or power consumption; the ARIMA model is suitable for analyzing periodic variables such as load or network traffic. During the fault prediction process, the embodiments of the present invention dynamically adjust the weights according to the performance of each model on real-time data, and this weighting process can be based on indicators such as the prediction accuracy, error size, or confidence of each model in the past period of time.

[0056] Optionally, in some embodiments, the dynamic adjustment of model weights can be achieved through meta-heuristic algorithms such as Bayesian optimization or genetic algorithms. Embodiments of the present invention evaluate the historical performance of each model with the help of these algorithms and automatically update the weights to ensure the best combined effect of the model under specific conditions. More specifically, the weight adjustment mechanism optimizes through an objective function, assigns weights to the models that perform optimally in the current context, and guarantees the overall prediction accuracy. For example, if the recent historical data presents a strong time series pattern, embodiments of the present invention will assign a higher weight to the LSTM; if the discrete feature changes are more significant, the weight of the random forest will be increased. Through this adaptive weighting method, embodiments of the present invention can self-adjust under different data patterns and enhance its adaptability to the changing environment.

[0057] Exemplarily, take the server temperature prediction in a communication computer room as an example. Suppose it is necessary to predict the temperature of a server within the next hour to prevent possible overheating problems. In this specific embodiment, three basic models are selected: random forest, LSTM, and ARIMA, which are respectively suitable for different data feature analyses. To demonstrate the advantages of adaptive weighting, assume that various changes have occurred in the environmental conditions, such as the periodic changes in server load and the temperature and humidity in the computer room, as well as the abnormal state caused by a sudden increase in temporary load.

[0058] In the initial stage, equal weights are assigned to each model, that is, each model makes the same contribution to the prediction result. As time goes by, the LSTM model gradually captures the daily and nightly periodic changes in the temperature data, and the trend that the temperature rises with the load peak and drops with the load trough. In this context, the prediction accuracy of the LSTM is significantly better than that of other models. Therefore, the adaptive weighting model gradually increases the weight of the LSTM to ensure its greater contribution to the final prediction result. At the same time, since the ARIMA model can also identify periodic changes, the adaptive weighting model also assigns it a certain weight. The random forest is weak in capturing such long-term trends, so the adaptive weighting model reduces its weight allocation.

[0059] However, when the server load suddenly soars, causing the temperature to rise rapidly, the LSTM and ARIMA may be inaccurate in prediction because they fail to adapt to the sudden data changes in time. In this case, the random forest model shows faster responsiveness due to its ability to capture non-linear features and can quickly reflect the impact of emergencies on the temperature. Therefore, the adaptive weighting mechanism will increase the weight of the random forest model according to its performance, making its prediction dominant under sudden load conditions. Through the automatic adjustment of weights, the fault prediction model can quickly identify and respond to emergencies, effectively preventing the occurrence of overheating problems.

[0060] It can be seen from this that the embodiments of the present invention adopt an adaptive weighted integration model. The first feature data and the second feature data are input into a pre-trained adaptive weighted model to obtain the weights of each machine learning model of the fault prediction model. By dynamically allocating weights in different scenarios, the advantages of multiple machine learning models are combined, thereby achieving higher prediction accuracy and stability. This method can not only make full use of the specialties of different machine learning models, but also effectively cope with the complex operating environment and diverse data characteristics of the communication computer room. Compared with the model fusion method with fixed weights, the adaptive weighted method adopted in the embodiments of the present invention has stronger system adaptability through optimizing the weight allocation, and can still maintain high prediction accuracy under complex conditions such as data distribution changes and feature mutations. In practical applications, the embodiments of the present invention have significantly improved the prediction effect of various data of the target computer room through adaptive weight setting, and at the same time reduced the prediction inaccuracy caused by a single model's inability to adapt to all scenarios.

[0061] Optionally, in some embodiments, the training of the adaptive weighted model may include:

[0062] 1) Obtain the historical prediction results of the fault prediction model within a historical time interval, where the historical time interval is a time interval before the current moment;

[0063] 2) Evaluate the accuracy of the historical prediction results, and optimize the objective function of the adaptive weighted model according to the historical prediction results and the corresponding accuracy.

[0064] Optionally, in some embodiments, the fault prediction method may further include:

[0065] In response to a change in the device distribution or computer room environment of the target computer room, update the parameters of the fault prediction model according to the old data set and the new data set.

[0066] Wherein, the old data set is the data set before the change of the target computer room, and the new data set is the data set after the change of the target computer room.

[0067] According to the foregoing, the operating environment, load, and device type of the devices in the computer room may change, which may cause the data distribution to drift, thereby reducing the accuracy of the prediction results output by the model. For this reason, in the embodiments of the present invention, when the device distribution or computer room environment of the target computer room changes, the parameters of the fault prediction model are updated according to the data sets before and after the change of the target computer room, so that the fault prediction model can cope with the drift of the data distribution, that is, in the new device environment or operating conditions, the fault prediction model can quickly adapt and continuously optimize. This mechanism improves the stability of the fault prediction model, can effectively cope with the challenges brought by data changes, and reduces the negative impact of external environment changes on prediction accuracy.

[0068] Further, in some embodiments, updating the parameters of the fault prediction model according to the old dataset and the new dataset may include:

[0069] Adjusting the parameters of the fault prediction model based on transfer learning according to the old dataset and the new dataset;

[0070] Adjusting the parameters of the fault prediction model after transfer learning based on reinforcement learning according to the new dataset.

[0071] It can be understood that the embodiments of the present invention introduce transfer learning and incremental learning mechanisms (Transfer Learning and Incremental Learning) in the further fault prediction model update scheme. And in order to better cope with the problem of data distribution drift and improve the adaptability of the fault prediction model in different environments, the introduction of transfer learning and incremental learning is also refined. The combination of transfer learning and incremental learning can ensure that the model not only quickly adapts to new devices or environments, but also continuously optimizes its performance under gradually changing data distributions. Their applications and effects are illustrated below through the design of specific mechanisms and examples.

[0072] It should be noted that transfer learning generally performs fine-tuning based on a pre-trained model, that is, by adjusting the weights of the existing fault prediction model to make the fault prediction model adapt to the new environment. The basic principle of dynamically adjusting the weights is to retain most of the weights of the old model initially, because these weights represent the general characteristics of the device, and only adjust the specific data of the new environment. The update of the weights is usually controlled by the learning rate to balance the impact of new data on the model.

[0073] During the transfer learning process, the embodiments of the present invention combine the new dataset Dnew with the old dataset Dold and perform fine-tuning on the basis of the existing model. Assuming the parameters of the model are θ, the update of transfer learning can be expressed by the following formula:

[0074]

[0075] where η is the learning rate, is the gradient calculated on the new dataset Dnew. The selection of the learning rate η is crucial, which determines the influence size of the new dataset on the update of the model parameters. The embodiments of the present invention will select a relatively small learning rate during transfer learning, for example, η ∈ [10-4, 10-3], to avoid excessive interference of the new dataset on the existing model. It can be understood that too large a learning rate may cause the loss of old knowledge, while too small a learning rate may cause the adaptation process of new data to be too slow.

[0076] Exemplarily, assume that a fault prediction model is trained in the environment of Machine Room A, and the model performs well on the original features such as temperature, humidity, and server load. When the cooling system in Machine Room A changes, the temperature pattern will change. Use transfer learning to fine-tune the original model, set the learning rate to η = 10-4, and adjust the small sample data after the change of the cooling system based on the data in Machine Room A. In this way, the model can effectively adapt to the new temperature fluctuation pattern without losing the original knowledge.

[0077] In incremental learning, the fault prediction model needs to continuously adjust the weights according to the new data set to timely respond to the change of data distribution. In the embodiments of the present invention, a weighted average strategy based on historical data and new data in the new data set is usually used to dynamically update the weights of the model in incremental learning. Assume that at time step t, the parameters of the model are θt. When the new data Dt arrives, the update of incremental learning can be expressed as:

[0078]

[0079] where α is the learning rate of incremental learning, is the gradient calculated on the model parameters at time step t and the new data Dt. The key of incremental learning is the learning rate α, which is usually set to a small constant value, such as α ∈ [10-5, 10-3], to ensure the smoothness of model update. Incremental learning is particularly suitable for time series data, such as temperature and humidity, network traffic, etc., and it can ensure that the model is fine-tuned according to the latest operation data.

[0080] Exemplarily, assume that the temperature and humidity data in a certain machine room has seasonal fluctuations, and the fault prediction model is initially trained according to the temperature and humidity data in autumn and winter. As summer approaches, the temperature gradually rises, resulting in a change in the distribution of temperature and humidity in the machine room. To cope with this change, an incremental learning mechanism can be used to update the fault prediction model once a week, and update the model parameters θt through the temperature and humidity data Dt of the new week. If the learning rate α for each update is set to 10-4, the fault prediction model will gradually adapt to the temperature and humidity changes in the summer environment while avoiding being affected by the autumn and winter data.

[0081] The embodiments of the present invention combine transfer learning and incremental learning to dynamically adjust the weights of transfer learning and incremental learning. In practical applications, the combination of transfer learning and incremental learning can further enhance the adaptability and prediction ability of the fault prediction model. Especially when new devices are introduced or the environment changes greatly, rapid adaptation can be first achieved through transfer learning, and then the model can be updated in the long term using incremental learning to ensure that the model always maintains sensitivity to new data.

[0082] Specifically, the dynamic weight adjustment strategy for this combination is as follows:

[0083] Transfer learning stage: First, use transfer learning to fine-tune the existing fault prediction model. During the transfer learning process, use a relatively small learning rate, such as η = 10-4, so that the weights of the original fault prediction model will not change drastically due to new data. Adjust through a small amount of new data sets to make the fault prediction model adapt to the new feature distribution.

[0084] Incremental learning stage: When the fault prediction model adapts to the new environment, it enters the incremental learning stage. At this time, the embodiments of the present invention update the model weights regularly according to real-time data. The learning rate α during incremental learning is usually set to a value smaller than η (such as α = 10-5), so that the fault prediction model can be updated smoothly in a gradually changing environment.

[0085] Exemplarily, when a new high-performance server is added to the computer room, first use transfer learning to fine-tune the original fault and model to make it adapt to the operating characteristics of the new device. At this time, the embodiments of the present invention use the learning rate η = 10-4 of transfer learning for fine-tuning. Subsequently, as the device is put into use and new operation data is generated, the embodiments of the present invention adopt incremental learning to update the parameters of the fault prediction model. At this time, the learning rate is adjusted to a value smaller than η, such as α = 10-5, to gradually update the fault prediction model to adapt to the changes in long-term operation data.

[0086] According to the foregoing, through this dynamic adjustment mechanism that combines transfer learning and incremental learning, the fault prediction model of the embodiments of the present invention can quickly adapt when new devices or environments are introduced, and at the same time maintain the accuracy and stability of predictions when the data changes gradually. At the same time, this combined use method can greatly improve the adaptability of the fault prediction model in an uncertain environment and enable the fault prediction model to operate stably in a diverse computer room environment for a long time, especially suitable for computer room systems that are updated and expanded regularly.

[0087] Optionally, in some embodiments, the fault prediction method may further include:

[0088] 1) Perform fault alarm according to the fault prediction result and generate fault analysis information;

[0089] 2) Real-time display of environmental data, operation data, fault prediction results, and fault analysis information.

[0090] It can be understood that the embodiments of the present invention perform fault alarm according to the fault prediction result and generate fault analysis information to detect potential faults of the equipment in the target computer room and provide accurate fault diagnosis based on the fault prediction result. When a fault sign is detected by the fault prediction model, an alarm will be automatically triggered and detailed fault analysis information will be provided. Optionally, this information may include fault type, possible root cause, equipment status, and recommended handling steps.

[0091] Exemplarily, assume that the processor temperature of a server in the computer room is abnormal. The fault prediction model detects the trend of temperature increase, identifies the relevant historical fault data, and concludes that there may be a problem with the cooling system. At this time, an alarm message will be generated, including: "Warning: Abnormal CPU temperature of the server, possibly a cooling system failure. It is recommended to check the fan and radiator." Meanwhile, the relevant historical fault patterns and repair records of similar devices can also be displayed to help the operation and maintenance personnel quickly locate the problem.

[0092] In addition, in some embodiments, the fault alarm includes multi-level alarms, that is, different levels of alarms are triggered according to the severity or urgency of the fault. For example, a normal alarm is triggered when the temperature of the device exceeds the threshold, while an emergency alarm is triggered when the temperature continuously exceeds the safe range or the device shuts down. The alarm information is sent to the operation and maintenance personnel in the form of emails, text messages or system notifications to ensure timely response.

[0093] It can be understood that the embodiments of the present invention display the environmental data, operation data, fault prediction results and fault analysis information in real time, improving the intuitiveness of information presentation.

[0094] In some embodiments of the present invention, the environmental data, operation data, fault prediction results and fault analysis information are displayed in a graphical and intuitive manner, enabling the operation and maintenance personnel to more quickly understand the system health status and make effective decisions. The display interface usually includes modules such as device status monitoring dashboards, fault trend graphs, alarm history records, and real-time data stream displays.

[0095] Exemplarily, on a typical computer room visualization monitoring platform, the operation and maintenance personnel can see a real-time "dashboard" that shows the current status of each device (such as servers, UPSs, power supplies, etc.). If the health score of a device is lower than the preset threshold, the dashboard will automatically mark the device in red and provide detailed fault warning information. In addition, the fault trend graph shows the trend of device status changes over time. For example, if the temperature of a device has been continuously rising in the past few days, the trend graph will show an obvious upward curve to help the operation and maintenance personnel identify potential fault risks.

[0096] The alarm history record shows all the alarm events that have occurred and distinguishes different alarm levels through colors, icons, etc., enabling the operation and maintenance personnel to quickly review and analyze the fault history. The real-time data stream display provides the system real-time monitoring data, such as the temperature, humidity, voltage, load, etc. of the computer room. These data are combined with the operation status of the device and the fault prediction results to provide a global view.

[0097] In summary, in the embodiment of the present invention, the environmental data of the target computer room and the operation data of each device in the target computer room at the current moment are obtained, and the environmental data and the operation data are preprocessed to obtain the corresponding first feature data and second feature data. Then, the first feature data and the second feature data are input into a fault prediction model with a multi-level model architecture including a global model, a device type model, and a device individual model, and a fault prediction result is output. The fault prediction result includes the fault prediction result corresponding to the environmental data prediction result of the target computer room macroscopically, the fault prediction result corresponding to the target device type in the target computer room, and the fault prediction result corresponding to the target device in the target computer room. Therefore, the fault prediction method of the present invention can simultaneously focus on the macroscopic trend and microscopic characteristics of the target computer room equipment, pay attention to the operation status of specific device types and characteristic devices while warning the overall safety of the computer room system, improve the fault prediction accuracy when facing different devices and diverse and complex computer room environments, and has high adaptability.

[0098] The fault prediction method in the embodiment of the present invention introduces an adaptive weighted model, enabling the fault prediction model to automatically adjust the weights according to the real-time performance of each machine learning model, avoiding manual adjustment and complex fusion strategies. Through this dynamic optimization, the embodiment of the present invention simplifies the fusion process of multiple models, can effectively improve the prediction accuracy, and reduces the complexity and computational overhead that may occur in traditional model fusion.

[0099] When the device distribution or computer room environment in the target computer room changes, the fault prediction method in the embodiment of the present invention updates the parameters of the fault prediction model according to the data sets before and after the change in the target computer room, and further introduces a transfer learning and incremental learning mechanism, so that the fault prediction model can cope with the drift of the data distribution, that is, in a new device environment or operating condition, the fault prediction model can quickly adapt and continuously optimize. This mechanism improves the stability of the fault prediction model, can effectively cope with the challenges brought by data changes, and reduces the negative impact of external environment changes on the prediction accuracy.

[0100] This embodiment also provides a fault prediction device, which can be specifically integrated in a fault prediction device. For example, the pixel array of an image sensor chip includes a photosensitive area and an optical dark area, as Figure 2 shown, the fault prediction device may include:

[0101] An acquisition unit 201, configured to acquire the environmental data of the target computer room and the operation data of each device in the target computer room at the current moment, and preprocess the environmental data and the operation data to obtain first feature data corresponding to the environmental data and second feature data corresponding to the operation data of each device;

[0102] A prediction unit 202 is configured to input first feature data and second feature data into a preset fault prediction model to obtain a fault prediction result. The fault prediction result includes a first prediction result, a second prediction result, and a third prediction result. The fault prediction model includes a global model, a device type model, and a device individual model. The global model is configured to output an environmental data prediction result of a target computer room according to the first feature data, and perform fault prediction based on the environmental data prediction result to obtain the first prediction result. The device type model is configured to output a second prediction result corresponding to a target device type according to the second feature data corresponding to the target device type, where the target device type is preset. The device individual model is configured to output a third prediction result corresponding to a target device according to the second feature data corresponding to the target device, where the target device is preset.

[0103] As Figure 3 shown, Figure 3 FIG. is a schematic structural diagram of a fault prediction device provided by an embodiment of the present invention. The fault prediction device 1100 includes a processor 1101 having one or more processing cores, a memory 1102 having one or more computer-readable storage media, and a computer program stored on the memory 1102 and executable on the processor. Among them, the processor 1101 is electrically connected to the memory 1102. Those skilled in the art can understand that the structural diagram of the fault prediction device shown in the figure does not constitute a limitation on the fault prediction device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange different components.

[0104] The processor 1101 is the control center of the fault prediction device 1100, connects various parts of the entire fault prediction device 1100 through various interfaces and lines, runs or loads software programs and / or units stored in the memory 1102, and calls data stored in the memory 1102 to execute various functions of the fault prediction device 1100 and process data, thereby monitoring the fault prediction device 1100 as a whole. The processor 1101 may be a central processing unit (CPU), a graphics processing unit (GPU), a network processor (NP), etc., and may implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention.

[0105] In the embodiment of the present invention, the processor 1101 in the fault prediction device 1100 will load instructions corresponding to the processes of one or more application programs into the memory 1102 according to the following steps, and the processor 1101 will run the application programs stored in the memory 1102 to implement various functions. For details, reference may be made to the previous embodiments, which will not be elaborated here.

[0106] Optionally, as Figure 3As shown, the fault prediction device 1100 further includes: a touch display screen 1103, a radio frequency circuit 1104, an audio circuit 1105, an input unit 1106, and a power supply 1107. Among them, the processor 1101 is electrically connected to the touch display screen 1103, the radio frequency circuit 1104, the audio circuit 1105, the input unit 1106, and the power supply 1107 respectively. Those skilled in the art can understand that Figure 3 the structure of the fault prediction device shown in does not constitute a limitation on the fault prediction device, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0107] The touch display screen 1103 can be used to display a graphical user interface and receive operation instructions generated by a user acting on the graphical user interface. The touch display screen 1103 may include a display panel and a touch panel. Among them, the display panel can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of the fault prediction device. These graphical user interfaces can be composed of graphics, text, icons, videos, and any combination thereof. Optionally, the display panel can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. The touch panel can be used to collect touch operations of the user on or near it (such as operations of the user using a finger, a stylus, or any suitable object or accessory on or near the touch panel), and generate corresponding operation instructions, and the operation instructions execute the corresponding program. Optionally, the touch panel can include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch orientation of the user, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 1101, and can receive commands sent by the processor 1101 and execute them. The touch panel can cover the display panel. After the touch panel detects a touch operation on or near it, it transmits it to the processor 1101 to determine the type of touch event. Subsequently, the processor 1101 provides a corresponding visual output on the display panel according to the type of touch event. In the embodiments of the present invention, the touch panel and the display panel can be integrated into the touch display screen 1103 to implement input and output functions. However, in some embodiments, the touch panel and the touch panel can be implemented as two independent components to implement input and output functions. That is, the touch display screen 1103 can also be used as a part of the input unit 1106 to implement the input function.

[0108] The radio frequency circuit 1104 can be used to transmit and receive radio frequency signals to establish wireless communication with a network device or other fault prediction devices through wireless communication, and transmit and receive signals with the network device or other fault prediction devices.

[0109] The audio circuit 1105 can be used to provide an audio interface between the user and the fault prediction device through a speaker and a microphone. The audio circuit 1105 can transmit the electrical signal converted from the received audio data to the speaker, and the speaker converts it into a sound signal for output; on the other hand, the microphone converts the collected sound signal into an electrical signal, which is received by the audio circuit 1105 and then converted into audio data. After the audio data is output to the processor 1101 for processing, it is sent through the radio frequency circuit 1104 to, for example, another fault prediction device, or the audio data is output to the memory 1102 for further processing. The audio circuit 1105 may also include an earphone jack to provide communication between the peripheral earphone and the fault prediction device.

[0110] The input unit 1106 can be used to receive input digital, character information or user feature information (such as fingerprint, iris, face information, etc.), and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.

[0111] The power supply 1107 is used to supply power to each component of the fault prediction device 1100. Optionally, the power supply 1107 can be logically connected to the processor 1101 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 1107 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.

[0112] Although Figure 3 not shown in the figure, the fault prediction device 1100 may also include a camera, a sensor, a Wi-Fi unit, a Bluetooth unit, etc., which will not be elaborated here.

[0113] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not elaborated in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0114] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed through instructions, or through instructions to control relevant hardware. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0115] Therefore, an embodiment of the present invention provides a computer-readable storage medium, in which multiple computer programs are stored. The computer programs can be loaded by a processor to execute any fault prediction method provided by the embodiments of the present invention. The computer programs can execute the steps of the foregoing fault prediction method, and reference can be made to the previous embodiments, which will not be elaborated here.

[0116] Among them, the computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disc, etc.

[0117] Since the computer program stored in the computer-readable storage medium can execute any one of the fault prediction methods provided by the embodiments of the present invention, the beneficial effects achievable by any one of the fault prediction methods provided by the embodiments of the present invention can be realized. For details, refer to the previous embodiments and will not be elaborated here.

[0118] In the above embodiments of the fault prediction device, computer-readable storage medium, fault prediction equipment, and computer program product, the descriptions of each embodiment have their own emphases. For parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments. Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes and the beneficial effects brought by the above-described fault prediction device, computer-readable storage medium, computer program product, fault prediction equipment, and their corresponding units can refer to the description of the fault prediction method in the above embodiments and will not be elaborated here specifically.

[0119] The above has introduced in detail a fault prediction method, a fault prediction device, a fault prediction equipment, a computer-readable storage medium, and a computer program product provided by the embodiments of the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A fault prediction method, characterized in that: The method comprises: Acquire the current environmental data of the target computer room and the operating data of each device in the target computer room, and pre-process the environmental data and the operating data to obtain first characteristic data corresponding to the environmental data and second characteristic data corresponding to the operating data of each device; The first feature data and the second feature data are input into a preset fault prediction model to obtain a fault prediction result, wherein the fault prediction result includes a first prediction result, a second prediction result and a third prediction result, and the fault prediction model includes a global model, a device type model and a device individual model; the global model is configured to output the environmental data prediction result of the target computer room according to the first feature data, and perform fault prediction based on the environmental data prediction result to obtain the first prediction result; the device type model is configured to output the second prediction result corresponding to the target device type according to the second feature data corresponding to the target device type, and the target device type is preset; the device individual model is configured to output the third prediction result corresponding to the target device according to the second feature data corresponding to the target device, and the target device is preset.

2. The fault prediction method according to claim 1, characterized in that: The fault prediction model performs fault prediction based on multiple machine learning models; Before inputting the first characteristic data and the second characteristic data into a preset fault prediction model to obtain a fault prediction result, the method further includes: The first feature data and the second feature data are input into a pre-trained adaptive weighted model to obtain the weights of each machine learning model of the fault prediction model.

3. The fault prediction method according to claim 2, characterized in that: The training of the adaptive weighted model includes: Obtaining historical prediction results of the fault prediction model within a historical time interval, where the historical time interval is a time interval before the current moment; The accuracy of the historical prediction results is evaluated, and the objective function of the adaptive weighted model is optimized according to the historical prediction results and the corresponding accuracy.

4. The fault prediction method according to claim 1, characterized in that: The method further comprises: In response to a change in the equipment distribution or the environment of the target computer room, the parameters of the fault prediction model are updated according to an old data set and a new data set, wherein the old data set is a data set before the change occurs in the target computer room, and the new data set is a data set after the change occurs in the target computer room.

5. The fault prediction method according to claim 4, characterized in that: The updating of the parameters of the fault prediction model according to the old data set and the new data set includes: According to the old data set and the new data set, adjusting the parameters of the fault prediction model based on transfer learning; According to the new data set, the parameters of the fault prediction model after transfer learning are adjusted based on reinforcement learning.

6. The fault prediction method according to claim 1, characterized in that: The preprocessing of the environment data and the operation data to obtain first feature data corresponding to the environment data and second feature data corresponding to the operation data of each device includes: The environmental data and the operating data are respectively subjected to data cleaning, data standardization and feature extraction in sequence to obtain the first feature data and the second feature data.

7. The fault prediction method according to claim 1, characterized in that: The method further comprises: Performing a fault alarm according to the fault prediction result and generating fault analysis information; The environmental data, the operating data, the fault prediction results and the fault analysis information are displayed in real time.

8. A fault prediction device, characterized in that: The fault prediction device comprises: an acquisition unit, used to acquire the environmental data of the target computer room and the operating data of each device in the target computer room at the current moment, and pre-process the environmental data and the operating data to obtain first characteristic data corresponding to the environmental data and second characteristic data corresponding to the operating data of each device; A prediction unit is used to input the first feature data and the second feature data into a preset fault prediction model to obtain a fault prediction result, wherein the fault prediction result includes a first prediction result, a second prediction result and a third prediction result, and the fault prediction model includes a global model, a device type model and a device individual model; the global model is configured to output the environmental data prediction result of the target computer room according to the first feature data, and perform fault prediction based on the environmental data prediction result to obtain the first prediction result; the device type model is configured to output the second prediction result corresponding to the target device type according to the second feature data corresponding to the target device type, and the target device type is preset; the device individual model is configured to output the third prediction result corresponding to the target device according to the second feature data corresponding to the target device, and the target device is preset.

9. A fault prediction device, characterized in that: The invention comprises a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the fault prediction method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium comprises a computer program. When the computer program is run on an electronic device, the computer program is used to enable the electronic device to execute the steps of the fault prediction method according to any one of claims 1 to 7.