Fault prediction method, electronic equipment and storage medium

By obtaining ripple signals and operation data in real time in the server, using FPGA for data cleaning and feature extraction, and combining with random forest models for training, the problem of low accuracy in server power failure prediction in the existing technology is solved, more accurate and timely fault prediction is achieved, and the stability and operation and maintenance efficiency of the server system are improved.

CN120353327APending Publication Date: 2025-07-22INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510495043.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The method of judging server power failure based on threshold values in the prior art ignores the fluctuations in the operating environment, resulting in low accuracy of the fault prediction results and prone to false alarms or missed alarms.

Method used

By obtaining ripple signals and operation data in real time on the server, using FPGA for data cleaning and feature extraction, and combining with random forest models for training, a fault prediction model that is closer to the actual environment is generated, and failure prediction and alarm are carried out through DMA and I2C links and BMC for fault prediction and alarm.

Benefits of technology

It improves the accuracy and timeliness of fault prediction, reduces the false alarm rate, enhances the stability and operation and maintenance efficiency of the server system, and realizes early identification and rapid response to power failures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353327A_ABST
    Figure CN120353327A_ABST
Patent Text Reader

Abstract

The invention discloses a fault prediction method, electronic equipment and a storage medium, and the method comprises the steps: obtaining multi-dimensional power supply operation state information including a ripple signal and current operation data in real time, training a first prediction model which is read to the local of a target server, and obtaining a multi-dimensional power supply operation state information including the ripple signal and the current operation data; and the trained second prediction model is more in line with the actual operation environment of the server. And inputting the adjacent operation state information obtained in the next time interval into a local second prediction model of the server to predict the fault risk of the power supply of the server in the next time interval. According to the method, the technical problem of low accuracy caused by predicting the power supply fault of the server only by using the threshold value in the related technology is solved, and the technical effect of improving the accuracy of the fault prediction result is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and in particular, to a fault prediction method, an electronic device, and a storage medium. Background Art

[0002] In a server system, a power supply unit is one of the core components to ensure the stable operation of the system. To ensure the stability of this unit, in related technologies, data of sensors such as internal voltage, current, and power of the power supply are usually obtained, and a fault alarm is triggered when the operating parameters exceed a threshold.

[0003] However, the above method of simply judging power supply faults based on a defined threshold ignores the impact of fluctuations in the server operating environment on the operation of the server power supply. Therefore, situations such as false alarms or missed alarms of faults may occur, resulting in the technical problem of low accuracy of the fault prediction result of the server power supply. Summary of the Invention

[0004] This application provides a fault prediction method, an electronic device, and a storage medium to at least solve the problem of low accuracy caused by only using a threshold to predict server power supply faults in related technologies.

[0005] According to an aspect of an embodiment of this application, a fault prediction method is provided, including: when the target server is powered on, reading a first prediction model pre-stored in a baseboard management controller; obtaining current operating state information of the server power supply of the target server within a current time interval, where the current operating state information includes a current ripple signal and current operating data, and the current ripple signal is used to represent the operating state of a capacitor in the server power supply within the current time interval; training the first prediction model based on the current ripple signal and the current operating data to obtain a trained second prediction model; inputting adjacent operating state information of the server power supply within a next time interval into the second prediction model to obtain a fault prediction result, and uploading the fault prediction result to an operation and maintenance platform, where the adjacent operating state information includes an adjacent ripple signal and adjacent operating data, and the next time interval is a time interval adjacent to the current time interval.

[0006] According to another aspect of the embodiments of the present application, there is also provided a fault prediction device, including: a reading unit, configured to read a first prediction model pre-stored in a baseboard management controller when a target server is powered on; a first obtaining unit, configured to obtain current operating state information of a server power supply of the target server within a current time interval, where the current operating state information includes a current ripple signal and current operating data, and the current ripple signal is used to represent the operating state of a capacitor in the server power supply within the current time interval; a training unit, configured to train the first prediction model based on the current ripple signal and the current operating data to obtain a trained second prediction model; and a third processing unit, configured to input adjacent operating state information of the server power supply within a next time interval into the second prediction model to obtain a fault prediction result, and upload the fault prediction result to an operation and maintenance platform, where the adjacent operating state information includes an adjacent ripple signal and adjacent operating data, and the next time interval is a time interval adjacent to the current time interval.

[0007] According to yet another aspect of the embodiments of the present application, there is also provided an electronic device, including a memory and a processor, where a computer program is stored in the memory, and the processor is configured to execute the steps of any one of the above-mentioned fault prediction methods through the computer program.

[0008] According to yet another aspect of the embodiments of the present application, there is also provided a computer-readable storage medium, in which a computer program is stored, where the computer program is configured to execute the steps of any one of the above-mentioned fault prediction methods when running.

[0009] According to yet another aspect of the embodiments of the present application, there is provided a computer program product or a computer program, where the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps of any one of the above-mentioned fault prediction methods.

[0010] Through the above embodiments provided by the present application, by collecting and analyzing multi-dimensional power supply state information including ripple signals in real time, and combining with the first prediction model for local training and optimization of the server, a second prediction model closer to the actual operating environment of the server is obtained. The trained second prediction model can be used to identify power supply fault risks in advance, such as hidden problems like capacitor aging. Compared with traditional threshold alarms, the fault prediction results in the technical solution of the present application are more accurate and timely, solving the problems of false alarms or missed alarms of power supply faults caused by traditional fault prediction methods, achieving the technical effect of improving the accuracy of server power supply fault prediction results, and at the same time enhancing the stability and operation and maintenance efficiency of the server system. Description of the Drawings

[0011] To more clearly illustrate the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0012] Figure 1 It is a schematic diagram of an application scenario of a fault prediction method according to an embodiment of the present application.

[0013] Figure 2 It is a schematic flowchart of an optional fault prediction method according to an embodiment of the present application.

[0014] Figure 3 It is a schematic structural diagram of an optional server motherboard according to an embodiment of the present application.

[0015] Figure 4 It is a schematic diagram of an optional fault prediction result obtained based on a prediction model according to an embodiment of the present application.

[0016] Figure 5 It is a schematic overall flowchart of an optional fault prediction method according to an embodiment of the present application.

[0017] Figure 6 It is a structural block diagram of an optional fault prediction device according to an embodiment of the present application. Detailed implementation manners

[0018] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0019] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such a process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence.

[0020] To enable those skilled in the art of the present technology to better understand the solution of the present application, the following will further describe the present application in detail in conjunction with the accompanying drawings and specific implementation manners.

[0021] According to one aspect of the embodiments of the present application, a fault prediction method is provided. Optionally, in this embodiment, the above-mentioned fault prediction method can be but is not limited to being applied to a hardware scenario as shown in Figure 1 the following. The server device may include one or more (only one is shown in Figure 1 ) processors 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Among them, the above-mentioned server device may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown in the following is only schematic and does not limit the structure of the above-mentioned server device. For example, the server device may further include more or fewer components than those shown in Figure 1 the following, or have a different configuration from that shown in Figure 1 the following.

[0022] The memory 104 can be used to store computer programs. For example, software programs and modules of application software, such as the computer program corresponding to the fault prediction method in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, the above-mentioned method is implemented. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely set relative to the processor 102, and these remote memories can be connected to the server device through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0023] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by the communication provider of the server device. In one instance, the transmission device 106 includes a network adapter (abbreviated as NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0024] The embodiments of the present application can be, but are not limited to, applicable to a variety of scenarios that require high reliability and real-time fault warning, especially in the fields of data centers, cloud computing platforms, high-performance computing clusters, industrial automation, new energy vehicles, smart grids, and smart homes, to monitor the working status of devices in real time and predict potential faults. Specific examples of several application scenarios are given below:

[0025] (1) Data centers and cloud computing environments: Data centers usually have extremely high requirements for the stability and availability of servers. By deeply coordinating the FPGA (Field Programmable Gate Array) and BMC (Baseboard Management Controller) in the technical solution of the present application, the power status can be monitored in real time, power failures can be predicted in advance, and server downtime caused by power problems in the data center can be avoided, thereby ensuring the continuity of critical services and the security of data.

[0026] (2) Server maintenance in high-performance computing clusters: High-performance computing clusters are usually used to run intensive computing tasks, which cause great pressure on the power supply and cooling systems. The technology of the present application can capture high-frequency ripple changes in real time, predict potential problems such as capacitor aging and component failure, avoid computing interruptions and data loss caused by power failures, and at the same time, through a fast response mechanism, the cluster can immediately adjust the load to ensure the smooth progress of computing tasks.

[0027] (3) Industrial automation and intelligent manufacturing: Industrial automation equipment has strict requirements for operation stability and real-time performance. Through real-time fault prediction and fast interruption response, this solution can quickly isolate and handle sudden faults, avoid production line downtime, reduce production losses, and its dynamic adaptability enables it to cope with complex working conditions in industrial sites and improve the reliability of automation equipment.

[0028] (4) Network and communication infrastructure: Communication infrastructure (such as 5G base stations, network switches, etc.) requires high availability to ensure uninterrupted services. The technical solution of the present application can monitor the operation status of the server power supply in real time, discover potential fault risks, and ensure the continuous and stable operation of the equipment. In addition, by quickly responding to faults, data transmission interruptions or service quality degradation can be avoided, and the stability of the network and the user experience can be maintained.

[0029] The fault prediction method of the embodiments of the present application can be executed by a server device, or can be executed by at least one of a server device and a terminal device (which can also be understood as an input / output device 108). Among them, the execution of the fault prediction method of the embodiments of the present application by the terminal device can also be executed by a client installed thereon.

[0030] Taking the execution of the fault prediction method in this embodiment by the server as an example, Figure 2 It is a schematic flowchart of an optional fault prediction method according to an embodiment of the present application. As Figure 2 shown, the process of this method may include steps S202 to S208.

[0031] Step S202, when the target server is powered on, read the first prediction model pre-stored in the baseboard management controller.

[0032] Step S204, obtain the current operating state information of the server power supply of the target server within the current time interval, where the current operating state information includes the current ripple signal and the current operating data, and the current ripple signal is used to represent the operating state of the capacitor in the server power supply within the current time interval.

[0033] Step S206, based on the current ripple signal and the current operating data, train the first prediction model to obtain the trained second prediction model.

[0034] Step S208, by inputting the adjacent operating state information of the server power supply in the next time interval into the second prediction model, obtain the fault prediction result, and upload the fault prediction result to the operation and maintenance platform, where the adjacent operating state information includes the adjacent ripple signal and the adjacent operating data, and the next time interval is a time interval adjacent to the current time interval.

[0035] In a server system, the power supply unit (PSU) is one of the core components to ensure the stable operation of the system. The stable operation of the PSU is related to the stability of the entire server. Its design needs to balance power density, heat dissipation capacity and reliability, and through in-depth cooperation with subsystems such as BMC and VRM (Voltage Regulator Module), ensure the 7×24-hour uninterrupted operation of the data center.

[0036] Server power failure warning (which can also be understood as PSU failure warning) refers to monitoring multi-dimensional parameters such as voltage and current of the power module, identifying potential hazards such as capacitor aging, overload risk or component failure in advance, and triggering hierarchical alarms (such as log recording, redundant switching or active power-off protection), so as to avoid server downtime, data loss or hardware damage caused by power failure.

[0037] Among them, BMC is a small operating system independent of the server system and is a firmware system with an independent IP integrated on the motherboard. The main functions of BMC include remote access and control, hardware monitoring and alarming, remote virtual media, remote power control, and system health monitoring, etc. Through BMC, administrators can remotely access the server console through the network and perform various operations such as starting or shutting down the server, restarting and resetting the server, etc. At the same time, BMC can also monitor the hardware status of the server, such as CPU, memory, hard disk, power supply, etc., generate alarms or logs when problems or anomalies occur, and notify the administrator for handling and maintenance.

[0038] In the related art, BMC communicates with the server power supply based on the I2C link to obtain data of the internal voltage, current, power, and temperature sensors of the power supply, records the data exceeding the threshold, and triggers an alarm when abnormal data is detected multiple times, reminding the maintenance personnel that a server power supply failure has occurred and the device needs to be replaced.

[0039] However, the traditional fault warning method for server power supplies mainly makes a simple judgment based on thresholds and does not consider the impact of fluctuations in the server operating environment on the operation of the server power supply. Moreover, BMC can only monitor basic information such as voltage and current and cannot predict sudden failures caused by capacitor aging. Therefore, it is easy to cause the situation of missed fault reports.

[0040] At the same time, since the BMC polling time is usually 2 - 3 seconds, it is also impossible to isolate abnormal components in time, which easily leads to damage to the server motherboard and thus affects the stability of the server system.

[0041] To solve the above problems, a fault prediction method is proposed in the technical solution of this application. The corresponding hardware structure is as Figure 3 shown. A ripple detection chip (which can also be understood as a high - frequency ripple detection chip) is added to the server motherboard. Through this chip, high - frequency ripple data (which can also be understood as ripple signals) of the server power supply is obtained in real time. Among them, the ripple signal can reflect whether the capacitor working state of the server power supply is normal.

[0042] The server motherboard uses an FPGA chip to connect to the voltage sensor, current sensor, temperature sensor, power sensor, and high - frequency ripple detection chip of the server power supply, so as to obtain data such as power supply voltage, current, power, and ripple signals in real time.

[0043] Among them, FPGA is a programmable integrated circuit chip, which consists of a large number of programmable logic units and configurable interconnect resources. Users can implement various digital logic functions through programming. In this embodiment, the FPGA chip enables the DMA (Direct Memory Access) function to connect to the BMC, and the FPGA interrupt output pin is connected to the BMC interrupt input pin through the motherboard CPLD (Complex Programmable Logic Device).

[0044] The so-called DMA function is a technology that allows some hardware subsystems to directly read or write system memory independently of the central processor, which can improve data transfer efficiency, reduce the burden on the central processor, and thus enhance the overall performance of the system.

[0045] The so-called CPLD is an integrated circuit chip used to implement logic functions and control. It plays a key role in the server hardware system, and its main functions include system timing control, logic function implementation, hardware monitoring and management, fault detection and troubleshooting, etc.

[0046] In this embodiment, when the target server is powered on and booted, the FPGA reads the first prediction model (which can also be understood as the pre-trained model A) pre-stored in the BMC memory through the I2C link and copies the read first prediction model into the FPGA. At the same time, the current operating status information of the power supply of the target server within the current time interval is obtained through various sensors in the FPGA, such as status data such as temperature, voltage, current, power consumption, etc., and the Figure 3 shown ripple detection chip is used to obtain the ripple signal of each power supply in real time and transmit the ripple signal to the FPGA (which can also be understood as the first processing unit).

[0047] The first prediction model in the FPGA is trained using the obtained current ripple signal and current operating data to obtain the trained second prediction model, and the first prediction model in the FPGA is replaced with the second prediction model.

[0048] When the adjacent operating status information in the next time interval is obtained, the second prediction model is used to output the fault prediction result and upload the fault prediction result to the operation and maintenance platform of the data center so that maintenance personnel can take preventive measures in a timely manner.

[0049] It should be noted that a server may be configured with one or more server power supplies. Therefore, the number of server power supplies in this embodiment is not limited.

[0050] Obviously, the FPGA is a control component on the server motherboard. By utilizing the current operating status information within the current time interval, the secondary training of the local model can be performed locally on the target server to optimize the model parameters to adapt to the operating environment of the current server, and a more accurate second prediction model can be obtained.

[0051] The reason for uploading the pre-trained model (the first prediction model) to the BMC and then reading it into the FPGA on the server local and performing model training in the FPGA is that the BMC, as an important part of the server, is usually designed for remotely managing and monitoring the operating status of the server, and its computing power and memory resources are relatively limited. While the FPGA is a programmable integrated circuit, especially good at parallel processing and high-speed data throughput. Performing model training and inference in the FPGA can utilize its hardware concurrency advantage, significantly reducing the training time and inference latency, which is crucial for fault prediction with high real-time requirements and rapid response.

[0052] It can be seen that in the technical solution of this application, a dedicated high-frequency ripple detection chip is used to capture the PSU capacitor characteristics, and combined with the time-frequency domain data joint analysis of data such as voltage, current, and temperature, a core data feature set for fault prediction is constructed, that is, a cooperative monitoring mechanism of high-frequency ripple and multi-parameters is used to construct the data basis for fault prediction.

[0053] Secondly, the FPGA, as a computing module, realizes real-time data acquisition, cleaning, and model inference, while the BMC is responsible for alarm decision-making and remote communication. The two achieve high-speed cooperation through DMA and terminal signals, breaking through the traditional I2C bandwidth limitation, and at the same time not affecting the BMC utilization rate, ensuring the operation efficiency of other functions of the BMC.

[0054] Furthermore, the technical solution of this application will perform secondary training on the model based on the operating environment of the server to adapt to the working conditions of the server power supply in different environments, reducing the occurrence of false alarm problems caused by the general model and improving the dynamic adaptation ability of the fault prediction method.

[0055] By adopting the above method, through real-time collection and analysis of multi-dimensional power supply status information including ripple signals, and combined with the first prediction model for training and optimization on the server local, a second prediction model closer to the actual operating environment of the server is obtained. Using the trained second prediction model can identify the power supply fault risk in advance, such as hidden problems like capacitor aging. Compared with the traditional threshold alarm, the fault prediction result in the technical solution of this application is more accurate and more timely, solving the problem of false alarms or missed alarms of power supply faults caused by traditional fault prediction methods, achieving the technical effect of improving the accuracy of the server power supply fault prediction result, and at the same time enhancing the stability and operation and maintenance efficiency of the server system.

[0056] In an exemplary embodiment, obtaining the current operating status information of the server power supply of the target server within the current time interval includes: obtaining the current ripple signal of the server power supply within the current time interval through a ripple detection unit on the server motherboard, where the ripple detection unit is connected to a first processing unit on the server motherboard; connecting a set of status detection sensors in the server power supply through the first processing unit to obtain the current operating data.

[0057] In this embodiment, the acquisition mechanism of the ripple signal and operating data in the server power supply fault prediction method is further refined. Among them, a ripple detection unit (such as Figure 3 the shown ripple detection signal) is used to capture the high-frequency ripple signal in the PSU output, which is a hardware component connected to the first processing unit (i.e., FPGA) on the server motherboard to ensure that the high-frequency ripple signal can be transmitted to the FPGA for analysis in real time and with high fidelity. The presence of the ripple detection unit enables the system to obtain key information reflecting the working state of the capacitor, providing a richer data source for fault prediction.

[0058] Secondly, a set of sensors in the FPGA are distributed at key positions of the server power supply to be used for real-time monitoring of basic operating data such as voltage, current, temperature, and power. The output of the status detection sensors is directly connected to the first processing unit, i.e., FPGA, enabling the FPGA to synchronously collect and process these diverse operating parameters, further enriching the information basis for fault prediction.

[0059] When the target server is running within the current time interval, the ripple detection unit continuously monitors the ripple signal output by the PSU to capture the fluctuations and changes in the working state of the capacitor. At the same time, the status detection sensors real-time monitor the basic operating state of the power supply and collect real-time data of voltage, current, temperature, and power. These data reflect the working state of the PSU from different perspectives.

[0060] After the above data is transmitted to the FPGA, the FPGA uses this data and combines it with the first prediction model read from the BMC for local optimization and training to generate a second prediction model that is more suitable for the current server operating environment. The second prediction model not only includes the analysis of basic operating parameters, but more importantly, integrates the high-frequency ripple signal, which is a key indicator reflecting capacitor aging and internal component status, making the fault prediction more comprehensive and accurate.

[0061] By adopting the above method, through enhancing the comprehensiveness of data acquisition and comprehensively using basic operating parameters and high-frequency ripple signals, the trained second prediction model can be trained and predicted based on a more comprehensive and sensitive data source, reducing misjudgment and missed judgment and improving the prediction accuracy.

[0062] In an exemplary embodiment, the above-mentioned first prediction model is trained based on the current ripple signal and the current operating data to obtain the trained second prediction model, including: reading the first prediction model in the baseboard management controller into the first processing unit on the server mainboard; performing data cleaning on the current ripple signal and the current operating data through the first processing unit to obtain a cleaned ripple signal and cleaned operating data; performing feature extraction on the cleaned ripple signal and the cleaned operating data to obtain current operating state features; based on the current operating state features, training the first prediction model to obtain the second prediction model.

[0063] In the first processing unit (FPGA), a data cleaning algorithm is used to process data such as voltage (12V / 5V / 3.3V), current (0-100A), temperature (-40℃-125℃), power (0-1200W) and ripple signals, eliminating the influence of transmission error data or noise data, and using a random forest model for fault prediction. At the same time, the FPGA can capture the instantaneous fault of the PSU and directly send it to the BMC through the interrupt signal through the motherboard CPLD. After receiving the interrupt, the BMC immediately takes measures such as power limitation and isolation for the faulty PSU. The following will describe the specific real-time processing process of the instantaneous fault.

[0064] The main process of training the first prediction model is as follows Figure 4 As shown, including:

[0065] S11, data cleaning: perform mean-variance joint filtering on the collected power supply parameters and perform standardization processing at the same time;

[0066] In the data cleaning algorithm embedded in FPGA, mean-variance joint filtering is a key data preprocessing technology used to smooth signals and remove high-frequency noise, while identifying and removing outliers. The specific process is as follows:

[0067] (1) Real-time sampling and windowing: The FPGA collects ripple signals and operating data from the PSU in real time, such as voltage, current, temperature, and power. In order to apply the filtering algorithm, the data is segmented and placed into a moving window. The window size can be flexibly adjusted according to the signal characteristics. For example, it can be set to a data sample set within a few milliseconds to tens of milliseconds.

[0068] (2) Calculate mean and variance: For each set of data in the window, the FPGA calculates its mean and variance in real time. The mean reflects the central trend of the data, while the variance describes the degree of dispersion of the data distribution.

[0069] (3) Outlier detection and elimination: Using the calculated mean and variance, a reasonable range for the data can be defined. For example, if a data point is more than 3 times the variance (standard deviation) away from the mean, it is considered an outlier, which may be caused by measurement errors, circuit interference, or other abnormal phenomena. The FPGA eliminates the identified outliers or replaces them with the mean of other data within the window to maintain the stability of the data set;

[0070] (4) Dynamic update and filtering: As new data samples continuously enter the window and old data gradually moves out, the mean and variance need to be updated dynamically. In this way, the filter can continuously adapt to the changes in the signal and provide a smooth and denoised data stream.

[0071] (5) Standardization processing: Standardization processing is another important link in the data cleaning algorithm, aiming to eliminate the influence of data dimensions and make all input features have the same scale. This is crucial for using machine learning algorithms because if the magnitude differences between features are too large, it may affect the convergence speed and prediction accuracy of the model.

[0072] S12, Feature extraction: Extract features related to power supply failures, such as voltage fluctuation amplitude, current mutation frequency, ripple frequency, etc.;

[0073] S13, Model training: The FPGA is connected to the BMC through I2C. The pre-trained model A is uploaded to the memory of the BMC by an external network in advance. When the system starts running, the FPGA reads the pre-trained model A from the BMC and uses the PSU-related feature data extracted from this server to train the model A (the first prediction model) based on random forest to output a second prediction model, that is, the power supply failure prediction model, which is more in line with the current operating environment of the PSU of this server.

[0074] It should be noted that when the first prediction model is being trained in the FPGA and the training is not completed, the first prediction model in the BMC is used to predict the failure of the service power supply within the current time interval. Although the prediction result of the first prediction model is lower than that of the second prediction model, since the two models have the same structure but different structural parameters, the prediction results can still meet the expected requirements.

[0075] Through the in-depth data cleaning and feature extraction implemented by the FPGA, the accuracy and real-time performance of the server power supply failure prediction are significantly enhanced. The cleaned ripple signal and operation data are purer and more consistent, and as model inputs, they can avoid noise interference and improve the reliability of the prediction model.

[0076] The feature extraction process accurately captures key indicators such as voltage, current fluctuations, and ripple frequency, providing highly relevant data for model training, making the fault warning more sensitive and enabling early identification of hidden problems such as capacitor aging.

[0077] As another alternative implementation, when the server environment changes, secondary training is initiated to update the model, ensuring its real-time effectiveness and improving the accuracy of the prediction results. This process can be implemented in the following specific steps:

[0078] S21, Environment monitoring and change identification;

[0079] First, the FPGA continuously monitors the operating environment of the server, including but not limited to key parameters such as environmental temperature, humidity, power grid fluctuations, CPU load, and memory usage. When it detects that the environmental parameters exceed the preset normal range or observes significant changes, the FPGA records these changes and identifies that the environment has changed.

[0080] S22, Data collection and preprocessing;

[0081] After identifying the environmental change, the FPGA starts to collect operation data related to the new environment, including power supply voltage, current, temperature, and high-frequency ripple, etc. While collecting data, the FPGA applies data cleaning algorithms to remove transmission error data or noisy data, ensuring the quality and accuracy of the data. In addition, the data is standardized to ensure the consistency of the data format and numerical range, facilitating model training.

[0082] S23, Model update and secondary training;

[0083] The FPGA transfers the preprocessed data to the BMC through DMA (Direct Memory Access), and the BMC initiates the secondary training process. During this process, the BMC uses the second prediction model saved in the memory as the basis and combines the new data in the current server environment to train the second prediction model. The purpose is to better adapt to the current environmental changes and ensure the accuracy of the prediction results.

[0084] S24, Model verification and replacement;

[0085] After training the third prediction model, the BMC verifies the new model to check whether its prediction performance in the current environment is better than the old model (the second prediction model). After passing the verification, the second prediction model in the FPGA is replaced with the third prediction model, and at the same time, the third prediction model is sent back to the BMC to achieve model update and synchronization.

[0086] S25, Prediction and feedback mechanism;

[0087] The updated third prediction model runs in the FPGA to predict the real-time collected data. The prediction results are sent to the BMC via the I2C link, and the BMC generates corresponding warning information or fault handling instructions based on the prediction results. At the same time, the BMC also collects feedback information from the model prediction to evaluate the long-term performance of model C and adjust the frequency and strategy of model updates as needed.

[0088] S26, operation and maintenance notification and maintenance.

[0089] If the third prediction model predicts a potential power failure, the BMC will send warning information to the operation and maintenance management terminal through the network module, including the fault type, fault location, and fault prediction probability. Based on the information received, the operation and maintenance personnel can make maintenance plans in advance, such as replacing the PSU with a risk of failure.

[0090] Through the above steps, when the server environment changes, the deep collaboration between FPGA and BMC can start the secondary training and update of the model, which can effectively improve the real-time and accuracy of power failure prediction, reduce the false alarm rate caused by environmental changes, and provide more intelligent and reliable guarantees for the stable operation and efficient operation and maintenance of the data center. This dynamic model update mechanism adapts to the complex operating environment and working condition changes of the data center, and reflects the advancement and practicality of the server power failure early warning system.

[0091] In an exemplary embodiment, the method further includes: replacing the first prediction model in the first processing unit with a second prediction model, and transmitting the second prediction model back to the baseboard management controller through the target communication link.

[0092] This embodiment is about the model update and synchronization mechanism. First, the FPGA reads the first prediction model from the BMC, locally trains and optimizes the first prediction model based on the data in the current server operating environment, and determines the model at the end of the training as the second prediction model.

[0093] The first prediction model in the first processing unit is replaced by the second prediction model, and at the same time, Figure 3 The I2C link shown transmits the second prediction model back to the BMC, thereby ensuring the consistency of the models in the first processing unit and the BMC system, thereby ensuring that the entire system is always in an optimized state and has the latest and best prediction capabilities.

[0094] That is to say, by transmitting the trained second prediction model back to the BMC, not only is the trained second prediction model deployed in the FPGA for real-time prediction, but it is also transmitted back to the BMC through the target communication link (such as I2C). As the center of server cluster management and remote access control, the BMC can store and manage multiple optimized models to call the latest and most suitable model version when needed.

[0095] In this embodiment, replacing the first prediction model with the second prediction model based on the actual operation data of the server ensures that the fault prediction model can continuously adapt to environmental changes and improve the prediction accuracy; efficiently transmitting the second prediction model through the target communication link and using the DMA mechanism to avoid frequent CPU interruptions not only guarantees the speed of model update but also reduces the occupancy of network resources.

[0096] In addition, the updated second prediction model is directly deployed on the FPGA side, reducing the latency of model invocation and making the fault prediction and response mechanism more efficient and real-time.

[0097] Through local optimization, deployment, and cluster-level transmission of the model, the continuous evolution and intelligent management of the server power fault prediction system are realized, significantly improving the operation and maintenance efficiency of the data center and the optimization of the use of hardware resources, providing a solid technical guarantee for the long-term stable operation of the server.

[0098] In an exemplary embodiment, after obtaining the current operation status information of the server power supply of the target server within the current time interval, the above method further includes: obtaining the current prediction result of the server power supply within the current time interval by inputting the current ripple signal and the current operation data into the first prediction model in the first processing unit, where the first prediction model in the first processing unit is read from the baseboard management controller; sending the current prediction result to the baseboard management controller, and reporting the current prediction result to the operation and maintenance platform through the baseboard management controller.

[0099] Before the second prediction model is trained, the first prediction model in the BMC is first used for effective fault prediction, and its prediction result is timely fed back to the operation and maintenance platform to form an effective closed-loop of fault warning. In this process, the FPGA plays a key role in data processing and model inference, while the BMC is responsible for the final alarm decision and remote communication.

[0100] The reason for using the pre-trained model (the first prediction model) for fault prediction before the training is completed is that the training process takes a certain amount of time, and the server cannot wait for it to run. Therefore, the FPGA first analyzes the ripple signal and operation data collected in real time, and uses the first prediction model to perform fault prediction based on the processed data. Finally, the fault prediction result is transmitted to the BMC through the I2C link. This process ensures that basic fault warning services can be provided even in the initial stage of system startup.

[0101] During the reporting process of the fault prediction result, the prediction result obtained based on the first prediction model is reported to the BMC through the DMA mechanism, and the BMC is responsible for further analysis and alarm decision-making. The BMC determines whether to report the prediction result to the operation and maintenance platform according to the preset rules and strategies. This process ensures the accuracy and timeliness of the alarm information.

[0102] That is to say, using the cooperation mechanism between the BMC and the FPGA to build a dynamic and hierarchical fault prediction and alarm system, specifically including:

[0103] (1) Fault prediction in the initial stage: In the initial stage of server startup, the first prediction model read from the BMC is used to perform preliminary fault prediction on the ripple signal and operation data collected in real time. Although the second prediction model is not trained yet at this time, the first prediction model can already provide certain prediction capabilities, ensuring the safety of the system startup stage.

[0104] (2) Model transition and result feedback: As the FPGA obtains more data, the second prediction model is gradually trained. During this period, the prediction results based on the first prediction model will be continuously sent to the BMC. Once the second prediction model meets the expected training requirements and prediction accuracy, the first prediction model in the FPGA is automatically replaced with the second prediction model, and the second prediction model is used for fault prediction. Finally, the prediction results are updated and sent to the BMC.

[0105] (3) Alarm decision-making and reporting of the BMC: Regardless of whether the FPGA uses the first prediction model or the second prediction model for prediction, the BMC will decide whether to report the current prediction result to the operation and maintenance platform based on the received prediction results and its own preset alarm strategy. This mechanism ensures that remote communication is only triggered when the fault warning is truly necessary, saving network resources and improving the accuracy of the alarm at the same time.

[0106] In this embodiment, by introducing a warning module to analyze the fault prediction results, if the prediction result indicates that the power supply is about to fail, communication is carried out with the BMC via I2C. The BMC is responsible for generating corresponding warning information and sending it to the operation and maintenance management terminal through the network module of the BMC. The warning information includes the fault type, fault location, fault probability, etc.; if a sudden fault occurs in the PSU, the FPGA notifies the BMC through an interrupt, and the interrupt handling function of the BMC performs operations such as current limiting or isolation on the PSU and sends an alarm.

[0107] By flexibly applying the pre-trained model A (the first prediction model) and the dynamically optimized model B (the second prediction model), a highly intelligent fault warning system is formed. This system can provide basic warning services at the initial stage of server startup and gradually improve the accuracy and response efficiency of warnings with the local training of the model, providing strong support for the efficient operation and maintenance of the data center and the optimal configuration of hardware resources.

[0108] In an exemplary embodiment, after obtaining the current operating status information of the server power supply of the target server within the current time interval, the above method further includes: in the case where the amplitude change of the current ripple signal exceeds the first threshold, determining that a first instantaneous fault occurs and generating a first interrupt signal at the same time; in the case where the current power supply temperature in the current operating data increases to the second threshold within a preset duration, determining that a second instantaneous fault occurs and generating a second interrupt signal at the same time; in the case where the current voltage value in the current operating data increases to the third threshold within a preset duration, determining that a third instantaneous fault occurs and generating a third interrupt signal at the same time; in the case where the current current value in the current operating data increases to the fourth threshold within a preset duration, determining that a fourth instantaneous fault occurs and generating a fourth interrupt signal at the same time; reporting the first interrupt signal, the second interrupt signal, the third interrupt signal, and the fourth interrupt signal to the baseboard management controller; transmitting the first interrupt signal, the second interrupt signal, the third interrupt signal, and the fourth interrupt signal to the baseboard management controller to perform corresponding fault handling operations on each component in the server power supply.

[0109] This embodiment deeply elaborates on how the FPGA can quickly respond to and capture the instantaneous faults of the server power supply in addition to data cleaning and model prediction. By generating and reporting interrupt signals to the BMC, an immediate fault handling mechanism is triggered. This process emphasizes the FPGA's ability to monitor sudden faults, not only limited to predicting potential long-term risks through models, but also being able to respond immediately to sudden abnormal situations, such as sharp changes in voltage and current, or abnormal increases in power supply temperature.

[0110] First, a variety of different instantaneous fault situations and corresponding interrupt signal generation mechanisms are defined. For example, through a special high-frequency ripple detection chip, the FPGA can obtain the high-frequency noise level in the PSU output voltage. Abnormal changes in the ripple signal, especially a sudden increase in the ripple amplitude, may indicate that the power supply capacitor may be aging or faulty. At this time, the first instantaneous fault is determined and the first interrupt signal is generated.

[0111] For another example, the FPGA is connected to a temperature sensor deployed inside the server power supply to monitor the power supply temperature information in real time. When the inside of the PSU overheats, it may be a sign of thermal management failure or component failure. Sustained high temperature may damage electronic components, leading to a decline or even failure of the server power supply performance. Therefore, when it is detected that the power supply temperature rapidly increases to the second threshold within a preset duration, the FPGA determines the second instantaneous fault (such as a failure of the cooling system) and then generates the second interrupt signal.

[0112] For another example, the FPGA is connected to the voltage sensor of the PSU and continuously monitors the output voltage. If the voltage suddenly fluctuates and is lower than the set minimum threshold or higher than the maximum threshold, it may indicate the occurrence of an instantaneous fault. For example, a power supply short circuit or overload. The FPGA can immediately identify the third instantaneous fault and automatically generate the third interrupt signal.

[0113] For another example, by connecting the FPGA to a current sensor deployed inside the server power supply, the FPGA can monitor the power supply current flow in real time. A sudden surge in current, especially a current peak that exceeds the normal value range, is usually a sign of overheating or short circuit of the components inside the PSU. Then, the fourth instantaneous fault is determined and the fourth interrupt signal is generated.

[0114] Through the first processing unit, the first interrupt signal, the second interrupt signal, the third interrupt signal, and the fourth interrupt signal are reported to the BMC. The interrupt mechanism ensures the immediate transmission of fault information and overcomes the delay problem in the traditional polling mode.

[0115] Through the capture of instantaneous faults by the first processing unit and the addition of the interrupt mechanism, the millisecond-level response to instantaneous faults is ensured, effectively preventing system damage caused by delays. By immediately identifying and processing instantaneous faults, the affected PSU components can be quickly isolated, preventing the fault from spreading to the entire server or cluster and reducing the operation risk of the data center.

[0116] In an exemplary embodiment, transmitting the first interrupt signal, the second interrupt signal, the third interrupt signal, and the fourth interrupt signal to the baseboard management controller includes: transmitting the first interrupt signal, the second interrupt signal, the third interrupt signal, and the fourth interrupt signal to a second processing unit on the server motherboard through the interrupt output pins of the first processing unit, where the second processing unit is connected to the first processing unit; performing signal conversion on the first interrupt signal, the second interrupt signal, the third interrupt signal, and the fourth interrupt signal through the second processing unit to obtain a converted interrupt signal; transmitting the converted interrupt signal to the interrupt input pin of the baseboard management controller and triggering an interrupt handling program.

[0117] This embodiment focuses on how, after the FPGA detects an instantaneous fault and generates an interrupt signal, signal conversion and transparent transmission are performed through a second processing unit (usually a CPLD, complex programmable logic device) on the server motherboard to ensure that the interrupt signal can be accurately received by the BMC and trigger an immediate fault handling program. The CPLD acts as a signal bridge here, ensuring that different types and forms of interrupt signals can be smoothly transmitted to the BMC without being hindered by hardware compatibility or signal matching.

[0118] Specifically, the FPGA sends the interrupt signals generated by four identified instantaneous faults (ripple amplitude mutation, power supply temperature surge, abnormal voltage rise, sharp increase in current) to the CPLD through its built-in interrupt output pins. As a signal converter, the CPLD performs necessary format conversion and level matching on the received interrupt signals to adapt to the interrupt input pins of the BMC, ensuring the integrity and effectiveness of the signals.

[0119] Transmit the converted interrupt signal to the interrupt input pin of the BMC. After receiving the interrupt signal, the BMC immediately responds, calls a preset interrupt handling program, and performs immediate fault handling operations on the corresponding components in the server power supply, such as current limiting, power-off protection, logging, or notifying the operation and maintenance personnel.

[0120] Send the interrupt signal to the CPLD on the server motherboard through the interrupt output pin of the FPGA. As a signal conversion device, the CPLD performs necessary format conversion and level matching on the interrupt signal. The signal conversion function of the CPLD eliminates the hardware interface differences between the FPGA and the BMC, ensures the accurate transmission of the interrupt signal, and increases the stability and reliability of the system.

[0121] Through the signal transparent transmission of the CPLD, the BMC can directly receive the interrupt signal generated by the FPGA without additional signal parsing or conversion steps, simplifies the fault handling process, and improves the operation and maintenance efficiency.

[0122] Through the signal conversion and transparent transmission mechanism of the CPLD, an efficient and highly compatible instantaneous fault warning and processing system is constructed. This system not only improves the real-time performance and accuracy of the server power supply fault warning, but also optimizes the configuration of hardware resources, providing technical support for the stable operation and efficient operation and maintenance of the data center.

[0123] In an exemplary embodiment, the above-mentioned obtaining the fault prediction result by inputting the adjacent operating state information of the server power supply in the next time interval into the second prediction model includes: obtaining the fault prediction result by inputting the adjacent operating state information into the second prediction model, where the fault prediction result includes at least one fault label or at least one fault prediction probability of the server power supply having a power supply fault; sending at least one fault label or at least one fault prediction probability to the baseboard management controller through the target communication link, and generating corresponding alarm prompt information.

[0124] Based on the descriptions of the above embodiments, this embodiment further explains how to deeply analyze the operating state of the server power supply by using the second prediction model, so as to obtain specific fault prediction results. Among them, this embodiment not only includes single-type fault warning, but also can predict multiple types of faults.

[0125] When using the second prediction model for fault prediction, the output fault prediction result includes, but is not limited to, the label or occurrence probability of the possibility of each fault in at least one fault. For example, "(1, 0, 2)" means that the occurrence probability of current overload (fault 1) is medium, 0 means that the occurrence probability of capacitor aging (fault 2) is low, and 2 means that the occurrence probability of the heat dissipation system failure (fault 3) is high.

[0126] Or the fault prediction result may include the probability of occurrence of each fault in at least one fault. For example, (0.8, 0.6, 0.9) means that the occurrence probability of fault 1 is 0.8, the occurrence probability of fault 2 is 0.6, and the occurrence probability of fault 3 is 0.9.

[0127] After outputting the above fault prediction result, the fault prediction result can be sent to the BMC through the target communication link, but is not limited to this. After receiving the prediction result containing multiple fault labels or probabilities, the BMC generates corresponding alarm prompt information according to the preset risk level and strategy, and reports it to the operation and maintenance platform through the network module.

[0128] Compared with simple fault prediction, this embodiment can provide more detailed fault prediction results, including the probabilities of different types of faults, enabling the operation and maintenance personnel to have a clear understanding of potential multiple fault types and facilitating the formulation of comprehensive prevention and response measures.

[0129] The probability encoding in the fault prediction results allows the operation and maintenance platform to quantitatively evaluate the risk level of each fault type. For example, for a high-probability fault 3, the operation and maintenance personnel may give priority to preventive maintenance, while for a low-probability fault 2, they can observe it temporarily to save resources.

[0130] If the BMC receives results containing multiple fault prediction information, it can immediately generate alarm prompts according to the severity and urgency of each fault, and reasonably allocate maintenance resources according to the specific type and expected occurrence time of the fault to optimize the operation efficiency of the data center.

[0131] In addition, the early warning mechanism for multiple fault tags or probabilities in this embodiment supports the concept of early fault prediction and proactive maintenance. Based on the prediction results, the operation and maintenance personnel can intervene in advance to avoid service interruption and data loss caused by faults.

[0132] By introducing a second prediction model for multi-dimensional fault probability prediction and through refined alarm information encoding, a more intelligent and comprehensive server power supply fault early warning system is constructed. This system can not only provide accurate risk assessment, but also support rapid response and resource optimization, laying a solid foundation for the stable operation of the data center and the improvement of operation and maintenance efficiency. Through this innovation, not only the information dimension of fault early warning is enriched, but also the fault management of the data center is brought with foresight and initiative, promoting the development of data center operation and maintenance towards the direction of intelligence and refinement.

[0133] After the FPGA reports the fault prediction results to the BMC, the BMC will take a series of processing operations according to the preset strategy to protect the server hardware and data security. The following are examples of the processing operations that the BMC may take in different situations:

[0134] (1) Power current overload

[0135] Alarm notification: Through the network module, the BMC will immediately send alarm information to the operation and maintenance management platform, including the fault type, location, and abnormal value of the current, to remind the operation and maintenance personnel to pay attention.

[0136] Immediate response: The BMC may activate the current limiting function to immediately reduce the current input to the server and avoid hardware damage that may be caused by overcurrent.

[0137] Long-term adjustment: If the current overload situation persists, the BMC may adjust the working mode of the server, such as reducing the CPU load, to reduce the current demand until the problem is solved.

[0138] Notify maintenance: At the same time, the BMC will notify the maintenance personnel to check the power supply unit (PSU) and related hardware to find and eliminate the cause of overcurrent, which may be hardware failure, overload, or power quality problems.

[0139] (2) Conditions of excessive power consumption

[0140] The BMC will evaluate the overall status of the server to determine whether the high power consumption is caused by an increase in task load or a hardware failure. Based on the evaluation results, the BMC may intelligently schedule the server resources, such as reducing power consumption through the Dynamic Voltage and Frequency Scaling (DVFS) technology.

[0141] If it is due to an overloaded task load, the BMC may offload some tasks to other servers to balance the power consumption. High power consumption is usually accompanied by high temperature, and the BMC will activate or strengthen the cooling system, such as increasing the fan speed, to prevent thermal damage.

[0142] In addition, the BMC will report the high power consumption situation to the operation and maintenance management platform and recommend a hardware check to avoid potential power supply or cooling system failures.

[0143] (3) Abnormal changes in ripple signals

[0144] The BMC will analyze the abnormal changes in the ripple signal to determine whether they are caused by capacitor aging, power fluctuations, or other factors. If the ripple signal is abnormal and may damage the server, the BMC will immediately take measures, such as active power-off protection, to avoid damage to the server hardware caused by power surges or ripples.

[0145] The BMC can isolate the problem power module to prevent the abnormal ripple signal from affecting the entire server system.

[0146] Based on the analysis results of the ripple signal, the BMC may predict the health status of the power module and notify the maintenance personnel in advance for preventive replacement or maintenance to avoid emergency shutdowns when faults occur.

[0147] Hardware diagnosis: The BMC may start a hardware diagnostic program to check the power-related hardware, such as capacitors, voltage regulators, etc., to determine the direct cause of the abnormal ripple.

[0148] When the BMC detects abnormalities in the current, power consumption, or ripple signal of the server power supply, it will take a series of processing operations from immediate response to long-term maintenance to ensure the stable operation of the server and the safety of the hardware. These operations are not limited to recording and alarming, but also include intelligent resource scheduling, adjustment of working modes, activation of cooling systems, and isolation of problem components, reflecting the core role and intelligent processing ability of the BMC in fault management.

[0149] To more clearly understand the above technical solution, the following further describes the fault prediction method in combination with Figure 5 the overall flowchart shown.

[0150] S502, the server boots and runs;

[0151] S504, the FPGA obtains the pre-trained model A from the BMC;

[0152] Among them, the FPGA is the first processing unit, and the pre-trained model A can also be understood as the first prediction model. The FPGA reads the pre-trained model A stored in the BMC memory through the I2C link.

[0153] S506, the FPGA obtains the PSU operation status information;

[0154] Among them, the obtained PSU operation status information includes but is not limited to the ripple signals of each PSU detected by the ripple detection unit (which can also be understood as the ripple detection unit) and the data of each sensor deployed in each PSU. For example, the temperature of PSU0 monitored by the temperature sensor, the voltage of PSU1 monitored by the voltage sensor, etc.

[0155] After obtaining the operation status information of each PSU, a data cleaning algorithm is used in the FPGA to process data such as voltage, temperature, power, and ripple signals. For example, the data is denoised through a mean-variance joint filter to obtain the processed data. Then the processed data is used to execute the following steps S508, step S514, and step S520.

[0156] S508, the FPGA can capture the instantaneous faults of the PSU;

[0157] S510, the captured instantaneous faults are transparently transmitted to the BMC through the main board CPLD via the interrupt signal;

[0158] S512, after receiving the interrupt signal, the BMC immediately takes corresponding measures for the faulty PSU and reports it to the operation and maintenance platform at the same time;

[0159] For example, measures such as power limitation and isolation are taken.

[0160] S514, before the model B is not trained using the processed data, first use the model A obtained in the FPGA for fault prediction;

[0161] Although the accuracy of the prediction result of model A is slightly lower than that of model B, it can still meet the basic fault prediction requirements.

[0162] S516, the result predicted by using model A is reported to the BMC through the DMA function;

[0163] S518, the result after analyzing the fault prediction result is reported to the operation and maintenance platform through the BMC;

[0164] S520, perform filtering and normalization on the processed data;

[0165] S524, use the normalized data as training data to train Model A, and obtain the trained Model B;

[0166] S526, replace Model A in the FPGA with Model B, and at the same time transmit Model B back to the BMC.

[0167] The purpose is to ensure the consistency of the models on the server local and the BMC side.

[0168] Adopting the technical solution provided by this application has at least the following beneficial effects:

[0169] (1) By combining the high-frequency ripple detection chip with the machine learning model (random forest model), potential fault features such as capacitor aging and component failure can be accurately captured. Compared with the traditional threshold warning method, it can predict power supply failures (such as sudden failures and overload risks) in advance, reduce the risk of server downtime caused by power supply problems, and improve the fault prediction ability.

[0170] (2) The collaborative design of the PGA and the BMC realizes a hierarchical processing mechanism: the FPGA triggers the BMC to quickly isolate sudden faults (such as instantaneous overcurrent and surges) within milliseconds through the interrupt signal. Compared with the traditional I2C polling mechanism (2 - 3s delay), the response speed is increased by a hundred times, effectively avoiding motherboard damage, and improving the real-time performance and response efficiency of fault response.

[0171] (3) Dynamic model adaptability

[0172] Adopting a hybrid training mode of the pre-trained Model A and the local Model B enables the fault prediction model to be dynamically adjusted according to the actual operating environment of the server (such as power grid fluctuations and heat dissipation conditions), reducing the false alarm rate and improving the environmental adaptability, especially suitable for complex working conditions in data centers.

[0173] (4) The FPGA embeds data cleaning and feature extraction algorithms to directly preprocess the original signal, reducing the computing power burden on the BMC; the DMA transmission mechanism avoids frequent CPU interrupts and reduces the low occupancy rate of hardware resources.

[0174] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.

[0175] According to another aspect of the embodiments of the present application, a fault prediction device is further provided. This fault prediction device can be used to implement the fault prediction method provided in the above embodiments, and those already described will not be repeated. As used hereinafter, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0176] Figure 6 is a structural block diagram of an optional fault prediction device according to an embodiment of the present application. As Figure 6 shown in the figure, the fault prediction device includes a reading unit 602, a first acquisition unit 604, a training unit 606, and a third processing unit 608.

[0177] The reading unit 602 is configured to read a first prediction model pre-stored in a baseboard management controller when the target server is powered on.

[0178] The first acquisition unit 604 is configured to acquire current operating state information of the server power supply of the target server within the current time interval. The current operating state information includes a current ripple signal and current operating data, and the current ripple signal is used to represent the operating state of the capacitor in the server power supply within the current time interval.

[0179] The training unit 606 is configured to train the first prediction model based on the current ripple signal and the current operating data to obtain a trained second prediction model.

[0180] The third processing unit 608 is configured to input adjacent operating state information of the server power supply in the next time interval into the second prediction model to obtain a fault prediction result, and upload the fault prediction result to an operation and maintenance platform, where the adjacent operating state information includes an adjacent ripple signal and adjacent operating data, and the next time interval is a time interval adjacent to the current time interval.

[0181] It should be noted that the reading unit 602 in this embodiment can be used to execute the above step S202, the first acquisition unit in this embodiment can be used to execute the above step S204, the training unit 606 in this embodiment can be used to execute the above step S206, and the third processing unit 608 is used to execute the above step S208.

[0182] By collecting and analyzing multi-dimensional power supply status information including ripple signals in real time, and combining with the first prediction model for local training and optimization on the server, a second prediction model closer to the actual operating environment of the server is obtained. The trained second prediction model can identify power failure risks in advance, such as hidden problems like capacitor aging. Compared with traditional threshold alarms, the fault prediction results in the technical solution of this application are more accurate and timely, solving the problems of false alarms or missed alarms of power failures caused by traditional fault prediction methods, achieving the technical effect of improving the accuracy of server power failure prediction results, and at the same time enhancing the stability and operation and maintenance efficiency of the server system.

[0183] In an exemplary embodiment, the above-mentioned first acquisition unit 604 includes:

[0184] The first acquisition module is used to acquire the current ripple signal of the server power supply within the current time interval through the ripple detection unit on the server motherboard, where the ripple detection unit is connected to the first processing unit on the server motherboard;

[0185] The second acquisition module is used to acquire the current operation data by connecting a group of status detection sensors in the server power supply through the first processing unit.

[0186] In an exemplary embodiment, the above-mentioned training unit 606 includes:

[0187] The reading module is used to read the first prediction model in the baseboard management controller into the first processing unit on the server motherboard;

[0188] The first processing module is used to perform data cleaning on the current ripple signal and the current operation data through the first processing unit to obtain the cleaned ripple signal and the cleaned operation data;

[0189] The feature extraction module is used to extract features from the cleaned ripple signal and the cleaned operation data to obtain the current operation status features;

[0190] The first training module is used to train the first prediction model based on the current operation status features to obtain the second prediction model.

[0191] In an exemplary embodiment, the above-mentioned device further includes:

[0192] The fourth processing unit is used to replace the first prediction model in the first processing unit with the second prediction model and send the second prediction model back to the baseboard management controller through the target communication link.

[0193] In an exemplary embodiment, the above-mentioned device further includes:

[0194] A fifth processing unit, configured to, after obtaining the current operating status information of the server power supply of the target server within the current time interval, obtain the current prediction result of the server power supply within the current time interval by inputting the current ripple signal and the current operating data into the first prediction model in the first processing unit, where the first prediction model in the first processing unit is read from the baseboard management controller;

[0195] A sixth processing unit, configured to send the current prediction result to the baseboard management controller, and report the current prediction result to the operation and maintenance platform through the baseboard management controller.

[0196] In an exemplary embodiment, the above device further includes:

[0197] A seventh processing unit, configured to, after obtaining the current operating status information of the server power supply of the target server within the current time interval, determine that a first instantaneous fault occurs and generate a first interrupt signal when the amplitude change of the current ripple signal exceeds a first threshold;

[0198] An eighth processing unit, configured to determine that a second instantaneous fault occurs and generate a second interrupt signal when the current power supply temperature in the current operating data increases to a second threshold within a preset duration;

[0199] A ninth processing unit, configured to determine that a third instantaneous fault occurs and generate a third interrupt signal when the current voltage value in the current operating data increases to a third threshold within a preset duration;

[0200] A tenth processing unit, configured to determine that a fourth instantaneous fault occurs and generate a fourth interrupt signal when the current current value in the current operating data increases to a fourth threshold within a preset duration;

[0201] An eleventh processing unit, configured to report the first interrupt signal, the second interrupt signal, the third interrupt signal, and the fourth interrupt signal to the baseboard management controller;

[0202] A twelfth processing unit, configured to transmit the first interrupt signal, the second interrupt signal, the third interrupt signal, and the fourth interrupt signal to the baseboard management controller to perform corresponding fault handling operations on each component in the server power supply.

[0203] In an exemplary embodiment, the above twelfth processing unit includes:

[0204] A first transmission module, configured to transmit the first interrupt signal, the second interrupt signal, the third interrupt signal, and the fourth interrupt signal to the second processing unit on the server motherboard through the interrupt output pin of the first processing unit, where the second processing unit is connected to the first processing unit;

[0205] A conversion module, configured to perform signal conversion on a first interrupt signal, a second interrupt signal, a third interrupt signal, and a fourth interrupt signal through a second processing unit to obtain a converted interrupt signal;

[0206] An input module, configured to transmit the converted interrupt signal to an interrupt input pin of a baseboard management controller and trigger an interrupt handling program.

[0207] In an exemplary embodiment, the above-mentioned third processing unit 608 includes

[0208] A second processing module, configured to obtain a fault prediction result by inputting adjacent operation state information into a second prediction model, where the fault prediction result includes at least one fault label or at least one fault prediction probability of a power failure of a server power supply;

[0209] A sending module, configured to send at least one fault label or at least one fault prediction probability to a baseboard management controller through a target communication link and generate a corresponding alarm prompt message.

[0210] It should be noted that the above-mentioned respective modules can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited to this: the above-mentioned modules are all located in the same processor; or, the above-mentioned respective modules are located in different processors in any combination form.

[0211] According to another aspect of the embodiments of the present application, an electronic device is further provided, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned embodiments of the fault prediction method.

[0212] According to another aspect of the embodiments of the present application, a computer-readable storage medium is further provided. A computer program is stored in the computer-readable storage medium, where the computer program is configured to execute the steps in any of the above-mentioned embodiments of the fault prediction method when running.

[0213] In an exemplary embodiment, the above-mentioned computer-readable storage medium may include, but is not limited to: a USB flash drive, a read-only memory (ROM for short), a random access memory (RAM for short), a mobile hard disk, a magnetic disk, or an optical disc and other various media that can store a computer program.

[0214] According to another aspect of the embodiments of the present application, a computer program product is further provided. The above-mentioned computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-mentioned embodiments of the fault prediction method.

[0215] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in any of the above-described embodiments of the fault prediction method are implemented.

[0216] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0217] The above has introduced in detail a fault prediction method provided by the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

Claims

1. A fault prediction method, characterized in that, It includes: When the target server is powered on, read the first prediction model pre-stored in the baseboard management controller; Obtain the current operating state information of the server power supply of the target server within the current time interval, wherein the current operating state information includes a current ripple signal and current operating data, and the current ripple signal is used to represent the operating state of the capacitor in the server power supply within the current time interval; Based on the current ripple signal and the current operating data, train the first prediction model to obtain a trained second prediction model; By inputting the adjacent operating state information of the server power supply within the next time interval into the second prediction model, obtain a fault prediction result, and upload the fault prediction result to the operation and maintenance platform, wherein the adjacent operating state information includes an adjacent ripple signal and adjacent operating data, and the next time interval is a time interval adjacent to the current time interval.

2. The method according to claim 1, characterized in that, The obtaining the current operating state information of the server power supply of the target server within the current time interval includes: Obtain the current ripple signal of the server power supply within the current time interval through a ripple detection unit on the server motherboard, wherein the ripple detection unit is connected to a first processing unit on the server motherboard; Connect a group of state detection sensors in the server power supply through the first processing unit to obtain the current operating data.

3. The method according to claim 1, characterized in that, The training the first prediction model based on the current ripple signal and the current operating data to obtain a trained second prediction model includes: Read the first prediction model in the baseboard management controller into the first processing unit on the server motherboard; Through the first processing unit, perform data cleaning on the current ripple signal and the current operating data to obtain a cleaned ripple signal and cleaned operating data; Extract features from the cleaned ripple signal and the cleaned operating data to obtain current operating state features; Based on the current operating state features, train the first prediction model to obtain the second prediction model.

4. The method according to claim 3, characterized in that, The method further includes: Replace the first prediction model in the first processing unit with the second prediction model, and transmit the second prediction model back to the baseboard management controller through a target communication link.

5. The method according to any one of claims 1 to 4, characterized in that, After obtaining the current operating state information of the server power supply of the target server within the current time interval, the method further includes: By inputting the current ripple signal and the current operation data into the first prediction model in the first processing unit, the current prediction result of the server power supply within the current time interval is obtained, where the first prediction model in the first processing unit is read from the baseboard management controller; Send the current prediction result to the baseboard management controller, and report the current prediction result to the operation and maintenance platform through the baseboard management controller.

6. The method according to any one of claims 1 to 4, wherein After obtaining the current operation status information of the server power supply of the target server within the current time interval, the method further includes: When the amplitude change of the current ripple signal exceeds a first threshold, it is determined that a first instantaneous fault occurs, and a first interrupt signal is generated at the same time; When the current power supply temperature in the current operation data increases to a second threshold within a preset duration, it is determined that a second instantaneous fault occurs, and a second interrupt signal is generated at the same time; When the current voltage value in the current operation data increases to a third threshold within a preset duration, it is determined that a third instantaneous fault occurs, and a third interrupt signal is generated at the same time; When the current current value in the current operation data increases to a fourth threshold within a preset duration, it is determined that a fourth instantaneous fault occurs, and a fourth interrupt signal is generated at the same time; Report the first interrupt signal, the second interrupt signal, the third interrupt signal, and the fourth interrupt signal to the baseboard management controller; Transmit the first interrupt signal, the second interrupt signal, the third interrupt signal, and the fourth interrupt signal to the baseboard management controller to perform corresponding fault handling operations on each component in the server power supply.

7. The method according to claim 6, wherein The transmitting the first interrupt signal, the second interrupt signal, the third interrupt signal, and the fourth interrupt signal to the baseboard management controller includes: Transmit the first interrupt signal, the second interrupt signal, the third interrupt signal, and the fourth interrupt signal to a second processing unit on the server motherboard through the interrupt output pin of the first processing unit, where the second processing unit is connected to the first processing unit; Perform signal conversion on the first interrupt signal, the second interrupt signal, the third interrupt signal, and the fourth interrupt signal through the second processing unit to obtain the converted interrupt signal; Transmit the converted interrupt signal to the interrupt input pin of the baseboard management controller and trigger an interrupt handling program.

8. The method according to claim 1, wherein The obtaining the fault prediction result by inputting the adjacent operation status information of the server power supply in the next time interval into the second prediction model includes: By inputting the adjacent operating state information into the second prediction model, the fault prediction result is obtained, where the fault prediction result includes at least one fault label or at least one fault prediction probability of a power failure of the server power supply; Via a target communication link, the at least one fault label or the at least one fault prediction probability is sent to the baseboard management controller, and corresponding alarm prompt information is generated.

9. An electronic device, characterized in that, Comprising: A memory for storing a computer program; A processor for implementing the steps of the fault prediction method according to any one of claims 1 to 8 when executing the computer program.

10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, where the computer program implements the steps of the fault prediction method according to any one of claims 1 to 8 when executed by a processor.

11. A computer program product, comprising a computer program, characterized in that, The computer program implements the steps of the fault prediction method according to any one of claims 1 to 8 when executed by a processor.

Citation Information

Cited By

  • Equipment fault prediction method and device, electronic equipment and storage medium

    CN121166493A