Fault prediction system and method of server and storage medium

By introducing a fault prediction system in the baseboard management controller that allows the first and second processing cores to work together, the problem of excessive processing core load is solved, efficient server fault prediction is achieved, and the business continuity of the server is ensured.

CN121008992APending Publication Date: 2025-11-25SHANDONG YUNHAI GUOCHUANG CLOUD COMPUTING EQUIP IND INNOVATION CENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511393434.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-26
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

The processing core of the baseboard management controller is complex in handling various component indicators, communication, and fault prediction model algorithms, resulting in low execution efficiency, inability to detect and handle server hardware faults in a timely manner, and impacting business continuity.

Method used

The system employs a collaborative approach between a first processing core and a second processing core. The first processing core acquires and stores the operational data of the server components, while the second processing core performs feature extraction and fault prediction. Shared memory is used for data transmission and processing, reducing the burden on the processing cores.

Benefits of technology

This improves the efficiency of the baseboard management controller's processing core in executing other management and monitoring tasks, preventing server hardware failures caused by processing core anomalies from going undetected and unhandled in a timely manner, thus ensuring business continuity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121008992A_ABST
    Figure CN121008992A_ABST
Patent Text Reader

Abstract

The invention discloses a server fault prediction system and method and a storage medium, and relates to the technical field of servers, the system comprises a substrate management controller, the substrate management controller comprises a first processing core, a second processing core and a shared memory, and the shared memory is in communication connection with the first processing core and the second processing core; the first processing core is used for acquiring operation data of a plurality of components included in the server and storing the operation data into the shared memory; the second processing core is used for acquiring the operation data from the shared memory and performing feature extraction processing on the operation data to obtain operation feature data; and based on the operation characteristic data, predicting the fault of the server to obtain a fault prediction result of each component of the server. According to the method and the device, the problems that the execution efficiency of the processing core of the substrate management controller on other management monitoring tasks is low, and server hardware faults cannot be found and processed in time in related technologies can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of servers, and particularly relates to a server fault prediction system and method and a storage medium. BACKGROUND

[0002] With the development of cloud computing and big data technology, servers are increasingly widely applied in various fields. In modern data centers, hardware faults of servers can cause business interruption, data loss and increased operation and maintenance costs. Therefore, it is very important to monitor faults of various components of servers.

[0003] In related technologies, a baseboard management controller (BMC) is usually used to monitor faults of various components of servers, mainly relying on passive monitoring, monitoring indexes of various components, such as processor usage, memory occupancy, disk input / output performance indexes, network traffic, or analyzing related logs of servers, such as system logs, application logs and error logs, and relying on fixed threshold alarms or preset models to predict faults of servers. In this way, the processing core of the baseboard management controller needs to simultaneously process the collection, communication and processing functions of the indexes of various components, and the fault prediction model algorithm is complex, which reduces the execution efficiency of the processing core of the baseboard management controller on other management and monitoring tasks, causes the baseboard management controller to be abnormal, and further causes server hardware faults to be unable to be discovered and processed in time, finally causing the server to be shut down and affecting the continuity of business. SUMMARY

[0004] The present application provides a server fault prediction system, method and storage medium to at least solve the problem that the execution efficiency of the processing core of the baseboard management controller on other management and monitoring tasks is low and server hardware faults cannot be discovered and processed in time in related technologies.

[0005] The present application provides a server fault prediction system, which comprises a baseboard management controller including a first processing core, a second processing core and a shared memory, and the shared memory is in communication connection with the first processing core and the second processing core. The first processing core is configured to acquire running data of a plurality of components included in the server and store the running data in the shared memory.

[0006] The second processing core is configured to acquire the running data from the shared memory, perform feature extraction processing on the running data to obtain running feature data, and predict faults of the server based on the running feature data to obtain fault prediction results of each component of the server.

[0007] The application provides a server fault prediction method, which is applied to a first processing core of a server fault prediction system. The server fault prediction system comprises a first processing core, a second processing core and a shared memory included in a baseboard management controller. The shared memory is in communication connection with the first processing core and the second processing core. The method comprises the following steps: obtaining running data of a plurality of components included in the server and storing the running data in the shared memory. The running data is used for the second processing core to predict the fault of the server based on the running data, obtain a fault prediction result of each component of the server, and store the fault prediction result in the shared memory. The fault prediction result is obtained from the shared memory and sent to a client.

[0008] The application provides a server fault prediction method, which is applied to a second processing core of a server fault prediction system. The server fault prediction system comprises a first processing core, a second processing core and a shared memory included in a baseboard management controller. The shared memory is in communication connection with the first processing core and the second processing core. The method comprises the following steps: obtaining running data from the shared memory. The running data is subjected to feature extraction processing to obtain running feature data. The running feature data is used to indicate the trend of the running state of each component changing with time during the running process. The fault of the server is predicted based on the running feature data to obtain a fault prediction result of each component of the server.

[0009] The application further provides a server fault prediction device, which is applied to a first processing core of a server fault prediction system. The server fault prediction system comprises a first processing core, a second processing core and a shared memory included in a baseboard management controller. The shared memory is in communication connection with the first processing core and the second processing core. The server fault prediction device comprises a first obtaining module, which is used for obtaining running data of a plurality of components included in the server and storing the running data in the shared memory. The running data is used for the second processing core to predict the fault of the server based on the running data, obtain a fault prediction result of each component of the server, and store the fault prediction result in the shared memory. The fault prediction result is obtained from the shared memory and sent to a client.

[0010] The application further provides a server fault prediction device, which is applied to the second processing core of the server fault prediction system.

[0011] The application further provides an electronic device, which comprises a memory for storing a computer program and a processor for executing the computer program to implement the steps of the server fault prediction method.

[0012] The application further provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the steps of the server fault prediction method.

[0013] The application further provides a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement the steps of the server fault prediction method.

[0014] According to the application, the processing core of the baseboard management controller does not need to simultaneously acquire the running data of the indicators of each component, and the running feature data is used to predict the server fault and obtain the fault prediction result of each component of the server, so that the influence of the processing core of the baseboard management controller on the execution efficiency of other management and monitoring tasks can be reduced, the abnormality of the baseboard management controller can be avoided, the server hardware fault caused by the abnormality of the baseboard management controller can be discovered and processed in time, and the problem that the server is shut down and the continuity of the business is affected can be solved. BRIEF DESCRIPTION OF DRAWINGS

[0015] In order to more clearly illustrate the embodiments of the application, the drawings needed in the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative effort.

[0016] Figure 1 A topology structure diagram of a server fault prediction system is provided for the embodiments of the application.

[0017] Figure 2A start process schematic diagram of a system of the first processing core and the second processing core provided for the embodiment of the present application;

[0018] Figure 3 A flow schematic diagram of a server fault prediction method provided for the embodiment of the present application;

[0019] Figure 4 A flow schematic diagram of another server fault prediction method provided for the embodiment of the present application;

[0020] Figure 5 A flow schematic diagram of another server fault prediction method provided for the embodiment of the present application;

[0021] Figure 6 A device structure block diagram of a server fault prediction device provided for the embodiment of the present application;

[0022] Figure 7 A device structure block diagram of another server fault prediction device provided for the embodiment of the present application;

[0023] Figure 8 A hardware structure schematic diagram of an electronic device provided for the embodiment of the present application. DETAILED DESCRIPTION

[0024] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the protection scope of the present application.

[0025] It should be noted that, in the description of the present application, the terms "comprise", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. The terms "first", "second" and the like in the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence.

[0026] In order for those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the drawings and specific embodiments.

[0027] The embodiments of the present application are applied to the scenario of baseboard management controller (BMC) predicting server faults.

[0028] As an independent management subsystem, the Baseboard Management Controller (BMC) can collect data from various sensors. In related technologies, BMC fault prediction primarily relies on monitoring indicators and log analysis, depending on fixed threshold alarms, and cannot predict progressive faults. Alternatively, the BMC runs a machine learning model on its processing core. This involves introducing machine learning algorithms, training offline based on historical data collected by the BMC, and importing the trained model into the BMC firmware for fault prediction using the BMC's processing core. In this approach, the BMC's processing core needs to simultaneously handle the acquisition, communication, and processing of various component indicators. Furthermore, the fault prediction model algorithm is complex, reducing the efficiency of the BMC's processing core in executing other management and monitoring tasks. This can lead to BMC malfunctions, resulting in the inability to detect and handle server hardware faults in a timely manner, ultimately causing server downtime and impacting business continuity.

[0029] To address the aforementioned technical problems, this application provides a server fault prediction system. The system includes a baseboard management controller comprising a first processing core, a second processing core, and shared memory. The shared memory is communicatively connected to the first and second processing cores. The first processing core acquires operational data from multiple components of the server and stores this data in the shared memory. The second processing core acquires the operational data from the shared memory and performs feature extraction processing on the data to obtain operational feature data. This operational feature data indicates the trend of the operational status of each component changing over time. Based on the operational feature data, server faults are predicted to obtain fault prediction results for each component of the server.

[0030] In this way, the first and second processing cores in the baseboard management controller each perform their own operations. This means that the processing cores of the baseboard management controller do not need to simultaneously handle the acquisition, communication, and processing functions of various component indicators, as well as the fault prediction model algorithm. This improves the efficiency of the baseboard management controller's processing cores in performing other management and monitoring tasks, and avoids the problem that server hardware failures caused by abnormalities in the baseboard management controller cannot be detected and handled in a timely manner, which could ultimately lead to server downtime and affect business continuity.

[0031] This application provides a server fault prediction system, such as... Figure 1 As shown, Figure 1 This is a topology diagram of a server fault prediction system provided in an embodiment of this application. Figure 1 In the system, the server fault prediction system 100 includes a baseboard management controller 101 and a client 102.

[0032] The substrate management controller 101 can be a board-level management controller, which is usually integrated in servers, network devices, industrial control systems, etc. The BMC is responsible for managing and monitoring hardware resources and providing remote management functions. The BMC includes a first processing core, a second processing core, and a shared memory.

[0033] The shared memory can be a shared memory of the first processing core and the second processing core in the BMC. The shared memory is in communication connection with the first processing core and the second processing core.

[0034] The Linux system runs on the first processing core, and the Real-Time Operating System (RTOS) runs on the second processing core. After the Linux system is started, it is used to run basic BMC service programs; after the RTOS system is started, it is used to train a preset prediction model and predict the failure of the server.

[0035] As shown in Figure 2 , Figure 2 the first processing core and the second processing core each system provided by the embodiment of the present application are shown in the starting process diagram; in Figure 2 , specifically includes: powering on the first processing core; the first processing core runs a specified program in the Boot Read-Only Memory (BootRom), loads the RTOS system to start; the first processing core runs a specified program in the BootRom, loads the RTOS system to start; the RTOS system calls a Secondary Program Loader (SPL) during the starting process; the SPL stage guides the Universal BootLoader (Uboot) to start; the uboot stage loads the Linux kernel and starts the BMC service program.

[0036] The client 102 can be any device with communication and display functions. For example, the client 102 can be a client of the substrate management controller, which is used to display the failure prediction result or prompt information of the server.

[0037] Based on the above-mentioned server failure prediction system, the first processing core is configured to acquire running data of a plurality of components included in the server, and store the running data into the shared memory.

[0038] The second processing core is configured to acquire the running data from the shared memory, and perform feature extraction processing on the running data to obtain running feature data; based on the running feature data, predict the failure of the server to obtain the failure prediction result of each component of the server.

[0039] The server may include multiple components such as processor, memory, disk, baseboard management controller, graphics card, and independent disk redundant array card.

[0040] The operational data of the multiple components comprising a server can be the operational status data of each component during operation. For example, the operational data of multiple components could include processor resource utilization, memory usage, disk I / O operations per second, bandwidth rate, network rate, disk utilization, etc. Furthermore, the BMC can monitor data from various sensors; the operational data can also include the temperature, power consumption, memory error count, hard drive SMART health status, and power supply voltage fluctuations monitored by various sensors (e.g., temperature sensors, voltage sensors, fan speed sensors) for each component. The operational data of the multiple components comprising a server includes the hardware data of each of the multiple components.

[0041] Operational characteristic data is used to indicate the trend of the operating status of each component changing over time during operation. Operational characteristic data includes dynamic time-series characteristics, discrete event density characteristics, and periodic fluctuation characteristics.

[0042] Dynamic time-series characteristics are used to indicate the operational data characteristics of each component in a continuous time series.

[0043] Discrete event density features are used to indicate the distribution density of abnormal events in each component over time.

[0044] Periodic fluctuation characteristics are used to indicate the characteristics of individual components that repeat over a fixed time period.

[0045] Specifically, the second processing core retrieves runtime data from shared memory and performs feature extraction processing on the runtime data to obtain runtime feature data, including: for the retrieved runtime data, using the sliding window averaging method to eliminate instantaneous noise (such as voltage spikes), and using interpolation algorithms to complete missing data caused by sensor failures; calculating time-series indicators such as temperature change rate, power consumption fluctuation standard deviation, and fan speed trend slope to obtain dynamic time-series features; statistically analyzing discrete event densities such as the hourly frequency of memory error correction codes and hard disk read / write error rates to obtain discrete event density features; and eliminating high-frequency noise interference through low-pass filtering, or using Fourier transform to extract periodic fluctuation features.

[0046] Specifically, the second processing core is used to input dynamic time series features, discrete event density features, and periodic fluctuation features into a preset prediction model and output the server's fault prediction results.

[0047] The preset prediction model can be a prediction model built based on historical operational feature data.

[0048] In one example, the first processing core is further configured to acquire historical operating data of multiple components included in the server and store the historical operating data in shared memory. The second processing core is further configured to acquire historical operating data from shared memory, perform feature extraction processing on the historical operating data to obtain historical operating feature data; and train a model on the historical operating feature data based on time series prediction algorithms and classification algorithms to obtain a preset prediction model.

[0049] Among them, the time series prediction algorithm includes the Long Short-Term Memory Network algorithm and the Autoregressive Integral Moving Average algorithm. The Long Short-Term Memory Network algorithm is used to capture the long-term dependencies between various features in the historical operating characteristic data, while the Autoregressive Integral Moving Average algorithm is used to analyze the periodic fluctuations of various features in the historical operating characteristic data.

[0050] Classification algorithms are used for fault identification based on historical operational feature data and alarm thresholds. These algorithms can be random forests or support vector machines. Random forests are used for fault type discrimination based on multi-dimensional feature fusion (e.g., hard drive failure / fan failure), while support vector machines are suitable for anomaly detection in scenarios with small sample sizes.

[0051] Historical operational data can be partitioned into training, validation, and test sets according to a preset ratio. For example, a preset ratio of 7:2:1 can be used. This partitioning method helps avoid overfitting. Optionally, the model's decision-making logic can be explained using SHapley Additive explanatory values ​​(SHAPs). The preset prediction model supports incremental training and dynamic model updates to adapt to environmental changes such as hardware aging.

[0052] Furthermore, the second processing core is also used to store the fault prediction results in shared memory; the first processing core is also used to retrieve the fault prediction results from the shared memory and send the fault prediction results to the client.

[0053] Optionally, the first processing core is also used to send heartbeat information to the second processing core; when the first processing core does not receive the response information based on the heartbeat information from the second processing core within a preset time period, it is determined that the thread corresponding to the preset prediction model on the second processing core is deadlocked, triggering the second processing core to perform a restart operation and sending an exception message to the client. The exception message is used to indicate that the baseboard management controller cannot predict the server failure.

[0054] Understandably, the heartbeat information monitoring mechanism can promptly detect deadlocks in the preset prediction model thread on the second processing core, preventing the fault prediction function from failing for a long time due to thread deadlock and ensuring the continuous availability of the server's fault prediction capability.

[0055] Then, the first processing core is also used to compare the hardware data of each component with the corresponding fault threshold to determine the hardware fault result of each component; compare the fault prediction result of each component with the hardware fault result; if the fault prediction result is inconsistent with the hardware fault result, the fault prediction result is determined to be a false alarm event; based on the false alarm event, the alarm threshold corresponding to the preset prediction model is adjusted.

[0056] Specifically, the first processing core is used to obtain the number of consecutive false alarm events within the first time period; if the number of consecutive false alarm events is greater than or equal to a preset threshold, the alarm threshold is increased or decreased.

[0057] Understandably, by comparing fault prediction results with actual hardware fault results, false alarms can be accurately identified, and alarm thresholds can be dynamically adjusted based on the false alarm situation. This allows the prediction model to continuously adapt to actual operating scenarios and reduce invalid alarms. Thresholds are only adjusted when consecutive false alarms reach a preset threshold, avoiding frequent parameter changes caused by occasional false alarms, while ensuring timely correction when systemic deviations exist, thus balancing stability and adaptability. Reducing false alarms significantly reduces unnecessary maintenance responses, preventing maintenance personnel from being distracted by invalid information and allowing them to focus on genuine fault handling, thereby improving maintenance efficiency. Furthermore, dynamically adjusting thresholds avoids misjudgments caused by fixed thresholds; continuously optimizing the prediction model adapts to different hardware environments; and supporting preventative maintenance can prevent sudden server downtime.

[0058] Through the aforementioned server fault prediction system, the first processing core of the baseboard management controller can acquire the operating data of each component of the server and store the operating data in shared memory; the second processing core can retrieve the operating data from the shared memory and perform feature extraction processing on the operating data to obtain operating feature data. Based on the operating feature data, server faults are predicted to obtain fault prediction results for each component of the server.

[0059] Since the processing core of the baseboard management controller does not need to simultaneously acquire the operating data of various component indicators and predict server failures based on operating characteristic data, thus obtaining the failure prediction results of each component of the server, the impact of the processing core of the baseboard management controller on the execution efficiency of other management and monitoring tasks can be reduced. This avoids the problem that server hardware failures caused by abnormalities in the baseboard management controller cannot be detected and handled in a timely manner, ultimately leading to server downtime and affecting business continuity.

[0060] Embodiments of this application provide a server fault prediction method, which can be applied to the aforementioned server fault prediction system. The server fault prediction system includes a baseboard management controller comprising a first processing core, a second processing core, and shared memory. The shared memory is communicatively connected to the first and second processing cores.Figure 3 As shown, Figure 3 A flowchart illustrating a server fault prediction method provided in this application embodiment; the specific processing steps of the server fault prediction method may include:

[0061] S301, the first processing core, is used to acquire the operating data of multiple components included in the server and store the operating data in shared memory.

[0062] S302, the second processing core, is used to obtain running data from shared memory and perform feature extraction processing on the running data to obtain running feature data. The running feature data is used to indicate the trend of the running status of each component changing over time during operation. Based on the running feature data, the failure of the server is predicted to obtain the failure prediction result of each component of the server.

[0063] The specific implementation process of S301 and S302 can be referred to the execution process in the above-mentioned server fault prediction system, and will not be elaborated here.

[0064] Embodiments of this application provide yet another method for predicting server failures, which can be applied to the first processing core of the aforementioned server failure prediction system, such as... Figure 4 As shown, Figure 4 A flowchart illustrating another server fault prediction method provided in this application embodiment; the specific processing steps of the server fault prediction method may include:

[0065] S401 retrieves the operating data of multiple components included in the server and stores the operating data in shared memory.

[0066] The second processing core uses the running data to predict server faults, obtains fault prediction results for each component of the server, and stores the fault prediction results in shared memory.

[0067] S402 retrieves the fault prediction results from shared memory and sends the fault prediction results to the client.

[0068] In some optional implementations, the first processing core is also used to compare the hardware data of each component with the corresponding fault threshold to determine the hardware fault result of each component; compare the fault prediction result of each component with the hardware fault result; if the fault prediction result is inconsistent with the hardware fault result, the fault prediction result is determined to be a false alarm event; and based on the false alarm event, adjust the alarm threshold corresponding to the preset prediction model.

[0069] In one example, the first processing core obtains the number of consecutive false alarm events within a first time period; if the number of consecutive false alarm events is greater than or equal to a preset threshold, the alarm threshold is increased or decreased.

[0070] In some alternative implementations, the first processing core acquires historical operational data from multiple components of the server and stores the historical operational data in shared memory.

[0071] Embodiments of this application provide another method for server fault prediction, which can be applied to the second processing core of the aforementioned server fault prediction system. The second processing core is communicatively connected to the first processing core of the baseboard management controller and shared memory. Figure 5 As shown, Figure 5 A flowchart illustrating another server fault prediction method provided in this application embodiment; the specific processing steps of the server fault prediction method may include:

[0072] S501 retrieves runtime data from shared memory.

[0073] S502 performs feature extraction processing on the running data to obtain running feature data.

[0074] Among them, the operational characteristic data is used to indicate the trend of the operational status of each component changing over time during operation.

[0075] S503 predicts server failures based on operational characteristic data, obtaining failure prediction results for each component of the server.

[0076] In some optional implementations, the second processing core can input dynamic time-series features, discrete event density features, and periodic fluctuation features into a preset prediction model and output the server's fault prediction results.

[0077] In some optional implementations, the second processing core can obtain historical running data from shared memory and perform feature extraction processing on the historical running data to obtain historical running feature data; based on time series prediction algorithms and classification algorithms, the historical running feature data is used to train a model to obtain a preset prediction model.

[0078] In some alternative implementations, the second processing core can store the fault prediction results in shared memory.

[0079] For a description of the features in the embodiments corresponding to the above-mentioned server fault prediction method, please refer to the relevant descriptions in the embodiments corresponding to the server fault prediction system, which will not be repeated here.

[0080] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0081] Embodiments of this application also provide a server fault prediction device, applied to the first processing core of the aforementioned server fault prediction system; such as Figure 6 As shown, Figure 6 A structural block diagram of a server fault prediction device provided in this application embodiment; the server fault prediction device includes:

[0082] The first acquisition module 601 is used to acquire the operating data of multiple components included in the server and store the operating data in shared memory. The operating data is used by the second processing core to predict the failure of the server based on the operating data, obtain the failure prediction result of each component of the server, and store the failure prediction result in shared memory; retrieve the failure prediction result from shared memory and send the failure prediction result to the client.

[0083] The server fault prediction device also includes a first processing module 602.

[0084] In some optional implementations, the first processing module 602 is used to compare the hardware data of each component with the corresponding fault threshold to determine the hardware fault result of each component; compare the fault prediction result of each component with the hardware fault result; if the fault prediction result is inconsistent with the hardware fault result, then determine that the fault prediction result is a false alarm event; and adjust the alarm threshold corresponding to the preset prediction model based on the false alarm event.

[0085] In one example, the first acquisition module 601 is further configured to acquire the number of consecutive false alarm events occurring within a first time period; the first processing module 602 is further configured to increase or decrease the alarm threshold if the number of consecutive false alarm events is greater than or equal to a preset threshold.

[0086] In some optional implementations, the first acquisition module 601 is further configured to acquire historical operating data of multiple components included in the server and store the historical operating data in shared memory.

[0087] Embodiments of this application also provide yet another server fault prediction device, applied to the second processing core of the aforementioned server fault prediction system; such as Figure 7 As shown, Figure 7 A structural block diagram of a server fault prediction device provided in an embodiment of this application; the server fault prediction device includes:

[0088] The second acquisition module 701 is used to acquire runtime data from shared memory.

[0089] The second processing module 702 is used to perform feature extraction processing on the operating data to obtain operating feature data. The operating feature data indicates the trend of the operating status of each component changing over time during operation. Based on the operating feature data, server faults are predicted, and fault prediction results for each component of the server are obtained.

[0090] In some optional implementations, the second processing module 702 is specifically used to input dynamic time series features, discrete event density features and periodic fluctuation features into a preset prediction model and output the server's fault prediction results.

[0091] In some optional implementations, the second acquisition module 701 is further configured to acquire historical running data from shared memory and perform feature extraction processing on the historical running data to obtain historical running feature data; the second processing module 702 is further configured to train a model on the historical running feature data based on a time series prediction algorithm and a classification algorithm to obtain a preset prediction model.

[0092] In some alternative implementations, the second processing module 702 is also used to store the fault prediction results into shared memory.

[0093] For a description of the features in the embodiment corresponding to the server fault prediction device, please refer to the relevant description of the embodiment corresponding to the server fault prediction method, which will not be repeated here.

[0094] Embodiments of this application also provide an electronic device, such as... Figure 8 As shown, Figure 8 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. The electronic device includes a processor 10 and a memory 20, in which a computer program is stored. The processor 10 is configured to run the computer program to execute the steps in any of the above-described embodiments of the server fault prediction method. The electronic device may be a baseboard management controller in a server fault prediction system.

[0095] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described server fault prediction method embodiments when running.

[0096] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0097] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described server fault prediction method embodiments.

[0098] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described server fault prediction method embodiments.

[0099] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0100] The foregoing has provided a detailed description of a server fault prediction system, method, and storage medium provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only intended to aid in understanding the method and core ideas of this application. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A server fault prediction system, characterized in that, The server's fault prediction system includes a baseboard management controller comprising a first processing core, a second processing core, and shared memory, wherein the shared memory is communicatively connected to the first processing core and the second processing core; The first processing core is used to acquire the operating data of multiple components included in the server and store the operating data in shared memory; The second processing core is used to obtain the running data from the shared memory and perform feature extraction processing on the running data to obtain running feature data. The running feature data is used to indicate the trend of the running status of each component changing over time during the running process. Based on the operational characteristic data, the failures of the server are predicted, and the failure prediction results for each component of the server are obtained.

2. The system according to claim 1, characterized in that, The operational characteristic data includes dynamic time-series characteristics, discrete event density characteristics, and periodic fluctuation characteristics; the dynamic time-series characteristics are used to indicate the operational data characteristics of each component in a continuous time series; the discrete event density characteristics are used to indicate the distribution density characteristics of abnormal events occurring in each component in the time dimension; the periodic fluctuation characteristics are used to indicate the characteristics of each component that repeat with a fixed time period. The second processing core is specifically used to input the dynamic time series features, the discrete event density features, and the periodic fluctuation features into a preset prediction model, and output the fault prediction result of the server.

3. The system according to claim 2, characterized in that, The server's fault prediction system also includes a client; The second processing core is also used to store the fault prediction result into the shared memory; The first processing core is further configured to obtain the fault prediction result from the shared memory and send the fault prediction result to the client.

4. The system according to claim 3, characterized in that, The first processing core is also used to send heartbeat information to the second processing core; when the first processing core does not receive response information from the second processing core based on the heartbeat information within a preset time period, it determines that the thread corresponding to the preset prediction model on the second processing core is deadlocked, triggers the second processing core to perform a restart operation, and sends an exception message to the client, the exception message being used to indicate that the baseboard management controller cannot predict the server's failure.

5. The system according to any one of claims 1-4, characterized in that, The first processing core is also used to acquire historical operating data of multiple components included in the server and store the historical operating data in shared memory; The second processing core is further configured to obtain the historical running data from the shared memory, and perform feature extraction processing on the historical running data to obtain historical running feature data; and perform model training on the historical running feature data based on time series prediction algorithm and classification algorithm to obtain a preset prediction model; The time-series prediction algorithm includes a long short-term memory network algorithm and an autoregressive integral moving average algorithm. The long short-term memory network algorithm is used to capture the long-term dependencies between various features in the historical operating feature data, and the autoregressive integral moving average algorithm is used to analyze the periodic fluctuations of various features in the historical operating feature data. The classification algorithm is used to identify faults in the historical operating feature data based on alarm thresholds.

6. The system according to claim 5, characterized in that, The server includes the operational data of its multiple components, including the hardware data of each component. The first processing core is further configured to compare the hardware data of each component with the corresponding fault threshold to determine the hardware fault result of each component; and to compare the fault prediction result and the hardware fault result of each component. If the fault prediction result is inconsistent with the hardware fault result, then the fault prediction result is determined to be a false alarm event. Based on the false alarm event, adjust the alarm threshold corresponding to the preset prediction model.

7. The system according to claim 6, characterized in that, The first processing core is specifically used to obtain the number of consecutive false alarm events occurring within a first time period; if the number of consecutive false alarm events is greater than or equal to a preset threshold, the alarm threshold is increased or decreased.

8. A method for predicting server failures, characterized in that, A second processing core is applied to a baseboard management controller, the baseboard management controller including a first processing core, a second processing core, and shared memory, the shared memory being communicatively connected to the first processing core and the second processing core; the method includes: The running data is obtained from the shared memory, and the running data is the running data of multiple components of the server obtained by the first processing core; The operational data is processed by feature extraction to obtain operational feature data, which is used to indicate the trend of the operational status of each component changing over time during operation. Based on the operational characteristic data, the failures of the server are predicted, and the failure prediction results for each component of the server are obtained.

9. The method according to claim 8, characterized in that, The operational characteristic data includes dynamic time-series characteristics, discrete event density characteristics, and periodic fluctuation characteristics; the dynamic time-series characteristics are used to indicate the operational data characteristics of each component in a continuous time series; the discrete event density characteristics are used to indicate the distribution density characteristics of abnormal events occurring in each component in the time dimension; the periodic fluctuation characteristics are used to indicate the characteristics of each component that repeat with a fixed time period. The process of predicting server failures based on the operational characteristic data, and obtaining failure prediction results for each component of the server, includes: The dynamic time series features, the discrete event density features, and the periodic fluctuation features are input into a preset prediction model, and the fault prediction results of the server are output.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, it implements the steps of the fault prediction method for the server as described in any one of claims 8-9.