Fault monitoring method and device

By collecting and comparing reference and real-time operation data during the startup process of the basic input and output system, and identifying numerical and timing abnormalities, the problem of low accuracy of fault monitoring of basic input and output system is solved, and the stable operation of the server is achieved.

CN120353668BActive Publication Date: 2025-08-29INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510848429.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-08-29
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

In the prior art, the accuracy of basic input and output system fault monitoring is not high, especially in the problems of poor dynamic adaptability and difficult to identify hidden faults.

Method used

By collecting the benchmark operation data set and the real-time operation data set in the target startup stage, combining the dynamic comparison of trigger signals, the whole machine power consumption time sequence data and the chassis temperature time sequence data, numerical and timing abnormalities are identified and fault warning prompts are generated.

Benefits of technology

It improves the accuracy and timeliness of fault monitoring of basic input and output systems, avoids missed judgments of implicit faults, and ensures the stable operation of the server.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353668B_ABST
    Figure CN120353668B_ABST
Patent Text Reader

Abstract

The present application discloses a fault monitoring method and apparatus, relating to the field of fault monitoring technology, including obtaining a benchmark operating data set corresponding to a target startup phase under normal circumstances, including a benchmark trigger signal, benchmark whole-machine power consumption timing data, and benchmark chassis temperature timing data; further, in the process of starting a server in real time through a basic input / output system, collecting a real-time operating data set of the target startup phase, and comparing the data results of the real-time operating data set with the benchmark operating data set. Then, based on the dynamic matching of the timing data, the pertinence of the target startup phase, and in combination with the triggering time point of the trigger signal, accurate dynamic timing comparison can be performed to identify when a numerical anomaly occurs in the current target startup phase during the startup process, and whether the timing conforms to the rules of the benchmark data; further, the present application can improve the accuracy and timeliness of basic input / output system startup fault monitoring to ensure stable operation of the server.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of fault monitoring, and in particular to a fault monitoring method and device. Background Art

[0002] With the advancement of computing, Basic Input / Output System (BIOS) fault diagnosis is crucial for ensuring stable device operation. Current BIOS fault monitoring solutions rely on static threshold monitoring of hardware parameters, comparing them against fixed power consumption limits. However, this approach still has limitations.

[0003] First, dynamic adaptability is poor. In different business scenarios (for example, when a server's central processing unit (CPU) is idle or fully loaded), the current consumption during the basic input and output system startup process varies greatly, and the resulting power consumption inevitably varies. Traditional static threshold comparison methods cannot adapt to these dynamic changes, making it difficult to accurately reflect the health status of the hardware and prone to missed or misjudgment.

[0004] On the other hand, hidden faults are difficult to identify. Non-overcurrent faults such as chip latch-up and partial short circuits, due to the normal current consumption, result in power consumption that is difficult to trigger the fixed power consumption limit. Therefore, traditional solutions cannot detect them, causing the fault to spread latently.

[0005] Therefore, how to improve the accuracy of basic input and output system fault monitoring and ensure the stable operation of the computer has become an urgent problem to be solved. Summary of the Invention

[0006] The present application provides a fault monitoring method and apparatus to at least solve the problem of low accuracy in basic input and output system fault monitoring in the related art.

[0007] This application provides a fault monitoring method applied to a baseboard management controller, comprising:

[0008] During the process of starting the server through the basic input and output system, a benchmark operation data set corresponding to a normal target startup phase is collected; the benchmark operation data set includes a benchmark trigger signal, benchmark whole-machine power consumption time series data, and benchmark chassis temperature time series data for the target startup phase; the benchmark trigger signal is used to obtain the time for collecting the benchmark whole-machine power consumption time series data and the benchmark chassis temperature time series data;

[0009] In response to a server startup instruction, a real-time operation data set corresponding to the target startup phase is acquired; the real-time operation data set includes a real-time trigger signal corresponding to the target startup phase, real-time whole machine power consumption time series data, and real-time chassis temperature time series data;

[0010] Based on a data comparison result of the real-time operation data set corresponding to the target startup phase and the benchmark operation data set, determining whether an abnormality occurs in the target startup phase of the basic input and output system;

[0011] If so, a fault warning prompt for the target startup phase is generated.

[0012] The present application also provides a fault monitoring device, comprising:

[0013] A collection unit is configured to collect a benchmark operating data set corresponding to a normal target startup phase during the process of starting the server through a basic input and output system; the benchmark operating data set includes a benchmark trigger signal, benchmark whole-machine power consumption time series data, and benchmark chassis temperature time series data during the target startup phase; the benchmark trigger signal is used to obtain a time for collecting the benchmark whole-machine power consumption time series data and the benchmark chassis temperature time series data;

[0014] an acquisition unit, configured to acquire, in response to a server startup instruction, a real-time operation data set corresponding to the target startup phase; the real-time operation data set including a real-time trigger signal corresponding to the target startup phase, real-time whole-machine power consumption time series data, and real-time chassis temperature time series data;

[0015] a judging unit, configured to judge whether an abnormality occurs in the target startup phase of the basic input / output system based on a data comparison result of the real-time operation data set corresponding to the target startup phase and the benchmark operation data set;

[0016] A generating unit is configured to generate a fault warning prompt for the target startup phase when an abnormal situation occurs in the target startup phase of the basic input and output system.

[0017] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned fault monitoring methods when executing the computer program.

[0018] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned fault monitoring methods are implemented.

[0019] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned fault monitoring methods when executed by a processor.

[0020] Through the present application, the basic input and output system is divided into startup stages during the process of starting a computer, and a benchmark operating data set corresponding to each startup stage under normal circumstances is collected; for the target startup stage, the benchmark operating data set includes a benchmark trigger signal, a benchmark whole-machine power consumption timing data, and a benchmark chassis temperature timing data; further, in the process of starting the server in real time through the basic input and output system, the real-time operating data set of the target startup stage is collected, and through the data comparison results of the real-time operating data set and the benchmark operating data set, dynamic timing comparison can be accurately performed based on the dynamic matching of the timing data, the pertinence of the target startup stage, and the trigger time point of the trigger signal to identify when the numerical abnormality occurs in the current target startup stage during the startup process and whether the timing conforms to the rules of the benchmark data; therefore, the present application combines the timing data and trigger signal of the startup process to perform fault monitoring on the target startup process, which can break the limitation of the traditional static threshold that can only identify fixed upper limit faults, and avoid the problem that hidden faults are difficult to detect, thereby greatly improving the accuracy and timeliness of basic input and output system startup fault monitoring to ensure stable operation of the server. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0022] Figure 1 A schematic diagram of the architecture of a fault monitoring method provided in an embodiment of the present application;

[0023] Figure 2 One of the flow charts of a fault monitoring method provided in an embodiment of the present application;

[0024] Figure 3 A second flow chart of a fault monitoring method provided in an embodiment of the present application;

[0025] Figure 4 A schematic diagram of the structure of a fault monitoring device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0026] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0027] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.

[0028] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0029] The schematic diagram of the server hardware management architecture that the fault monitoring method provided in this application relies on is shown in FIG. Figure 1 As shown, the architecture is described here. The architecture includes:

[0030] The central processing unit (CPU) 11 is the core computing unit of the server, used to process key tasks such as system instructions and data operations.

[0031] Platform Controller Hub (PCH) 12: Responsible for managing I / O devices (such as USB and SATA), integrating peripheral functions, coordinating communication between the CPU and other low-speed peripherals, and also assuming some system management responsibilities, such as power management and clock control.

[0032] The Complex Programmable Logic Device (CPLD) 13 can be flexibly programmed and configured to implement customized logic control. In this application, it mainly plays the role of signal switching, logic processing and distribution, and coordinates the signal interaction between different modules.

[0033] The Baseboard Management Controller (BMC) 14 runs independently of the main system and is responsible for hardware status monitoring (such as temperature and voltage), fault diagnosis, and remote management (such as remote power on and off, and logging), ensuring stable server operation and manageability.

[0034] Sensors 15 are distributed throughout the server (such as the CPU, chassis, etc.) to collect hardware status data such as temperature and voltage, and transmit them to the BMC through the bus to provide basic data for system monitoring and management.

[0035] The power supply unit (PSU) 16 provides stable power to the server hardware modules and can feed back its own status (such as power and voltage output) to the BMC through the bus to facilitate monitoring of the power supply health.

[0036] Specifically, the CPU and PCH communicate via the Direct Media Interface (DMI) protocol. The CPU transmits high-speed data and instructions to the PCH, which then transmits relevant information such as peripherals back, collaboratively completing system operations and input / output (I / O) management. The PCH and CPLD interact via GPIO (general-purpose input / output pins). The PCH sends control signals, status information, and other information to the CPLD, and the CPLD can also feedback logically processed signals to the PCH, achieving functional coordination between the two. The CPLD and BMC are also connected via GPIO. The CPLD transmits processed signals (such as signals received and converted from the PCH) to the BMC, helping the BMC to obtain the status of different parts of the system. It can also receive control commands issued by the BMC to control other modules.

[0037] Sensors, PSUs, and BMCs are connected through an integrated circuit bus. Sensors collect temperature, voltage, and other data, and the PSU transmits its own power status data (such as output voltage, current, and power) to the BMC on a regular or on-demand basis. The BMC monitors hardware health based on this data and performs fault warnings and other operations.

[0038] This application focuses on server hardware management, with BMC as the management core. It connects CPU, PCH, CPLD, sensors, PSU and other modules through corresponding interfaces and buses to realize the architecture of hardware status collection, control and management, ensuring stable operation and maintainability of the server.

[0039] The embodiment of the present application provides a fault monitoring method, and the method is described in detail in conjunction with the execution process of the fault monitoring method. Figure 2 FIG. 1 is a flow chart of a fault monitoring method for a baseboard management controller provided by the present application, and the specific steps include the following:

[0040] S201 : In the process of starting a server through a basic input and output system, a benchmark operation data set corresponding to a normal target startup phase is collected.

[0041] Among them, the benchmark operation data set includes the benchmark trigger signal of the target startup phase, the benchmark whole machine power consumption timing data and the benchmark chassis temperature timing data; the benchmark trigger signal is used to obtain the time for collecting the benchmark whole machine power consumption timing data and the benchmark chassis temperature timing data.

[0042] In some embodiments, the Basic Input / Output System (BIOS) is responsible for hardware initialization and firmware loading when the server is started, and has a fixed timing.

[0043] In the embodiment of the present application, the basic input and output system is divided into three startup phases, namely: a CPU initialization phase, a memory training phase, and a high-speed peripheral component interconnect (PCIe) device enumeration phase; wherein the CPU initialization phase includes activating the basic operating capabilities of the CPU, completing hardware reset and basic parameter configuration, initializing the internal cache and on-chip bus; and setting up the interrupt controller and exception handling mechanism to lay the foundation for subsequent hardware interaction; the memory training phase includes scanning the memory slots, reading the SPD chip to obtain standard timing parameters, and verifying the physical connection; at the same time, it also performs full address range read and write tests, marks bad blocks and maps spare addresses to ensure memory reliability; the PCIe device enumeration phase includes scanning the PCIe link, reading the device identification (ID) and negotiating transmission parameters, allocating memory addresses and interrupt resources, completing device driver initialization; and testing key device functions to ensure the normal operation of peripherals.

[0044] Specifically, the benchmark operation data set includes normal standard data collected in advance when the server is normally started through the basic input and output system, which is used for subsequent comparison with the real-time operation data set for fault monitoring.

[0045] Among them, the reference trigger signal is used to mark the characteristic signal of the start of pre-collection of power consumption timing data and temperature timing data in the target startup phase, and then obtain the collection time point of power consumption timing data and temperature timing data through the reference trigger signal, so as to facilitate the subsequent comparison of the reference trigger signal with the real-time collected trigger signal to obtain whether the time of the collected data is aligned, so as to assist in fault monitoring.

[0046] Baseline system power consumption time series data is pre-collected data showing the time-varying power consumption of the server during the target startup phase. This data is typically collected through the power supply unit. It should be noted that in addition to power consumption time series data, baseline time series data corresponding to voltage and current can also be pre-collected, recorded at a millisecond frequency (e.g., every 10ms). This data reflects the dynamic characteristics of power consumption during the current startup phase and is used for subsequent data comparison to determine whether real-time power consumption deviates from the baseline. Furthermore, by comparing the baseline time series data for voltage and current with the real-time data, anomalies can be more accurately identified.

[0047] Baseline chassis temperature time series data is pre-collected data on chassis temperature changes over time, collected in real time by temperature sensors. This data is then combined with power consumption data to analyze hardware heating and heat dissipation status (for example, whether the temperature rises synchronously when power consumption suddenly increases), assisting in the identification of hidden faults.

[0048] It should be noted that the method for obtaining the benchmark operating data set can be to collect the benchmark operating data set for each startup stage multiple times, and then take the average value of the data collected multiple times as the final benchmark operating data set, and associate the benchmark operating data set corresponding to each startup stage with the corresponding stage identifier, so as to facilitate the subsequent comparison of the benchmark operating data set of the corresponding stage with the real-time operating data set through the stage identifier.

[0049] S202 : In response to a computer startup instruction issued by a user, when starting a target startup phase by starting a server in real time through a basic input and output system, obtaining a real-time running data set corresponding to the target startup phase.

[0050] The real-time operation data set includes the real-time trigger signal corresponding to the target startup phase, the real-time whole machine power consumption time series data, and the real-time chassis temperature time series data.

[0051] When the server receives the computer startup instruction from the user, such as pressing the computer power button, the server will enter the BIOS real-time startup process. At this time, any of the above startup stages can be used as the target startup stage. Then, when the target startup stage starts, the real-time running data set corresponding to the target startup stage is synchronously collected, which is convenient for subsequent comparison with the benchmark running data set corresponding to the target startup stage.

[0052] Specifically, the real-time trigger signal is a characteristic signal used to mark the start of real-time collection of power consumption timing data and temperature timing data during the target startup phase. Real-time whole-machine power consumption timing data is a sequence of data showing the real-time changes in the power consumption of the server during the target startup phase over time. It is usually collected through the power supply unit (PSU). It should be noted that in addition to power consumption timing data, timing data corresponding to real-time voltage and current can also be obtained. These data are recorded at the same frequency as the benchmark operating data set, i.e., at a millisecond frequency (e.g., once every 10ms). Furthermore, the timing data corresponding to the real-time voltage and current can be compared with the timing data corresponding to the benchmark voltage and current to provide comprehensive data comparison for fault monitoring and improve the accuracy of fault location. Real-time chassis temperature timing data can be data collected in real time by a temperature sensor on the changes in chassis temperature over time.

[0053] Based on the above real-time collected data, a real-time operation data set corresponding to the target startup stage is generated. When generating the real-time operation data set, in order to quickly and accurately locate which startup stage it is, it is necessary to associate the real-time trigger signal, real-time whole machine power consumption timing data and real-time chassis temperature timing data with the stage identifier of the corresponding target startup stage to facilitate subsequent data comparison operations.

[0054] Specifically, the above S202 can be further refined into the following steps 1 and 2:

[0055] Step 1: Receive hardware interrupt signal.

[0056] Among them, the hardware interrupt signal is sent when the basic input and output system reads the real-time trigger signal.

[0057] After receiving the startup command from the user, the BIOS executes the startup process. When the BIOS receives a real-time trigger signal, indicating that it is about to enter a certain startup phase, it needs to start collecting real-time data, that is, start collecting real-time whole-machine power consumption timing data and real-time chassis temperature data.

[0058] The signal to the BMC to start data collection is transmitted by sending a hardware interrupt signal. Specifically, the hardware interrupt signal is used to quickly transmit key events at the hardware level, allowing the system to respond in a timely manner. The hardware interrupt signal can be understood as an emergency message between hardware modules. In the embodiment of the present application, when the BIOS recognizes the real-time trigger signal during the startup phase, it will immediately send an electrical signal through the hardware circuit, namely the hardware interrupt signal, to notify the BMC to collect data.

[0059] It should be noted that the mechanism for generating a real-time trigger signal involves inserting the phase identifier corresponding to the target startup phase into the code start position corresponding to the target startup phase before the BIOS begins initializing the target startup phase. Then, when the BIOS reads the phase identifier at the code start position corresponding to the target startup phase, it obtains the phase identifier corresponding to the target startup phase to generate a real-time trigger signal.

[0060] Specifically, the stage identifier POST CODE can be understood as a software mark, which can be a hexadecimal identification code; the corresponding stage identifiers can be set in advance for different startup stages of BIOS, such as the CPU initialization stage, the memory training stage, and the PCIe device enumeration stage, such as setting the stage identifier of the CPU initialization stage to 0xA1, the stage identifier of the memory training stage to 0xB1, and the stage identifier of the PCIe device enumeration stage to 0xC1; so as to distinguish these three startup stages through the stage identifiers.

[0061] By setting a corresponding stage identifier for each startup phase, a corresponding relationship is established between code execution and hardware signal triggering, achieving precise coordination between hardware and software. That is, when code execution reaches a specific stage, a trigger signal is immediately generated to trigger the BMC to collect data. At the same time, the stage identifier can accurately locate the stage where a fault occurs during the fault monitoring phase, which helps to determine the cause of the fault.

[0062] Step 2: In response to the hardware interrupt signal, collect the real-time running data set corresponding to the target startup phase.

[0063] The real-time operation data set includes the real-time trigger signal corresponding to the target startup phase, the real-time whole machine power consumption time series data, and the real-time chassis temperature time series data.

[0064] Specifically, after the BMC responds to a hardware interrupt signal, it immediately initiates the real-time data set collection process to collect real-time trigger signals, real-time system power consumption time series data, and real-time chassis temperature time series data. The real-time trigger signal can be understood as the time point when the hardware interrupt signal is generated, which is subsequently used to align with the baseline trigger signal. The real-time system power consumption time series data forms a power consumption curve for subsequent data comparison. Similarly, the real-time chassis temperature time series data can also form a temperature curve for subsequent data comparison.

[0065] The hardware logic for generating and responding to the aforementioned hardware interrupt signal is as follows: When the BIOS boots the server, it executes the code corresponding to each boot phase, executing each phase's code sequentially. Because the corresponding phase identifier is embedded at the beginning of each boot phase's code, the BIOS reads the phase identifier when it begins executing the code for each boot phase. The PCH, acting as the platform control hub, monitors the BIOS phase status via GPIO pins. Upon detecting that the BIOS has read the phase identifier, the PCH sends the read phase identifier to the CPLD via the GPIO pins. Upon receiving the phase identifier, the CPLD immediately generates a trigger signal, generates a hardware interrupt signal based on the trigger signal, and sends the hardware interrupt signal to the BMC via the GPIO pins. For example, the hardware interrupt signal can be edge-triggered (e.g., level transition, such as from low to high) or level-triggered (continuously high level) to ensure a timely BMC response. Upon detecting the hardware interrupt signal from the CPLD, the BMC immediately initiates the data collection process.

[0066] In the embodiment of the present application, when the basic input / output system reads the real-time trigger signal, a hardware interrupt signal is issued, thereby triggering a hardware interrupt process, so that the BMC responds to the hardware interrupt signal to synchronously collect the power consumption timing data, temperature timing data, and trigger signal of the corresponding target startup phase, accurately captures the operating data of the startup phase, and provides accurate real-time data for fault diagnosis.

[0067] S203 : Based on the data comparison result of the real-time operation data set corresponding to the target startup phase and the benchmark operation data set, it is determined whether an abnormality occurs in the target startup phase of the basic input and output system.

[0068] In an embodiment of the present application, through multi-dimensional comparison of the real-time running data set and the benchmark running data set, anomalies in the BIOS startup phase are identified, such as numerical anomalies (such as power consumption and temperature exceeding the standard) and timing anomalies (such as stage delays and trigger signal misalignment), and then it is determined whether the startup phase is abnormal.

[0069] It should be noted that, in the embodiment of the present application, in the fault diagnosis during the server startup phase, the collected trigger signals and temperature timing data are mainly used to assist in determining whether the real-time power consumption timing data is abnormal.

[0070] For the target startup phase, the real-time power consumption time series data is aligned with the power consumption time series data of the benchmark running data set, starting from the timestamp of the trigger signal. For example, when the target startup phase is the memory training phase, according to the benchmark running data set, the power consumption in the 100ms of this phase is 60W. Then, it is necessary to compare the real-time power consumption corresponding to the same time point in the real-time running data set to see if it is 60W; then perform numerical anomaly judgment and timing anomaly judgment.

[0071] Furthermore, to determine numerical anomalies, the real-time system power consumption time series data can be compared with the benchmark system power consumption time series data at a single time point. This means traversing the real-time system power consumption time series data and comparing it with the benchmark system power consumption time series data at the corresponding time point. If the real-time data deviation exceeds a preset threshold range (e.g., ±5%) from the benchmark data, the power consumption time series data corresponding to that time point is marked as an abnormal data point. For example, the power consumption at the 200ms mark during the memory training phase in the benchmark data set is 50W, while the power consumption at the 200ms mark during the memory training phase in the real-time data set is 60W. This deviation is determined to be 10W, and the deviation from the benchmark data is 10%, which exceeds the preset threshold range of 5%, thus determining that the power consumption is abnormal.

[0072] To identify timing anomalies, we can compare the sequence duration of real-time time series data with the sequence duration of benchmark time series data. If the sequence duration of the real-time time series data exceeds the sequence duration threshold of the benchmark time series data (e.g., ±5%), a timing anomaly is determined. For example, if the sequence duration of the memory training phase in the benchmark dataset lasts for 500ms, while the sequence duration in the real-time dataset lasts for 700ms, the deviation is 40% of the benchmark sequence duration, which is greater than 5%. Therefore, it is determined that the duration of this phase is too long and inconsistent with the normal duration, thus determining a timing anomaly.

[0073] You can also check whether the trigger timing of the real-time trigger signal is consistent with the baseline trigger signal. If the timestamp deviation between the real-time trigger signal and the baseline trigger signal exceeds a threshold (e.g., ±50ms), a timing anomaly is determined. For example, the memory training phase in the benchmark dataset should be triggered and executed 1000ms after startup, but the real-time dataset is triggered and executed 1200ms after startup. The deviation reaches 200ms, which is much greater than 50ms. This indicates a trigger delay and a timing anomaly. This situation may be caused by BIOS execution lag or aging server hardware causing slow operation.

[0074] You can also first generate the corresponding real-time power consumption change curve and benchmark power consumption change curve based on the real-time whole-machine power consumption timing data and the benchmark whole-machine power consumption timing data, and then compare the overall trends of the two curves. If the benchmark power consumption change curve in a certain startup phase shows a trend of first rising and then stabilizing, while the real-time power consumption change curve shows a continuous increase, and the two trends are obviously different, then it is judged that a timing anomaly has occurred; further, the abnormal timing segment can be determined based on the slope deviation of the two curves at the same time point, and the abnormal timing segment can be intercepted for subsequent fault monitoring.

[0075] During server startup, fault diagnosis uses temperature time series data to determine if the real-time power consumption time series data is abnormal. By referring to the benchmark temperature curve corresponding to the benchmark time series temperature data, we can measure whether the real-time temperature curve corresponding to the real-time temperature data is consistent in trend. If the real-time temperature curve deviates from the benchmark temperature curve, this indicates abnormal power consumption or a localized heat dissipation problem, potentially leading to power consumption anomalies and the generation of fault warnings to prevent missed faults. Sudden temperature changes may also occur in real time, allowing us to locate transient high power consumption not captured by the power consumption data through the real-time temperature change points, thus addressing the shortcomings of simply determining power consumption thresholds.

[0076] Furthermore, through the judgment of the above-mentioned numerical anomalies, time series anomalies, and trend anomalies, combined with the assistance of time series temperature data, the embodiment of the present application can capture invisible anomalies that cannot be identified by fixed thresholds, and discover potential hardware risks by analyzing the trend of data time series changes. Compared with the traditional judgment based on fixed thresholds, the embodiment of the present application monitors the entire time series data for anomalies, and finally combines the stage identifier with the time series data to identify the periodic anomalies of a specific stage to construct a dynamic time series fault monitoring system, thereby improving the comprehensiveness and accuracy of fault monitoring.

[0077] After the above S203 (determining whether an abnormality occurs during the target startup phase of the basic input and output system) is executed, if an abnormality occurs during the target startup phase of the basic input and output system, the following S204 is executed:

[0078] S204: Generate a fault warning prompt for the target startup phase.

[0079] When an abnormality is detected in the target startup phase of the BIOS, the system will generate a fault warning prompt corresponding to the target startup phase to remind the operation and maintenance personnel that an abnormality has occurred in the target startup phase and that inspection and repair are required.

[0080] Specifically, early server failures can be intercepted by sending fault warning prompts to operation and maintenance personnel, and hidden hardware defects (such as CPU cache degradation) can be monitored when the server is started daily (such as restarting or turning on the computer) but no business is loaded. Subsequent business operations can be blocked in a timely manner to avoid wasting resources and time.

[0081] Through the present application, the basic input and output system is divided into startup stages during the process of starting a computer, and a benchmark operating data set corresponding to each startup stage under normal circumstances is collected; for the target startup stage, the benchmark operating data set includes a benchmark trigger signal, a benchmark whole-machine power consumption timing data, and a benchmark chassis temperature timing data; further, in the process of starting the server in real time through the basic input and output system, the real-time operating data set of the target startup stage is collected, and through the data comparison results of the real-time operating data set and the benchmark operating data set, dynamic timing comparison can be accurately performed based on the dynamic matching of the timing data, the pertinence of the target startup stage, and the trigger time point of the trigger signal to identify when the numerical abnormality occurs in the current target startup stage during the startup process and whether the timing conforms to the rules of the benchmark data; therefore, the present application combines the timing data and trigger signal of the startup process to perform fault monitoring on the target startup process, which can break the limitation of the traditional static threshold that can only identify fixed upper limit faults, and avoid the problem that hidden faults are difficult to detect, thereby greatly improving the accuracy and timeliness of basic input and output system startup fault monitoring to ensure stable operation of the server.

[0082] As an extension and refinement of the above embodiment, refer to Figure 3 As shown, the present application also provides a fault monitoring method, the specific steps of which include the following:

[0083] S301 : In the process of starting a server through a basic input and output system, a benchmark operation data set corresponding to a normal target startup phase is collected.

[0084] Among them, the benchmark operation data set includes the benchmark trigger signal of the target startup phase, the benchmark whole machine power consumption timing data and the benchmark chassis temperature timing data; the benchmark trigger signal is used to obtain the time for collecting the benchmark whole machine power consumption timing data and the benchmark chassis temperature timing data.

[0085] Specifically, since server hardware components age over time, to avoid baseline data deviations caused by hardware aging, the baseline operation data set corresponding to each startup phase can be periodically corrected based on the server's hardware aging coefficient to improve the accuracy of fault monitoring.

[0086] The aging coefficient characterizes the degree of hardware performance degradation over time and is typically expressed as a baseline parameter deviation rate. For example, the aging coefficient of server memory can be calculated by subtracting 1 from the ratio of the current baseline power consumption to the initial baseline power consumption. If the hardware has not aged, the current baseline power consumption and the initial baseline power consumption are theoretically equal. In this case, the ratio of the current baseline power consumption to the initial baseline power consumption is 1, and subtracting 1 yields 0, indicating that the memory has not aged. The initial baseline power consumption can be obtained from the server's factory data.

[0087] Specifically, the benchmark operation data set can be corrected once at a preset period (for example, every 3 months) to reduce the false alarm / missed alarm rate; then, through dynamic benchmark correction, fault diagnosis can be adapted to hardware performance degradation, and hardware aging coefficients can be periodically collected. The numerical and time series data of the benchmark operation data set can be dynamically corrected according to the startup phase, so that the benchmark operation data set continuously matches the real-time status of the hardware, ensuring the accuracy of fault monitoring.

[0088] S302: In response to a computer startup instruction issued by a user, a hardware interrupt signal is received.

[0089] Among them, the hardware interrupt signal is sent when the basic input and output system reads the real-time trigger signal.

[0090] S303 : In response to the hardware interrupt signal, collect a real-time operation data set corresponding to the target startup phase.

[0091] The real-time operation data set includes the real-time trigger signal corresponding to the target startup phase, the real-time whole machine power consumption time series data, and the real-time chassis temperature time series data.

[0092] In the embodiment of the present application, the above S303 (collecting a real-time running data set corresponding to the target startup phase in response to a hardware interrupt signal) can be further divided into the following steps 1 to 3:

[0093] Step 1: In response to a hardware interrupt signal, real-time whole-machine power consumption time series data and real-time chassis temperature time series data corresponding to a target startup phase are collected.

[0094] Specifically, the real-time whole-machine power consumption timing data and real-time chassis temperature timing data corresponding to the above-mentioned target startup phase can be the real-time whole-machine power consumption timing data of the power supply unit read based on the first bus, and the real-time chassis temperature timing data collected by the temperature sensor read based on the second bus.

[0095] Specifically, the first and second buses can be I2C buses. These two independent hardware buses are used to collect real-time power consumption and temperature data in parallel to ensure timing consistency. The BMC connects to the PSU via the first bus to read PSU registers to obtain real-time power consumption timing data. The BMC connects to the temperature sensor via the second bus to read sensor registers to obtain temperature timing data. It should be noted that the BMC is equipped with a clock source to synchronize the start time of the two bus acquisitions.

[0096] Step 2: Receive the phase identifier and real-time trigger signal corresponding to the target startup phase.

[0097] In an embodiment of the present application, the CPLD sends the phase identifier and the trigger signal to the BMC, and the BMC parses the phase identifier to clarify which startup phase the current startup phase is, and then associates all the data corresponding to the startup phase through the phase identifier to generate a real-time operation data set corresponding to the target startup phase, so as to facilitate the subsequent use of the phase identifier to accurately find the benchmark operation data set corresponding to the current target startup phase, which is helpful for subsequent data comparison.

[0098] Step 3: Correlate the real-time whole-machine power consumption time series data, real-time chassis temperature time series data, stage identifier, and real-time trigger signal corresponding to the target startup phase to generate a real-time operation data set.

[0099] Then, the real-time whole machine power consumption time series data, real-time chassis temperature time series data, stage identifier and real-time trigger signal corresponding to the target startup phase are correlated to obtain a real-time operation data set.

[0100] The embodiment of the present application associates four types of data, namely, real-time whole-machine power consumption time series data, real-time chassis temperature time series data, stage identification and trigger signal, through data association to construct a high-dimensional, structured real-time operation data set, thereby facilitating the subsequent comparison of the real-time operation data set with the benchmark operation data set, thereby accurately locating the startup stage where the fault occurs, and at the same time improving the efficiency of fault monitoring.

[0101] S304: Extract real-time key time series features from the real-time running data set.

[0102] Specifically, extracting the real-time key timing features from the real-time running data set can be understood as screening the real-time key timing features that can represent the running status of the startup phase from the real-time running data set.

[0103] This is because time series data (such as power consumption values ​​collected every 10ms) is a continuous curve, making direct comparison inefficient. By extracting key real-time timing features (such as the power consumption ramp rate and temperature stability), complex curves can be described with a small number of key features, reducing the computational complexity of subsequent data comparisons.

[0104] In some embodiments, key features such as the maximum, minimum, and average power consumption corresponding to each startup stage, the slope of the curve (such as the rate of temperature rise over time), the fluctuation period (such as the oscillation frequency of the power consumption curve), and the mutation characteristic step point (such as the power consumption suddenly jumps by 5W at a certain moment) can be extracted as real-time key timing features.

[0105] S305 : Combining the real-time key timing features with the benchmark key timing features corresponding to the benchmark operation data set, the real-time operation data set is compared with the benchmark operation data set to generate a data comparison result.

[0106] In some embodiments, corresponding baseline key timing features can be extracted from the baseline operation data set in advance, and then the extracted real-time key timing features can be used to compare the baseline key timing features, thereby quickly determining whether the real-time operation data deviates from the baseline operation data, so as to assist the data comparison process and improve the accuracy and efficiency of data comparison.

[0107] S306: Based on the data comparison result, determine whether an abnormality occurs during the target startup phase of the basic input and output system.

[0108] After the above S306 (determining whether an abnormality occurs during the target startup phase of the basic input / output system based on the data comparison result) is executed, if an abnormality occurs during the target startup phase of the basic input / output system, the following S307 is executed:

[0109] S307: Generate a fault warning prompt for the target startup phase.

[0110] In the embodiment of the present application, the real-time chassis temperature time series data of the target startup phase can be compared with the fixed safety temperature threshold to determine whether an abnormality occurs in the current target startup phase. The specific steps include the following:

[0111] Step A: Compare the real-time chassis temperature time series data with the preset chassis temperature data range.

[0112] Specifically, the preset chassis temperature data range can be understood as a fixed safety temperature threshold range, and then by comparing the real-time chassis temperature time series data with the preset chassis temperature data range, it is determined whether the real-time chassis temperature time series data exceeds the preset chassis temperature data range. When the real-time temperature time series data exceeds the preset chassis temperature data range, for example, the maximum temperature in the CPU startup phase should not exceed 60°C, but the maximum temperature in the real-time chassis temperature time series data has reached 70°C, this situation can directly determine that an abnormality has occurred in the target startup phase. It should be noted that different preset chassis temperature data ranges can be divided according to the startup phase. For example, the preset chassis temperature data range in the memory training phase is 30°C~60°C, and the preset chassis temperature data range in the CPU initialization phase is 40°C~55°C.

[0113] Step B: When the real-time chassis temperature time series data exceeds the preset chassis temperature data range, a fault warning prompt for the target startup phase is generated.

[0114] For example, when the preset chassis temperature data range corresponding to the memory training stage is 30℃~60℃, the highest temperature in the real-time running data set is 52℃ and the lowest temperature is 32℃, it means that the chassis temperature of the current memory training stage is within the safe range and there is no abnormality; otherwise, if it exceeds the preset chassis temperature data range, it is judged as abnormal.

[0115] As an extension and refinement of the above embodiment, after responding to the hardware interrupt signal, before collecting the real-time running data set corresponding to the target startup phase, it is also necessary to lock the speed of the fan in the basic input and output system from the intelligent speed regulation state to the preset speed.

[0116] In some embodiments, since the CPU, GPU and other chips in the server will trigger frequency reduction at high temperatures, resulting in performance degradation, long-term exposure to an overheated environment may cause accelerated hardware aging. In order to maintain the stable performance of hardware devices, fans will be installed in the chassis to dissipate heat and ensure that the hardware operates within the rated temperature range.

[0117] Therefore, in the embodiments of the present application, when collecting data during the startup phase, fluctuations in fan speed may cause chassis temperature changes, interfering with the accuracy of the test data. To reduce interference caused by the fan, the fan speed is locked from the intelligent speed regulation state to a preset speed before data collection is performed, maintaining a constant heat dissipation efficiency and reducing data errors caused by temperature fluctuations. It should be noted that pulse width modulation can be used to adjust the signal duty cycle to control the fan's output power, thereby achieving the purpose of controlling the fan speed. For example, setting the pulse width to 50% adjusts the fan speed to the preset speed.

[0118] Therefore, when pre-collecting the benchmark running data set, in order to maintain the relative stability of the data collection environment, the fan speed also needs to be controlled at a preset speed to reduce environmental variables and ensure the accuracy of the data comparison results.

[0119] It should also be noted that after determining whether an abnormality occurs in the target startup phase of the basic input and output system based on the data comparison result of the real-time operation data set corresponding to the target startup phase and the benchmark operation data set, if it is determined that no abnormality occurs in the target startup phase of the basic input and output system, the fan speed in the basic input and output system is restored from being locked to the preset speed to the intelligent speed regulation state.

[0120] Specifically, when the target startup phase is determined to be complete and no abnormalities occurred throughout the entire process, the fan speed can be restored from the preset speed to the intelligent speed control state. This is done until the next target startup phase begins data collection, at which point the fan speed in the basic input and output system is again locked from the intelligent speed control state to the preset speed. Furthermore, by flexibly adjusting the fan speed, the data collection environment is maintained in a stable state, improving data collection accuracy.

[0121] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0122] The embodiments of the present application also provide a fault monitoring device, which corresponds one-to-one with the method claims. Figure 4 A structural diagram of a data destruction device 400 provided by the present disclosure is shown as follows: Figure 4 As shown, the apparatus 400 of this embodiment includes:

[0123] The acquisition unit 41 is configured to acquire a benchmark operating data set corresponding to a normal target startup phase during the process of starting the server through the basic input and output system; the benchmark operating data set includes a benchmark trigger signal, benchmark whole-machine power consumption time series data, and benchmark chassis temperature time series data during the target startup phase; the benchmark trigger signal is used to obtain a time for acquiring the benchmark whole-machine power consumption time series data and the benchmark chassis temperature time series data;

[0124] An acquisition unit 42 is configured to acquire a real-time operation data set corresponding to the target startup phase in response to a server startup instruction; the real-time operation data set includes a real-time trigger signal corresponding to the target startup phase, real-time whole-machine power consumption time series data, and real-time chassis temperature time series data;

[0125] A judging unit 43 is configured to judge whether an abnormality occurs in the target startup phase of the basic input and output system based on a data comparison result of the real-time operation data set corresponding to the target startup phase and the benchmark operation data set;

[0126] The generating unit 44 is configured to generate a fault warning prompt for the target startup phase when an abnormality occurs in the target startup phase of the basic input and output system.

[0127] As an optional implementation of an embodiment of the present application, the acquisition unit 42 is specifically used to receive a hardware interrupt signal; the hardware interrupt signal is issued when the basic input and output system reads the real-time trigger signal; in response to the hardware interrupt signal, the real-time operation data set corresponding to the target startup phase is collected.

[0128] As an optional implementation of the embodiment of the present application, the acquisition unit 42 is further used to insert the stage identifier corresponding to the target startup stage into the code start position corresponding to the target startup stage before the target startup stage of the basic input and output system is initialized.

[0129] As an optional implementation of an embodiment of the present application, the acquisition unit 42 is also used to obtain the stage identifier corresponding to the target startup stage when the basic input and output system reads the stage identifier of the code start position corresponding to the target startup stage to generate a real-time trigger signal.

[0130] As an optional implementation of the embodiment of the present application, the acquisition unit 42 is further configured to lock the rotation speed of the fan in the basic input and output system from the intelligent speed regulation state to a preset rotation speed.

[0131] As an optional implementation of the embodiment of the present application, the acquisition unit 42 is also used to restore the fan speed in the basic input and output system from being locked to a preset speed to the intelligent speed regulation state when no abnormal situation occurs in the target startup phase of the basic input and output system.

[0132] As an optional implementation of an embodiment of the present application, the acquisition unit is specifically used to collect the real-time whole-machine power consumption timing data and the real-time chassis temperature timing data corresponding to the target startup phase in response to the hardware interrupt signal; receive the phase identifier and the real-time trigger signal corresponding to the target startup phase; associate the real-time whole-machine power consumption timing data, the real-time chassis temperature timing data, the phase identifier and the real-time trigger signal corresponding to the target startup phase to generate the real-time operation data set.

[0133] As an optional implementation of an embodiment of the present application, the acquisition unit 42 is specifically used to respond to the hardware interrupt signal, read the real-time whole-machine power consumption timing data of the power supply unit based on the first bus, and read the real-time chassis temperature timing data corresponding to the temperature sensor based on the second bus.

[0134] As an optional implementation of an embodiment of the present application, the judgment unit 43 is specifically used to extract real-time key timing features from the real-time operation data set; combine the real-time key timing features with the benchmark key timing features corresponding to the benchmark operation data set, compare the real-time operation data set with the benchmark operation data set, and generate the data comparison result; based on the data comparison result, determine whether an abnormality occurs in the target startup phase of the basic input and output system.

[0135] As an optional implementation of the embodiment of the present application, the acquisition unit 41 is further configured to periodically correct the benchmark operation data set corresponding to the target startup phase according to the hardware aging coefficient of the server.

[0136] As an optional implementation of an embodiment of the present application, the judgment unit 43 is also used to compare the real-time chassis temperature timing data with the preset chassis temperature data range; when the real-time chassis temperature timing data exceeds the preset chassis temperature data range, a fault warning prompt for the target startup phase is generated.

[0137] For the description of the features in the embodiment corresponding to the fault monitoring device, reference can be made to the relevant description of the embodiment corresponding to the fault monitoring method, which will not be repeated here.

[0138] An embodiment of the present application further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned fault monitoring method embodiments.

[0139] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps of any of the above-mentioned fault monitoring method embodiments when running.

[0140] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0141] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any one of the above-mentioned fault monitoring method embodiments are implemented.

[0142] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above-mentioned fault monitoring method embodiments are implemented.

[0143] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0144] The above is a detailed introduction to a fault monitoring method and device provided by the present application. This article uses specific examples to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core ideas of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.

Claims

1. A fault monitoring method, characterized in that: Applicable to baseboard management controllers, including: During the process of starting the server through the basic input and output system, a benchmark operation data set corresponding to a normal target startup phase is collected; the benchmark operation data set includes a benchmark trigger signal, benchmark whole-machine power consumption time series data, and benchmark chassis temperature time series data for the target startup phase; the benchmark trigger signal is used to obtain the time for collecting the benchmark whole-machine power consumption time series data and the benchmark chassis temperature time series data; In response to a server startup instruction, a real-time operation data set corresponding to the target startup phase is acquired; the real-time operation data set includes a real-time trigger signal corresponding to the target startup phase, real-time whole machine power consumption time series data, and real-time chassis temperature time series data; Based on a data comparison result of the real-time operation data set corresponding to the target startup phase and the benchmark operation data set, determining whether an abnormality occurs in the target startup phase of the basic input and output system; If so, a fault warning prompt for the target startup phase is generated.

2. The method according to claim 1, characterized in that The obtaining of a real-time operation data set corresponding to the target startup phase includes: receiving a hardware interrupt signal; the hardware interrupt signal is sent when the basic input and output system reads the real-time trigger signal; In response to the hardware interrupt signal, a real-time operation data set corresponding to the target startup phase is collected.

3. The method according to claim 1, characterized in that Before obtaining the real-time operation data set corresponding to the target startup phase, the method further includes: Before the target startup phase of the basic input / output system is initialized, a phase identifier corresponding to the target startup phase is inserted into a code start position corresponding to the target startup phase.

4. The method according to claim 3, characterized in that The method further comprises: When the basic input and output system reads the phase identifier of the code start position corresponding to the target startup phase, the basic input and output system obtains the phase identifier corresponding to the target startup phase to generate a real-time trigger signal.

5. The method according to claim 2, characterized in that After responding to the hardware interrupt signal, the method further includes: The speed of the fan in the basic input and output system is locked from the intelligent speed regulation state to the preset speed.

6. The method according to claim 5, characterized in that The method further comprises: If no abnormality occurs during the target startup phase of the basic input and output system, the fan speed in the basic input and output system is restored from being locked to the preset speed to the intelligent speed regulation state.

7. The method according to claim 2, characterized in that The step of collecting a real-time operation data set corresponding to the target startup phase in response to the hardware interrupt signal includes: In response to the hardware interrupt signal, collecting the real-time whole machine power consumption time series data and the real-time chassis temperature time series data corresponding to the target startup phase; receiving a phase identifier corresponding to the target startup phase and the real-time trigger signal; The real-time whole-machine power consumption time series data, the real-time chassis temperature time series data, the stage identifier, and the real-time trigger signal corresponding to the target startup phase are associated to generate the real-time operation data set.

8. The method according to claim 7, characterized in that The step of collecting the real-time whole-machine power consumption time series data and the real-time chassis temperature time series data corresponding to the target startup phase in response to the hardware interrupt signal includes: In response to the hardware interrupt signal, the real-time whole-machine power consumption time series data of the power supply unit is read based on the first bus, and the real-time chassis temperature time series data corresponding to the temperature sensor is read based on the second bus.

9. The method according to claim 1, characterized in that The determining whether an abnormality occurs in the target startup phase of the basic input / output system based on a data comparison result of the real-time operation data set corresponding to the target startup phase and the benchmark operation data set includes: Extracting real-time key timing features from the real-time running data set; combining the real-time key timing features with the benchmark key timing features corresponding to the benchmark operation data set, comparing the real-time operation data set with the benchmark operation data set to generate the data comparison result; According to the data comparison result, it is determined whether an abnormality occurs in the target startup phase of the basic input and output system.

10. The method according to claim 1, characterized in that The method further comprises: The benchmark operation data set corresponding to the target startup phase is periodically corrected according to the hardware aging coefficient of the server.

11. The method according to claim 1, wherein The method further comprises: Comparing the real-time chassis temperature time series data with a preset chassis temperature data range; When the real-time chassis temperature time series data exceeds the preset chassis temperature data range, a fault warning prompt for the target startup phase is generated.

12. A fault monitoring device, characterized in that: include: A collection unit, configured to collect a benchmark operation data set corresponding to a normal target startup phase during the process of starting the server through a basic input and output system; The benchmark operation data set includes a benchmark trigger signal, benchmark whole machine power consumption time series data and benchmark chassis temperature time series data of the target startup phase; The reference trigger signal is used to obtain the time for collecting the reference whole-machine power consumption time series data and the reference chassis temperature time series data; an acquiring unit, configured to acquire a real-time running data set corresponding to the target startup phase in response to a server startup instruction; The real-time operation data set includes a real-time trigger signal corresponding to the target startup phase, real-time whole machine power consumption time series data and real-time chassis temperature time series data; a judging unit, configured to judge whether an abnormality occurs in the target startup phase of the basic input / output system based on a data comparison result of the real-time operation data set corresponding to the target startup phase and the benchmark operation data set; A generating unit is configured to generate a fault warning prompt for the target startup phase when an abnormal situation occurs in the target startup phase of the basic input and output system.

13. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the fault monitoring method according to any one of claims 1 to 11 when executing the computer program.

14. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the fault monitoring method according to any one of claims 1 to 11.

15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the fault monitoring method according to any one of claims 1 to 11 are implemented.

Citation Information

Patent Citations

  • BIOS (basic input / output system) program startup monitoring method and electronic device

    CN107992397A

  • Substrate management controller adaptation method of basic input / output system and program product

    CN119201225A