Fault monitoring method and device

By collecting and comparing benchmark and real-time data during BIOS startup, BIOS failures are identified, and the problem of low accuracy of BIOS fault monitoring is solved, and the stable operation of the server is achieved.

CN120353668AActive Publication Date: 2025-07-22INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510848429.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-07-22
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

In the prior art, the basic input and output system (BIOS) fault monitoring is not accurate, and it is difficult to adapt to dynamically changing business scenarios and identify hidden faults, resulting in unstable operation of computer equipment.

Method used

By collecting the benchmark run data set and the real-time run data set in the BIOS startup stage, dynamic timing comparison is performed with the trigger signal, numerical and timing abnormalities are identified, and fault warning prompts are generated.

Benefits of technology

It improves the accuracy and timeliness of BIOS fault monitoring, avoids missed judgments of implicit faults, and ensures the stable operation of the server.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353668A_ABST
    Figure CN120353668A_ABST
Patent Text Reader

Abstract

The invention discloses a fault monitoring method and device, and relates to the technical field of fault monitoring, and the method comprises the steps: obtaining a corresponding reference operation data set under the normal condition of a target starting stage, and the reference operation data set comprises a reference trigger signal, reference complete machine power consumption time sequence data and reference case temperature time sequence data; furthermore, in the process of starting the server through the basic input and output system in real time, a real-time operation data set in a target starting stage is collected, and through a data comparison result of the real-time operation data set and the reference operation data set, the target starting stage can be started according to the dynamic matching performance of the time sequence data and the pertinence of the target starting stage; by combining the triggering time point of the triggering signal, dynamic time sequence comparison can be accurately carried out, so that when numerical value abnormity occurs in the current target starting stage in the starting process and whether the time sequence accords with the rule of reference data or not can be identified; therefore, the accuracy and timeliness of starting fault monitoring of the basic input and output system can be improved, so that stable operation of the server is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of fault monitoring, and in particular, to a fault monitoring method and device. Background Art

[0002] With the development of the computer field, the fault diagnosis of the Basic Input / Output System (BIOS) is crucial for ensuring the stable operation of devices. Currently, for the fault monitoring solution of the basic input / output system, it mostly relies on static threshold monitoring of hardware parameters, that is, by comparing with a fixed power consumption upper limit; however, the above method still has certain limitations.

[0003] On the one hand, the dynamic adaptability is poor. In different business scenarios (for example, when the Central Processing Unit (CPU) of a server is in an idle or fully loaded scenario), during the process of starting the server through the basic input / output system, the span of current consumption is large, and the generated power consumption will inevitably be different. Then, the traditional static threshold comparison method cannot adapt to this dynamic change, is difficult to accurately reflect the hardware health status, and is prone to missed judgments and false judgments.

[0004] On the other hand, it is difficult to identify hidden faults. For non-overcurrent faults such as chip latching effect and local short circuit, since the current consumption is within the normal range, the generated power consumption is difficult to trigger the fixed power consumption upper limit. Therefore, the traditional solution cannot detect it, resulting in the latent spread of faults.

[0005] Therefore, how to improve the accuracy of basic input / output system fault monitoring and ensure the stable operation of the computer has become an urgent problem to be solved. Summary of the Invention

[0006] This application provides a fault monitoring method and device to at least solve the problem of low accuracy in basic input / output system fault monitoring in related technologies.

[0007] This application provides a fault monitoring method, which is applied to a baseboard management controller and includes: During the process of starting the server through the basic input / output system, collect a reference operation data set corresponding to the normal situation in the target startup stage; the reference operation data set includes a reference trigger signal, reference overall machine power consumption time series data, and reference chassis temperature time series data in the target startup stage; the reference trigger signal is used to obtain the time for collecting the reference overall machine power consumption time series data and reference chassis temperature time series data; In response to a server startup instruction, obtain the real-time operation dataset corresponding to the target startup phase; the real-time operation dataset includes the real-time trigger signal, real-time overall machine power consumption time series data, and real-time chassis temperature time series data corresponding to the target startup phase; Based on the data comparison result between the real-time operation dataset corresponding to the target startup phase and the reference operation dataset, determine whether an abnormal situation occurs in the target startup phase of the basic input / output system; If so, generate a fault warning prompt for the target startup phase.

[0008] This application also provides a fault monitoring device, including: An acquisition unit, configured to acquire a reference operation dataset corresponding to a normal situation in a target startup phase during the process of starting a server through the basic input / output system; the reference operation dataset includes a reference trigger signal, reference overall machine power consumption time series data, and reference chassis temperature time series data corresponding to the target startup phase; the reference trigger signal is used to obtain the time for acquiring the reference overall machine power consumption time series data and reference chassis temperature time series data; An acquisition unit, configured to obtain the real-time operation dataset corresponding to the target startup phase in response to a server startup instruction; the real-time operation dataset includes the real-time trigger signal, real-time overall machine power consumption time series data, and real-time chassis temperature time series data corresponding to the target startup phase; A judgment unit, configured to determine whether an abnormal situation occurs in the target startup phase of the basic input / output system based on the data comparison result between the real-time operation dataset corresponding to the target startup phase and the reference operation dataset; A generation unit, configured to generate a fault warning prompt for the target startup phase when an abnormal situation occurs in the target startup phase of the basic input / output system.

[0009] This application also provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any of the above-mentioned fault monitoring methods when executing the computer program.

[0010] This application also provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of any of the above-mentioned fault monitoring methods are implemented.

[0011] This application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of any of the above-mentioned fault monitoring methods are implemented.

[0012] Through this application, the startup process of the basic input / output system for starting a computer is divided into startup phases, and the benchmark operation dataset corresponding to each normal startup phase is collected; for the target startup phase, the benchmark operation dataset includes a benchmark trigger signal, benchmark overall machine power consumption timing data, and benchmark chassis temperature timing data; furthermore, during the process of starting the server through the basic input / output system in real time, the real-time operation dataset of the target startup phase is collected. Through the data comparison result between the real-time operation dataset and the benchmark operation dataset, furthermore, according to the dynamic matching of the timing data, the pertinence of the target startup phase, and combined with the trigger time point of the trigger signal, dynamic timing comparison can be accurately performed to identify when numerical anomalies occur in the current target startup phase during the startup process and whether the timing conforms to the law of the benchmark data; therefore, this application combines the timing data and trigger signal of the startup process to monitor faults in the target startup process, which can break through the limitation that traditional static thresholds can only identify fixed upper limit faults and avoid the problem that hidden faults are difficult to detect. Furthermore, it can greatly improve the accuracy and timeliness of fault monitoring in the basic input / output system startup to ensure the stable operation of the server. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] In order to more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0014] Figure 1 Schematic diagram of the architecture of a fault monitoring method provided by an embodiment of the present application; Figure 2 Schematic diagram of the process of a fault monitoring method provided by an embodiment of the present application (one); Figure 3 Schematic diagram of the process of a fault monitoring method provided by an embodiment of the present application (two); Figure 4 Schematic diagram of the structure of a fault monitoring device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0015] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0016] It should be noted that, in the description of this application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects, and are not used to describe a specific order or sequence.

[0017] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.

[0018] The schematic diagram of the architecture of the server hardware management relied on by the fault monitoring method provided in this application is shown in FIG. Figure 1 As shown, the architecture is described here. The architecture includes: The central processing unit (CPU) 11 is the core computing unit of the server and is used to process key tasks such as system instructions and data operations.

[0019] Platform Controller Hub (PCH) 12: Responsible for managing I / O devices (such as USB, SATA, etc.), integrating peripheral functions, coordinating CPU and other low-speed peripheral communications, and also assuming some system management responsibilities, such as power management, clock control, etc.

[0020] The complex programmable logic device (CPLD) 13 can be flexibly programmed and configured to implement customized logic control. In this application, it mainly plays the role of signal switching, logic processing and distribution, and coordinates the signal interaction between different modules.

[0021] The baseboard management controller (BMC) 14 runs independently of the main system and is responsible for hardware status monitoring (such as temperature and voltage), fault diagnosis, remote management (such as remote power on and off, log recording), etc., to ensure stable operation and manageability of the server.

[0022] Sensors 15 are distributed in various locations of the server (such as the CPU, chassis, etc.) to collect hardware status data such as temperature and voltage, and transmit them to the BMC through the bus to provide basic data for system monitoring and management.

[0023] The Power Supply Unit (PSU) 16 provides stable power for each hardware module of the server, and at the same time can feedback its own status (such as power, voltage output, etc.) to the BMC through the bus, facilitating the monitoring of the power supply health status.

[0024] Specifically, the CPU and the PCH communicate through the Direct Media Interface (DMI) protocol. The CPU transmits high-speed data and instructions to the PCH, and the PCH then transmits relevant information such as peripherals back, collaborating to complete system operations and input / output (I / O) management. The PCH and the CPLD interact through GPIO (General Purpose Input / Output pins). The PCH sends control signals, status information, etc. to the CPLD, and the CPLD can also feedback signals processed logically to the PCH to achieve the collaborative function of the two. The CPLD and the BMC are also connected through GPIO. The CPLD transmits the processed signals (such as the signals received and converted from the PCH) to the BMC, helping the BMC obtain the status of different parts of the system, and can also receive control instructions issued by the BMC to regulate other modules.

[0025] The sensors, PSU, and BMC are connected through an integrated circuit bus. The sensors transmit data such as temperature and voltage collected, and the PSU transmits its own power supply status data (such as output voltage, current, power) to the BMC regularly or on demand. The BMC monitors the hardware health based on these data and performs operations such as fault warning.

[0026] This application focuses on server hardware management, with the BMC as the management core, and connects modules such as the CPU, PCH, CPLD, sensors, and PSU through corresponding interfaces and buses to implement an architecture for hardware status collection, control, and management, ensuring the stable operation and maintainability of the server.

[0027] Embodiments of this application provide a fault monitoring method. Combining the execution process of the fault monitoring method, the method is described in detail. Refer to Figure 2 As shown, it is a schematic flowchart of the fault monitoring method applied to the baseboard management controller provided by this application. The specific steps are as follows: S201. During the process of starting the server through the basic input / output system, collect the reference operation data set corresponding to the target startup stage under normal circumstances.

[0028] Among them, the reference operation data set includes the reference trigger signal, reference overall machine power consumption timing data, and reference chassis temperature timing data of the target startup stage; the reference trigger signal is used to obtain the time for collecting the reference overall machine power consumption timing data and the reference chassis temperature timing data.

[0029] In some embodiments, the Basic Input / Output System (BIOS) is responsible for hardware initialization and firmware loading during server startup and has fixed timing.

[0030] In the embodiments of the present application, the Basic Input / Output System is divided into three startup stages, which are: the CPU initialization stage, the memory training stage, and the Peripheral Component Interconnect Express (PCIe) device enumeration stage. Among them, the CPU initialization stage includes activating the basic operating capabilities of the CPU, completing hardware reset and basic parameter configuration, initializing the internal cache and on-chip bus; and setting the interrupt controller and exception handling mechanism to lay a foundation for subsequent hardware interaction. The memory training stage includes scanning the memory slots, reading the SPD chips to obtain standard timing parameters, and verifying the physical connection; at the same time, it also performs full address range read and write tests, marks bad blocks and maps spare addresses to ensure memory reliability. The PCIe device enumeration stage includes scanning the PCIe link, reading the device identification (ID) and negotiating transmission parameters, allocating memory addresses and interrupt resources, and completing device driver initialization; and testing the functions of key devices to ensure the normal operation of peripherals.

[0031] Specifically, the reference operation dataset includes normal standard data pre-collected when the server is normally started through the Basic Input / Output System, which is used for subsequent comparison with the real-time operation dataset for fault monitoring.

[0032] Among them, the reference trigger signal is a characteristic signal used to mark the start of pre-collecting the power consumption timing data and temperature timing data of the target startup stage. Then, the acquisition time points of the power consumption timing data and temperature timing data are obtained through the reference trigger signal, which is convenient for subsequent comparison between the reference trigger signal and the real-time collected trigger signal to obtain whether the acquisition time of the collected data is aligned to assist in fault monitoring.

[0033] The reference overall power consumption timing data is a sequence of data of the power consumption change over time of the server pre-collected during the target startup stage, which is usually collected by the power supply unit. It should be noted that in addition to the power consumption timing data, the reference timing data corresponding to voltage and current can also be pre-collected, which can be recorded at a millisecond-level frequency (such as once every 10 ms). Then, it reflects the dynamic characteristics of the power consumption in the current startup stage and is used to obtain whether the real-time power consumption deviates from the reference during subsequent data comparison. At the same time, by comparing the reference timing data of voltage and current with the real-time data of voltage and current, abnormal situations can be determined more accurately.

[0034] The reference chassis temperature time-series data is the data collected in advance for the change of the chassis temperature over time, which is collected in real time by a temperature sensor. Furthermore, by combining the power consumption data analysis, the heat generation and heat dissipation status of the hardware can be analyzed (such as whether the temperature rises synchronously when the power consumption suddenly increases), which helps to judge latent faults.

[0035] It should be noted that the acquisition method of the reference operation dataset can be to collect the reference operation datasets for each startup stage multiple times, and then take the average of the data collected multiple times as the final reference operation dataset, and associate the reference operation datasets corresponding to each startup stage with the corresponding stage identifiers, which is convenient for subsequent comparison of the reference operation dataset corresponding to the corresponding stage with the real-time operation dataset through the stage identifier.

[0036] S202. In response to the computer startup instruction issued by the user, when starting the server in real time through the basic input / output system to start the target startup stage, obtain the real-time operation dataset corresponding to the target startup stage.

[0037] Among them, the real-time operation dataset includes the real-time trigger signal corresponding to the target startup stage, the real-time overall machine power consumption time-series data, and the real-time chassis temperature time-series data.

[0038] After the server receives the computer startup instruction issued by the user, for example, when pressing the computer power-on button, the server will enter the BIOS real-time startup process. At this time, any one of the above startup stages can be used as the target startup stage. Then, at the beginning of the target startup stage, synchronously collect the real-time operation dataset corresponding to the target startup stage, which is convenient for subsequent comparison with the reference operation dataset corresponding to the target startup stage.

[0039] Specifically, the real-time trigger signal is a characteristic signal used to mark the start of real-time collection of the power consumption time-series data and temperature time-series data of the target startup stage. The real-time overall machine power consumption time-series data is the sequence data of the power consumption of the server changing over time in the target startup stage, which is usually collected by the power supply unit (PSU). It should be noted that in addition to the power consumption time-series data, the time-series data corresponding to the real-time voltage and current can also be obtained, and recorded at the same frequency as the reference operation dataset, that is, recorded at a millisecond-level frequency (such as once every 10 ms). Furthermore, the time-series data corresponding to the real-time voltage and current can be compared with the time-series data corresponding to the reference voltage and current, providing comprehensive data comparison for fault monitoring and improving the accuracy of fault location. The real-time chassis temperature time-series data can be the data collected in real time by a temperature sensor for the change of the chassis temperature over time.

[0040] Furthermore, based on the above real-time collected data, a real-time operation dataset corresponding to the target startup phase is generated. When generating the real-time operation dataset, in order to facilitate quick and accurate positioning of which startup phase it is, the real-time trigger signal, the real-time overall machine power consumption time series data, and the real-time chassis temperature time series data need to be associated with the phase identifier of the corresponding target startup phase, so as to facilitate subsequent data comparison operations.

[0041] Specifically, the above S202 can be further refined into the following Step 1 and Step 2: Step 1: Receive a hardware interrupt signal.

[0042] Among them, the hardware interrupt signal is sent when the basic input / output system reads the real-time trigger signal.

[0043] After receiving the startup instruction issued by the user, the BIOS executes the startup process. When the BIOS receives the real-time trigger signal, indicating that it is about to enter a certain startup phase, it is necessary to start collecting real-time data, that is, start collecting real-time overall machine power consumption time series data and real-time chassis temperature data.

[0044] By sending a hardware interrupt signal, a signal indicating that data collection needs to start is conveyed to the BMC; specifically, the hardware interrupt signal is a signal used to quickly transmit key events at the hardware level to enable the system to respond in a timely manner. The hardware interrupt signal can be understood as an emergency message between hardware modules. In the embodiment of the present application, when the BIOS recognizes the real-time trigger signal of the startup phase, it will immediately send an electrical signal through the hardware circuit, that is, the hardware interrupt signal, for the purpose of notifying the BMC to collect data.

[0045] It should be noted that the generation mechanism of the real-time trigger signal is to insert the phase identifier corresponding to the target startup phase into the start position of the code corresponding to the target startup phase before the initialization of the target startup phase of the basic input / output system. Then, when the basic input / output system reads the phase identifier at the start position of the code corresponding to the target startup phase, the phase identifier corresponding to the target startup phase is obtained to generate the real-time trigger signal.

[0046] Specifically, the phase identifier POST CODE can be understood as a software mark, which can be a hexadecimal identification code; corresponding phase identifiers can be set in advance for different startup phases of the BIOS, such as the CPU initialization phase, the memory training phase, and the PCIe device enumeration phase. For example, the phase identifier of the CPU initialization phase is set to 0xA1, the phase identifier of the memory training phase is set to 0xB1, and the phase identifier of the PCIe device enumeration phase is set to 0xC1; these three startup phases are distinguished by the phase identifier.

[0047] By setting the phase identifiers corresponding to each startup phase, a corresponding relationship is formed between code execution and hardware signal triggering, achieving precise coordination between hardware and software. That is, when the code executes to a specific phase, a trigger signal is immediately generated to trigger the BMC to collect data. At the same time, through the phase identifier, the phase where a fault occurs can be accurately located during the fault monitoring phase, which helps to determine the cause of the fault.

[0048] Step 2: In response to the hardware interrupt signal, collect the real-time operation data set corresponding to the target startup phase.

[0049] Among them, the real-time operation data set includes the real-time trigger signal corresponding to the target startup phase, the real-time overall machine power consumption time series data, and the real-time chassis temperature time series data.

[0050] Specifically, after the BMC responds to the hardware interrupt signal, it immediately starts the collection process of the real-time operation data set to collect the real-time trigger signal, the real-time overall machine power consumption time series data, and the real-time chassis temperature time series data. Among them, the real-time trigger signal can be understood as the time point when the hardware interrupt signal is generated, which is used to align with the reference trigger signal later; the power consumption change curve is formed through the real-time overall machine power consumption time series data, which is convenient for subsequent data comparison; similarly, the real-time chassis temperature time series data can also form a temperature change curve, which is convenient for subsequent data comparison.

[0051] The hardware logic of the generation and response process of the above hardware interrupt signal is specifically as follows: When the BIOS starts the server, it executes the code corresponding to each startup phase, that is, it sequentially executes the code of each startup phase according to the order of the startup phases. Since the corresponding phase identifier is embedded at the beginning of the code of each startup phase, when the BIOS starts to execute the code of each startup phase, the phase identifier of each startup phase will be read. At this time, the PCH, as the platform control center, monitors the BIOS phase status through the GPIO pin. That is, after the PCH detects that the BIOS reads the phase identification code, it will send the read phase identification code to the CPLD through the GPIO pin, so that after receiving the phase identification code, the CPLD immediately generates a trigger signal, and generates a hardware interrupt signal based on the trigger signal, and sends the hardware interrupt signal to the BMC through the GPIO pin. For example, it can be triggered by an edge (level jump trigger, such as from low to high) or a level (continuous high level trigger) to ensure that the BMC responds in time. After the BMC detects the hardware interrupt signal sent by the CPLD, it immediately starts the data collection process.

[0052] When the basic input / output system reads the real-time trigger signal in the embodiment of the present application, a hardware interrupt signal is issued, thereby triggering a hardware interrupt process, so that the BMC responds to the hardware interrupt signal to synchronously collect the power consumption timing data, temperature timing data, and trigger signal corresponding to the target startup phase, accurately capture the operation data in the startup phase, and provide accurate real-time data for fault diagnosis.

[0053] S203. Based on the data comparison result between the real-time operation data set corresponding to the target startup phase and the reference operation data set, determine whether an abnormal situation occurs in the target startup phase of the basic input / output system.

[0054] In the embodiment of the present application, through the multi-dimensional comparison between the real-time operation data set and the reference operation data set, the abnormalities in the BIOS startup phase are identified, such as numerical abnormalities (such as power consumption and temperature exceeding the standard) and timing abnormalities (such as stage delay and trigger signal misalignment), and then it is determined whether the startup phase is abnormal.

[0055] It should be noted that in the embodiment of the present application, in the fault diagnosis of the server startup phase, the collected trigger signal and temperature timing data are mainly used to assist in judging whether the real-time power consumption timing data is abnormal.

[0056] For the target startup phase, starting from the timestamp of the trigger signal, align the real-time power consumption timing data with the power consumption timing data in the reference operation data set. For example, when the target startup phase is the memory training phase, according to the reference operation data set, the power consumption at the 100th ms of this phase is 60W, then it is necessary to compare whether the real-time power consumption corresponding to the same time point in the real-time operation data set is 60W; and then perform numerical abnormality determination and timing abnormality determination.

[0057] Furthermore, for the determination of numerical abnormalities, the real-time overall machine power consumption timing data can be compared with the reference overall machine power consumption timing data at a single time point, that is, traverse the real-time overall machine power consumption timing data and compare it with the reference overall machine power consumption timing data at the corresponding time point. If the deviation of the real-time data exceeds the preset threshold range of the reference data (such as ±5%), then mark the power consumption timing data corresponding to this time point as an abnormal data point. For example, the power consumption at the 200th ms of the memory training phase in the reference operation data set is 50W, while the power consumption at the 200th ms in the real-time operation data set of the memory training phase is 60W, the determined deviation is 10W, and the deviation from the reference data reaches 10%, which exceeds the preset threshold range of 5%, so it is determined that the power consumption is abnormal.

[0058] For the determination of timing anomalies, the sequence duration of real-time timing data can be compared with that of the reference timing data. If the sequence duration of the real-time timing data exceeds the sequence duration threshold of the reference timing data (e.g., ±5%), it is determined that a timing anomaly has occurred. For example, if the sequence duration of the memory training phase in the reference operation dataset lasts for 500 ms and the real-time operation dataset lasts for 700 ms, the deviation is 40% of the reference sequence duration, which is greater than 5%. Thus, it is determined that the duration of this phase is too long and does not match the normal duration, and then it is judged that a timing anomaly has occurred.

[0059] It is also possible to check whether the triggering timing of the real-time trigger signal is consistent with that of the reference trigger signal. If the deviation between the timestamps of the real-time trigger signal and the reference trigger signal exceeds the threshold (e.g., ±50 ms), it is determined that a timing anomaly has occurred. For example, in the reference operation dataset, the memory training phase should be triggered to start execution 1000 ms after startup, but in the real-time operation dataset, it is triggered to start execution 1200 ms after startup, with a deviation of 200 ms, which is much greater than 50 ms. Then it is determined that there is a trigger delay and a timing anomaly has occurred. This situation may be caused by a delay in BIOS execution or slow operation due to server hardware aging.

[0060] It is also possible to first generate the corresponding real-time power consumption change curve and reference power consumption change curve based on the real-time overall machine power consumption timing data and the reference overall machine power consumption timing data, and then compare the overall trends of the two curves. If the reference power consumption change curve in a certain startup phase shows a trend of rising first and then stabilizing, while the real-time power consumption change curve shows a continuous increase, and the two trends are significantly different, it is judged that a timing anomaly has occurred. Further, the abnormal timing segment can be determined according to the slope deviation between the two curves at the same time point, and the abnormal timing segment can be intercepted for subsequent fault monitoring.

[0061] In the fault diagnosis of the server startup phase, the temperature timing data mainly helps to judge whether the real-time power consumption timing data is abnormal. Referring to the reference temperature curve corresponding to the reference timing temperature data can measure whether the real-time temperature curve corresponding to the real-time timing temperature data is consistent in terms of curve trend. If there is a deviation between the real-time temperature curve and the reference temperature curve, it indirectly reflects that there is an abnormal power consumption or local heat dissipation problem, and it is very likely that there is an abnormal power consumption. Then a fault warning prompt is generated to avoid missing the judgment of the fault. There may also be a sudden change in the real-time temperature, and then the transient high-power consumption situation that is not captured by the power consumption data can be assisted in positioning through the real-time temperature mutation point, thus making up for the deficiency of simply judging by the power consumption threshold.

[0062] Furthermore, through the above judgments of numerical anomalies, timing anomalies, and trend anomalies, with the assistance of timing temperature data, the embodiments of the present application can capture invisible anomalies that cannot be recognized by fixed thresholds. By analyzing the changing trend of data timing, potential hardware risks can be discovered. Compared with the traditional method of judging by fixed thresholds, the embodiments of the present application perform anomaly monitoring on all timing data. Finally, by combining the phase identifier and timing data, periodic anomalies in specific phases are identified to construct a dynamic timing fault monitoring system, improving the comprehensiveness and accuracy of fault monitoring.

[0063] After the above S203 (judging whether there is an abnormal situation in the target startup phase of the basic input / output system) is executed, if there is an abnormal situation in the target startup phase of the basic input / output system, execute the following S204: S204. Generate a fault warning prompt for the target startup phase.

[0064] When it is detected that there is an abnormality in the target startup phase of the BIOS, the system will generate a fault warning prompt corresponding to the target startup phase to prompt the operation and maintenance personnel that there is an abnormality in the target startup phase and inspection and repair are required.

[0065] Specifically, it is possible to intercept pre-failure of the server by sending a fault warning prompt to the operation and maintenance personnel. Furthermore, when the server is started daily (such as restarting or powering on) but the service has not been loaded, hidden hardware defects (such as CPU cache decay) can be detected; subsequent services can be blocked in time to avoid wasting resources and time.

[0066] Through the present application, the startup process of the basic input / output system for starting a computer is divided into startup phases, and a reference operation data set corresponding to each normal startup phase is collected; for the target startup phase, the reference operation data set includes a reference trigger signal, reference overall machine power consumption timing data, and reference chassis temperature timing data; furthermore, during the process of starting the server through the basic input / output system in real time, the real-time operation data set of the target startup phase is collected. Based on the data comparison result between the real-time operation data set and the reference operation data set, and then according to the dynamic matching of the timing data, the pertinence of the target startup phase, and in combination with the trigger time point of the trigger signal, dynamic timing comparison can be accurately performed to identify when numerical anomalies occur in the current target startup phase during the startup process and whether the timing conforms to the law of the reference data; therefore, the present application combines the timing data and trigger signal of the startup process to perform fault monitoring on the target startup process, which can break through the limitation that traditional static thresholds can only identify fixed upper limit faults and avoid the problem that hidden faults are difficult to detect, thereby greatly improving the accuracy and timeliness of the startup fault monitoring of the basic input / output system to ensure the stable operation of the server.

[0067] As an extension and refinement of the above embodiments, refer to Figure 3As shown in the figure, the present application also provides a fault monitoring method, and the specific steps are as follows: S301. During the process of starting the server through the basic input / output system, collect the reference operation data set corresponding to the normal situation in the target startup phase.

[0068] Among them, the reference operation data set includes the reference trigger signal, the reference overall machine power consumption time series data, and the reference chassis temperature time series data in the target startup phase; the reference trigger signal is used to obtain the time for collecting the reference overall machine power consumption time series data and the reference chassis temperature time series data.

[0069] Specifically, since the hardware components of the server will age with the usage time, in order to avoid the deviation of the reference data caused by hardware aging. The reference operation data set corresponding to each startup phase can be corrected periodically according to the hardware aging coefficient of the server, so as to improve the accuracy of fault monitoring.

[0070] Among them, the aging coefficient is used to characterize the attenuation degree of the hardware performance with the usage time, usually expressed by the reference parameter deviation rate. For example, the aging coefficient of the server memory can be obtained by subtracting 1 from the ratio of the current reference power consumption to the initial reference power consumption; if the hardware is not aged, the current reference power consumption and the initial reference power consumption are theoretically equal. At this time, the ratio of the current reference power consumption to the initial reference power consumption is 1, and after subtracting 1, 0 is obtained, indicating that the memory is not aged. Among them, the initial reference power consumption can be obtained from the factory data of the server.

[0071] Specifically, the reference operation data set can be corrected once every preset period (such as every 3 months) to reduce the false alarm / miss rate; furthermore, through dynamic reference correction, the fault diagnosis can adapt to the attenuation of the hardware performance, and the hardware aging coefficient is collected periodically. The values and time series data of the reference operation data set are dynamically corrected according to the startup phase, so that the reference operation data set continuously matches the real-time state of the hardware, ensuring the accuracy of fault monitoring.

[0072] S302. In response to the computer startup instruction issued by the user, receive the hardware interrupt signal.

[0073] Among them, the hardware interrupt signal is issued when the basic input / output system reads the real-time trigger signal.

[0074] S303. In response to the hardware interrupt signal, collect the real-time operation data set corresponding to the target startup phase.

[0075] Among them, the real-time operation data set includes the real-time trigger signal, the real-time overall machine power consumption time series data, and the real-time chassis temperature time series data corresponding to the target startup phase.

[0076] In the embodiment of the present application, the above S303 (collecting the real-time operation data set corresponding to the target startup phase in response to the hardware interrupt signal) can be refined into the following steps 1 to 3: Step 1, in response to the hardware interrupt signal, collect the real-time overall machine power consumption time series data and the real-time chassis temperature time series data corresponding to the target startup phase.

[0077] Specifically, the above-mentioned collection of the real-time overall machine power consumption time series data and the real-time chassis temperature time series data corresponding to the target startup phase may be to read the real-time overall machine power consumption time series data of the power supply unit based on the first bus, and read the real-time chassis temperature time series data collected by the temperature sensor based on the second bus.

[0078] Specifically, the first bus and the second bus may be I2C buses. These two independent hardware buses are used to collect the real-time data of power consumption and temperature in parallel to ensure the timing consistency. The BMC is connected to the PSU through the first bus and is used to read the PSU register to obtain the real-time power consumption time series data; it is connected to the temperature sensor through the second bus and reads the sensor register to obtain the temperature time series data. It should be noted that a clock source is set in the BMC to synchronize the start time of the data collection on the two buses.

[0079] Step 2, receive the phase identifier and the real-time trigger signal corresponding to the target startup phase.

[0080] In the embodiment of the present application, the CPLD sends the phase identifier and the trigger signal to the BMC. The BMC parses the phase identifier to clarify which startup phase the current startup phase is, and then associates all the data corresponding to the startup phase through the phase identifier to generate the real-time operation data set corresponding to the target startup phase, which is convenient for accurately finding the reference operation data set corresponding to the current target startup phase through the phase identifier later, and helps with subsequent data comparison.

[0081] Step 3, associate the real-time overall machine power consumption time series data, the real-time chassis temperature time series data, the phase identifier, and the real-time trigger signal corresponding to the target startup phase to generate a real-time operation data set.

[0082] Furthermore, associate the real-time overall machine power consumption time series data, the real-time chassis temperature time series data, the phase identifier, and the real-time trigger signal corresponding to the target startup phase to obtain a real-time operation data set.

[0083] In the embodiment of the present application, through data association, the real-time overall machine power consumption time series data, the real-time chassis temperature time series data, the phase identifier, and the trigger signal are associated to construct a high-dimensional and structured real-time operation data set. Then, when comparing the real-time operation data set with the reference operation data set later, it is convenient to accurately locate the faulty startup phase, and at the same time, improve the fault monitoring efficiency.

[0084] S304. Extract the real-time key timing features from the real-time operation dataset.

[0085] Specifically, extracting the real-time key timing features from the real-time operation dataset can be understood as screening the real-time key timing features that can represent the operation state of this startup phase from the real-time operation dataset.

[0086] This is because the timing data (such as the power consumption values collected every 10 ms) is a continuous curve, and the efficiency is relatively low when directly comparing. By extracting the real-time key timing features (such as the power consumption rising slope, temperature stability value), complex curves can be described with a small number of key features, reducing the computational amount of subsequent data comparison.

[0087] In some embodiments, the maximum value, minimum value, average value of the power consumption corresponding to each startup phase, the curve slope (such as the rising rate of temperature over time), the fluctuation period (such as the oscillation frequency of the power consumption curve); the mutation feature step point (such as the power consumption suddenly jumps by 5 W at a certain moment), etc. can be extracted as the real-time key timing features.

[0088] S305. Combine the real-time key timing features with the reference key timing features corresponding to the reference operation dataset, compare the real-time operation dataset with the reference operation dataset, and generate a data comparison result.

[0089] In some embodiments, the corresponding reference key timing features can be extracted from the reference operation dataset in advance, and then the extracted real-time key timing features are used to compare with the reference key timing features, so as to quickly determine whether the real-time operation data deviates from the reference operation data, to assist the data comparison process and improve the accuracy and efficiency of data comparison.

[0090] S306. According to the data comparison result, determine whether there is an abnormal situation in the target startup phase of the basic input / output system.

[0091] After the above S306 (According to the data comparison result, determine whether there is an abnormal situation in the target startup phase of the basic input / output system) is executed, if there is an abnormal situation in the target startup phase of the basic input / output system, execute the following S307: S307. Generate a fault warning prompt for the target startup phase.

[0092] In the embodiments of the present application, it is also possible to compare the real-time chassis temperature timing data of the target startup phase with a fixed safety temperature threshold to determine whether there is an abnormal situation in the current target startup phase. The specific steps are as follows: Step A. Compare the real-time chassis temperature timing data with the preset chassis temperature data range.

[0093] Specifically, the preset chassis temperature data range can be understood as a fixed safe temperature threshold range. Then, by comparing the real-time chassis temperature time-series data with the preset chassis temperature data range, it is determined whether the real-time chassis temperature time-series data exceeds the preset chassis temperature data range. When the real-time temperature time-series data exceeds the preset chassis temperature data range, for example, the maximum temperature during the CPU startup phase should not exceed 60°C, but the maximum temperature in the real-time chassis temperature time-series data has reached 70°C, this situation can be directly determined that an abnormality occurs in the target startup phase. It should be noted that different preset chassis temperature data ranges can be divided according to the startup phase. For example, the preset chassis temperature data range during the memory training phase is 30°C to 60°C, and the preset chassis temperature data range during the CPU initialization phase is 40°C to 55°C.

[0094] Step B: When the real-time chassis temperature time-series data exceeds the preset chassis temperature data range, generate a fault warning prompt for the target startup phase.

[0095] Exemplarily, when the preset chassis temperature data range corresponding to the memory training phase is 30°C to 60°C, the highest temperature in the real-time operation data set is 52°C, and the lowest temperature is 32°C, it indicates that the chassis temperature during the current memory training phase is within the safe range and there is no abnormality; on the contrary, if it exceeds the preset chassis temperature data range, it is determined to be abnormal.

[0096] As an extension and refinement of the above embodiments, after responding to the hardware interrupt signal and before collecting the real-time operation data set corresponding to the target startup phase, it is also necessary to lock the rotation speed of the fan in the basic input / output system from the intelligent speed regulation state to the preset speed.

[0097] In some embodiments, since chips such as CPUs and GPUs in the server will trigger frequency reduction at high temperatures, resulting in performance attenuation, and being in an overheated environment for a long time may cause the hardware to age faster. In order to maintain the performance stability of the hardware devices, a fan is set in the chassis to ensure that the hardware works within the rated temperature range.

[0098] Then, during the data collection of the startup phase in the embodiments of the present application, the fluctuation of the fan speed may cause the chassis temperature to change, interfering with the accuracy of the test data. To reduce the interference caused by the fan, by locking the rotation speed of the fan from the intelligent speed regulation state to the preset speed, and then collecting data, the constant heat dissipation efficiency is maintained, and the data error caused by temperature fluctuation is reduced. It should be noted that the pulse width modulation method can be used to adjust the signal duty cycle to control the output power of the fan, that is, to achieve the purpose of controlling the fan speed. For example, the pulse width is set to 50% to adjust the fan speed to the preset speed.

[0099] Then, when pre-collecting the reference operation dataset, in order to maintain the relative stability of the data collection environment, it is also necessary to control the fan speed at the preset speed to reduce environmental variables and ensure the accuracy of the data comparison result.

[0100] It should also be noted that after determining whether there is an abnormal situation in the target startup phase of the basic input / output system based on the data comparison result between the real-time operation dataset corresponding to the target startup phase and the reference operation dataset, if it is determined that there is no abnormal situation in the target startup phase of the basic input / output system, the fan speed in the basic input / output system is restored from being locked to the preset speed to the intelligent speed regulation state.

[0101] Specifically, when it is determined that the target startup phase is completed and there is no abnormal situation during the entire process of the target startup phase, the fan speed can be restored from the preset speed to the intelligent speed regulation state: until the data collection for the next target startup phase is entered, the fan speed in the basic input / output system is locked to the preset speed again from the intelligent speed regulation state. Furthermore, by flexibly adjusting the fan speed, the data collection environment is maintained in a stable state, improving the accuracy of data collection.

[0102] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.

[0103] The embodiments of the present application also provide a fault monitoring device, which corresponds one by one to the method claims. Figure 4 The following is a schematic structural diagram of a data destruction device 400 provided by the present disclosure, as Figure 4 shown, the device 400 of this embodiment includes: A collection unit 41, configured to collect a reference operation dataset corresponding to a normal situation in a target startup phase during the process of starting a server through a basic input / output system; the reference operation dataset includes a reference trigger signal, reference overall machine power consumption timing data, and reference chassis temperature timing data in the target startup phase; the reference trigger signal is used to obtain the time for collecting the reference overall machine power consumption timing data and the reference chassis temperature timing data. An acquisition unit 42, configured to obtain a real-time operation dataset corresponding to the target startup phase in response to a server startup instruction; the real-time operation dataset includes a real-time trigger signal, real-time overall machine power consumption timing data, and real-time chassis temperature timing data corresponding to the target startup phase. A judgment unit 43, configured to judge whether an abnormal situation occurs in the target startup stage of the basic input / output system based on a data comparison result between the real-time operation data set corresponding to the target startup stage and the reference operation data set. A generation unit 44, configured to generate a fault warning prompt for the target startup stage when an abnormal situation occurs in the target startup stage of the basic input / output system.

[0104] As an optional implementation manner of an embodiment of the present application, the obtaining unit 42 is specifically configured to receive a hardware interrupt signal; the hardware interrupt signal is sent when the basic input / output system reads the real-time trigger signal; in response to the hardware interrupt signal, collect the real-time operation data set corresponding to the target startup stage.

[0105] As an optional implementation manner of an embodiment of the present application, the obtaining unit 42 is further configured to insert a stage identifier corresponding to the target startup stage into a code start position corresponding to the target startup stage before the initialization of the target startup stage of the basic input / output system starts.

[0106] As an optional implementation manner of an embodiment of the present application, the obtaining unit 42 is further configured to obtain the stage identifier corresponding to the target startup stage when the basic input / output system reads the stage identifier at the code start position corresponding to the target startup stage, so as to generate a real-time trigger signal.

[0107] As an optional implementation manner of an embodiment of the present application, the obtaining unit 42 is further configured to lock the rotation speed of a fan in the basic input / output system from an intelligent speed regulation state to a preset rotation speed.

[0108] As an optional implementation manner of an embodiment of the present application, the obtaining unit 42 is further configured to restore the rotation speed of the fan in the basic input / output system from being locked to the preset rotation speed to the intelligent speed regulation state when no abnormal situation occurs in the target startup stage of the basic input / output system.

[0109] As an optional implementation manner of an embodiment of the present application, the obtaining unit is specifically configured to, in response to the hardware interrupt signal, collect the real-time overall machine power consumption time series data and the real-time chassis temperature time series data corresponding to the target startup stage; receive the stage identifier corresponding to the target startup stage and the real-time trigger signal; associate the real-time overall machine power consumption time series data, the real-time chassis temperature time series data, the stage identifier, and the real-time trigger signal corresponding to the target startup stage to generate the real-time operation data set.

[0110] As an optional implementation manner of the embodiment of the present application, the obtaining unit 42 is specifically configured to, in response to the hardware interrupt signal, read the real-time overall machine power consumption timing data of the power supply unit based on the first bus, and read the real-time chassis temperature timing data corresponding to the temperature sensor based on the second bus.

[0111] As an optional implementation manner of the embodiment of the present application, the judging unit 43 is specifically configured to extract the real-time key timing features in the real-time operation data set; compare the real-time operation data set with the reference operation data set by combining the real-time key timing features with the reference key timing features corresponding to the reference operation data set to generate the data comparison result; and judge whether an abnormal situation occurs in the target startup phase of the basic input / output system according to the data comparison result.

[0112] As an optional implementation manner of the embodiment of the present application, the collecting unit 41 is further configured to periodically correct the reference operation data set corresponding to the target startup phase according to the hardware aging coefficient of the server.

[0113] As an optional implementation manner of the embodiment of the present application, the judging unit 43 is further configured to compare the real-time chassis temperature timing data with a preset chassis temperature data range; and when the real-time chassis temperature timing data exceeds the preset chassis temperature data range, generate a fault warning prompt for the target startup phase.

[0114] For the description of the features in the corresponding embodiment of the fault monitoring device, reference may be made to the relevant description in the corresponding embodiment of the fault monitoring method, which will not be elaborated here one by one.

[0115] The embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above-mentioned embodiments of the fault monitoring method.

[0116] The embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any one of the above-mentioned embodiments of the fault monitoring method when running.

[0117] In an exemplary embodiment, the above-mentioned computer-readable storage medium may include, but is not limited to: various media such as a USB flash drive, a read-only memory (ROM for short), a random access memory (RAM for short), a mobile hard disk, a magnetic disk, or an optical disc that can store a computer program.

[0118] Embodiments of the present application also provide a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described embodiments of the fault monitoring method.

[0119] Embodiments of the present application also provide another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described embodiments of the fault monitoring method.

[0120] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0121] The above has introduced in detail a fault monitoring method and device provided by the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

Claims

1. A fault monitoring method, characterized in that, Applied to a baseboard management controller, including: During the process of starting the server through the basic input / output system, collect the reference operation data set corresponding to the target startup phase under normal circumstances; the reference operation data set includes the reference trigger signal, reference overall machine power consumption timing data, and reference chassis temperature timing data of the target startup phase; the reference trigger signal is used to obtain the time for collecting the reference overall machine power consumption timing data and reference chassis temperature timing data; In response to the server startup instruction, obtain the real-time operation data set corresponding to the target startup phase; the real-time operation data set includes the real-time trigger signal, real-time overall machine power consumption timing data, and real-time chassis temperature timing data corresponding to the target startup phase; Based on the data comparison result between the real-time operation data set and the reference operation data set corresponding to the target startup phase, determine whether an abnormal situation occurs in the target startup phase of the basic input / output system; If so, generate a fault warning prompt for the target startup phase.

2. The method according to claim 1, wherein The obtaining the real-time operation data set corresponding to the target startup phase; includes: Receive a hardware interrupt signal; the hardware interrupt signal is sent when the basic input / output system reads the real-time trigger signal; In response to the hardware interrupt signal, collect the real-time operation data set corresponding to the target startup phase.

3. The method according to claim 1, wherein Before obtaining the real-time operation data set corresponding to the target startup phase; the method further includes: Before the initialization of the target startup phase of the basic input / output system starts, insert the phase identifier corresponding to the target startup phase into the start position of the code corresponding to the target startup phase.

4. The method according to claim 3, characterized in that, The method further includes: When the basic input / output system reads the phase identifier at the start position of the code corresponding to the target startup phase, obtain the phase identifier corresponding to the target startup phase to generate a real-time trigger signal.

5. The method according to claim 2, wherein After responding to the hardware interrupt signal, the method further includes: Lock the rotation speed of the fan in the basic input / output system from the intelligent speed regulation state to a preset speed.

6. The method according to claim 5, wherein The method further includes: If no abnormal situation occurs in the target startup phase of the basic input / output system, restore the rotation speed of the fan in the basic input / output system from being locked to the preset speed to the intelligent speed regulation state.

7. The method according to claim 2, wherein The responding to the hardware interrupt signal and collecting the real-time operation data set corresponding to the target startup phase; includes: In response to the hardware interrupt signal, collect the real-time overall machine power consumption timing data and real-time chassis temperature timing data corresponding to the target startup phase; Receive the phase identifier and the real-time trigger signal corresponding to the target startup phase; Associate the real-time overall machine power consumption timing data, real-time chassis temperature timing data, phase identifier, and real-time trigger signal corresponding to the target startup phase to generate the real-time operation data set.

8. The method according to claim 7, characterized in that, The responding to the hardware interrupt signal and collecting the real-time overall machine power consumption timing data and real-time chassis temperature timing data corresponding to the target startup phase; includes: In response to the hardware interrupt signal, read the real-time overall machine power consumption timing data of the power supply unit based on the first bus, and read the real-time chassis temperature timing data corresponding to the temperature sensor based on the second bus.

9. The method according to claim 1, characterized in that Judging whether an abnormal situation occurs in the target startup stage of the basic input / output system based on the data comparison result between the real-time operation data set corresponding to the target startup stage and the reference operation data set includes: Extract the real-time key timing features in the real-time operation data set; Combine the real-time key timing features with the reference key timing features corresponding to the reference operation data set, compare the real-time operation data set with the reference operation data set, and generate the data comparison result; Judge whether an abnormal situation occurs in the target startup stage of the basic input / output system according to the data comparison result.

10. The method according to claim 1, characterized in that The method further includes: Periodically correct the reference operation data set corresponding to the target startup stage according to the hardware aging coefficient of the server.

11. The method according to claim 1, characterized in that, The method further includes: Compare the real-time chassis temperature timing data with a preset chassis temperature data range; When the real-time chassis temperature timing data exceeds the preset chassis temperature data range, generate a fault warning prompt for the target startup stage.

12. A fault monitoring device, characterized in that, It includes: An acquisition unit for acquiring a reference operation data set corresponding to a normal situation in a target startup stage during the startup of a server through a basic input / output system; The reference operation data set includes a reference trigger signal, reference overall machine power consumption timing data, and reference chassis temperature timing data in the target startup stage; The reference trigger signal is used to obtain the time for acquiring the reference overall machine power consumption timing data and the reference chassis temperature timing data; An acquisition unit for acquiring a real-time operation data set corresponding to the target startup stage in response to a server startup instruction; The real-time operation data set includes a real-time trigger signal, real-time overall machine power consumption timing data, and real-time chassis temperature timing data corresponding to the target startup stage; A judgment unit for judging whether an abnormal situation occurs in the target startup stage of the basic input / output system based on the data comparison result between the real-time operation data set corresponding to the target startup stage and the reference operation data set; A generation unit for generating a fault warning prompt for the target startup stage when an abnormal situation occurs in the target startup stage of the basic input / output system.

13. An electronic device, characterized in that, It includes: A memory for storing a computer program; A processor for implementing the steps of the fault monitoring method according to any one of claims 1 to 11 when executing the computer program.

14. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program implements the steps of the fault monitoring method according to any one of claims 1 to 11 when executed by a processor.

15. A computer program product, comprising a computer program, characterized in that, The computer program implements the steps of the fault monitoring method according to any one of claims 1 to 11 when executed by a processor.

Citation Information

Patent Citations

  • BIOS (basic input / output system) program startup monitoring method and electronic device

    CN107992397A

  • Substrate management controller adaptation method of basic input / output system and program product

    CN119201225A

  • Metric-based anomaly detection system with evolving mechanism in large-scale cloud

    US20200142763A1