Fault-recoverable firmware detection system and method, storage medium, and server

By obtaining fault information in the server and analyzing it, combined with the heartbeat request mechanism of the virtual external device, accurately detecting whether the fault can cause the operating system to crash, solving the uncertainty problem of fault detection in the prior art.

WO2025123552A1PCT designated stage expired Publication Date: 2025-06-19INSPUR SUZHOU INTELLIGENT TECH CO LTD

Patent Information

Application Number
PCT/CN2024/089627
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-13
Filing Date
2024-04-24
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

When detecting recoverable faults, there are problems of fault missed and fault false alarms, and it is difficult to accurately determine whether recovery faults will cause the server operating system to go down.

Method used

The target fault information in the fault register is obtained through the basic input and output system BIOS and sent to the substrate management controller BMC. The BMC analyzes the received fault information, judges its fault type, and when it is determined that it can be recovered, it sends a heartbeat request to the server operating system through a virtual external device to monitor its response to determine whether it is down.

Benefits of technology

Accurate detection of whether recoverable faults cause operating system downtime, reduce fault missed and false alarms, and promptly remind users to deal with them.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024089627_19062025_PF_FP_ABST
    Figure CN2024089627_19062025_PF_FP_ABST
Patent Text Reader

Abstract

A fault-recoverable firmware detection system and method, a storage medium, and a server. The method comprises: acquiring target fault information stored in a fault register, and sending the target fault information to a BMC (201); upon receiving the target fault information sent by a BIOS, the BMC analyzing the target fault information, and determining a fault type (202); when the fault type of the target fault information is a recoverable fault, the BMC controlling a virtual external device to send a target heartbeat request to a server operating system according to a preset sending period, and receiving response data fed back by the server operating system on the basis of the target heartbeat request (203); and within a preset duration during which the BMC sends the target heartbeat request to the server operating system, if an interruption occurs in the response data fed back by the operating system, determining that the server operating system is crashed, and on the basis of the target fault information, determining a faulty component in a server (204).
Need to check novelty before this filing date? Find Prior Art

Description

Fault-recoverable firmware detection system, method, storage medium, and server

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to the Chinese patent application filed with the China Patent Office on December 13, 2023, with application number 202311709048.5, and entitled “Recoverable Fault Firmware Detection System, Method, Storage Medium and Server,” the entire contents of which are incorporated herein by reference. Technical Field

[0003] The present application relates to the field of computer technology, and in particular to a firmware detection system, method, storage medium, and server capable of recovering faults. Background Art

[0004] RAS-based servers have mature and reliable detection methods for both catastrophic and fatal failures. The server firmware can monitor the status of the corresponding fault personal identification number (PIN) signal to determine the type of failure that has occurred in the current server system.

[0005] However, even if the server supports a corresponding PIN signal to indicate the occurrence of a recoverable fault, there is a high degree of uncertainty as to whether the recoverable fault will cause the server operating system to crash. If a recoverable fault causes a crash and the firmware does not detect it, it will result in a missed fault report. If a recoverable fault does not cause a crash but the firmware reports the fault, it will result in a false fault report.

[0006] Based on this, there is an urgent need for a firmware detection method for recoverable faults to accurately detect whether a recoverable fault will cause the operating system to crash, and to promptly remind the user to handle it when the operating system crashes.

[0007] Summary of the Invention

[0008] In a first aspect, the present application provides a method for detecting a recoverable fault in firmware, comprising:

[0009] The basic input and output system BIOS obtains the target fault information stored in the fault register and sends the target fault information to the baseboard management controller BMC; when the baseboard management controller BMC receives the target fault information sent by the basic input and output system BIOS, the baseboard management controller BMC parses the target fault information and determines the fault type of the target fault information; when the fault type of the target fault information is a recoverable fault, the baseboard management controller BMC controls the virtual external device to send a target heartbeat request to the server operating system according to a preset sending cycle, and receives response data fed back by the server operating system based on the target heartbeat request; within the preset time period of the baseboard management controller sending the target heartbeat request to the server operating system, if the baseboard management controller BMC detects that the response data fed back by the operating system is interrupted, it determines that the server operating system is down, and determines the faulty component in the server based on the target fault information; wherein, the target fault information is generated when the memory and PCIe device in the server fail; the virtual external device is a universal serial bus USB device virtualized by the baseboard management controller BMC.

[0010] Optionally, the fault register includes: a UNCERRSTS register, a DEVSTS register, and a STATUS register; the basic input and output system BIOS obtains the target fault information stored in the fault register, including: the basic input and output system BIOS detects the fault information generated in the DEVSTS register according to a preset detection cycle, and / or, the fault information generated in the STATUS register, and / or, the fault information generated in the UNCERRSTS register.

[0011] Optionally, the target fault information is fault information generated when a memory fault occurs; when the baseboard management controller BMC receives the target fault information sent by the basic input and output system BIOS, it parses the target fault information and determines the fault type of the target fault information, including: the baseboard management controller BMC parses the target fault information, and when the field included in the target fault information meets the first preset rule, determines that the fault type of the target fault information is a recoverable fault; otherwise, determines that the fault type of the target fault information is an unrecoverable fault.

[0012] Optionally, the baseboard management controller BMC parses the target fault information, and when the fields included in the target fault information meet the first preset rule, determines that the fault type of the target fault information is a recoverable fault, including: the baseboard management controller BMC parses the target fault information, and when the target fault information indicates that a fault is recorded in the STATUS register and the target fault information contains a preset field, determines that the fault type of the target fault information is a recoverable fault.

[0013] Optionally, the target fault information is fault information generated when a PCIe device fails; when the baseboard management controller BMC receives the target fault information sent by the basic input and output system BIOS, it parses the target fault information and determines the fault type of the target fault information, including: the baseboard management controller BMC parses the target fault information, and when the field included in the target fault information meets the second preset rule, determines that the fault type of the target fault information is a recoverable fault.

[0014] Optionally, when the fields contained in the target fault information meet the second preset rule, the fault type of the target fault information is determined to be a recoverable fault, including: when the baseboard management controller BMC records an unrecoverable fault in the UNCERRSTS register and a non-fatal fault is recorded in the DEVSTS register, the fault type of the target fault information is determined to be a recoverable fault.

[0015] Optionally, the system also includes: a platform path controller PCH; a baseboard management controller BMC controls the virtual external device to send a target heartbeat request to the server operating system according to a preset sending period, and receives response data fed back by the server operating system based on the target heartbeat request, including: the baseboard management controller BMC controls the virtual external device to send a target heartbeat request to the platform path controller PCH according to a preset sending period, and receives response data fed back by the platform path controller PCH based on the target heartbeat request.

[0016] Optionally, the target heartbeat request is a HID_GET_REPORT request constructed using the bmRequest field; the baseboard management controller BMC controls the virtual peripheral device to send the target heartbeat request to the server operating system according to a preset sending period, including: the baseboard management controller BMC performs an initialization operation, and after the initialization operation is completed, creates a downtime status detection task process; the initialization operation includes: initializing related library functions and configuring the system clock; the baseboard management controller BMC executes the downtime status detection task process, calls the USBD_LL_SetupStage function to process the SETUP stage of the virtual peripheral device, and calls the USBD_LL_DataInStage function to process During the IN phase of the virtual external device, the USBD_LL_DataOutStage function is called to process the OUT phase of the virtual external device; when the baseboard management controller BMC determines that the virtual external device is configured successfully, it constructs the relevant parameters of the HID_GET_REPORT request; after completing the construction of the HID_GET_REPORT request, the baseboard management controller BMC calls the usb_control_msg function to send the HID_GET_REPORT request; the baseboard management controller BMC calls the USBD_HID_GetReport function in the peripheral interrupt processing function to receive the response data fed back by the server operating system based on the target heartbeat request.

[0017] Optionally, the baseboard management controller BMC determines the faulty component in the server based on the target fault information, including: when the baseboard management controller BMC determines that the server operating system is down, it analyzes the target fault information, determines the faulty component, and records the downtime event and the faulty component in the system environment log SEL to prompt the user of the diagnosed faulty component.

[0018] In a second aspect, the present application further provides a firmware detection system capable of recovering faults, comprising:

[0019] A basic input / output system (BIOS), a server operating system, a fault register set in a central processing unit (CPU), a baseboard management controller (BMC), memory, and a high-speed serial computer expansion bus (PCIe) standard device; the basic input / output system (BIOS) is used to obtain target fault information stored in the fault register and send the target fault information to the baseboard management controller (BMC); the baseboard management controller (BMC) is used to parse the target fault information and determine the fault type of the target fault information upon receiving the target fault information sent by the basic input / output system (BIOS); the baseboard management controller (BMC) is further used to control a virtual external device to send a target heartbeat request to the server operating system according to a preset sending period, and receive response data fed back by the server operating system based on the target heartbeat request, if the response data fed back by the server operating system is interrupted within a preset duration of the baseboard management controller sending the target heartbeat request to the server operating system, and determine the faulty component in the server based on the target fault information; wherein the target fault information is generated when a fault occurs in the memory and PCIe device in the server; and the virtual external device is a universal serial bus (USB) device virtualized by the baseboard management controller (BMC).

[0020] Optionally, the fault register includes: an UNCERRSTS register, a DEVSTS register, and a STATUS register; a basic input and output system BIOS, which is specifically used to detect the fault information generated in the DEVSTS register according to a preset detection cycle, and / or the fault information generated in the STATUS register, and / or the fault information generated in the UNCERRSTS register.

[0021] Optionally, the target fault information is fault information generated when a memory fault occurs; the baseboard management controller BMC is specifically used to parse the target fault information, and when the fields contained in the target fault information meet the first preset rule, determine that the fault type of the target fault information is a recoverable fault; otherwise, determine that the fault type of the target fault information is an unrecoverable fault.

[0022] Optionally, the baseboard management controller BMC is specifically configured to parse the target fault information, and determine that the fault type of the target fault information is a recoverable fault when the target fault information indicates that a fault is recorded in the STATUS register and the target fault information includes a preset field.

[0023] Optionally, the target fault information is fault information generated when a PCIe device fails; the baseboard management controller BMC is specifically used to parse the target fault information, and when the fields contained in the target fault information meet the second preset rule, determine that the fault type of the target fault information is a recoverable fault.

[0024] Optionally, the baseboard management controller BMC is specifically configured to determine that the fault type of the target fault information is a recoverable fault when the target fault information indicates that an unrecoverable fault is recorded in the UNCERRSTS register and a non-fatal fault is recorded in the DEVSTS register.

[0025] Optionally, the system also includes: a platform path controller PCH; a baseboard management controller BMC, specifically used to control the virtual external device to send a target heartbeat request to the platform path controller PCH according to a preset sending cycle, and receive response data feedback from the platform path controller PCH based on the target heartbeat request.

[0026] Optionally, the target heartbeat request is a HID_GET_REPORT request constructed using the bmRequest field; the baseboard management controller BMC is specifically used to perform an initialization operation and, after the initialization operation is completed, create a downtime status detection task process; the initialization operation includes: initializing related library functions and configuring the system clock; the baseboard management controller BMC is also specifically used to execute a downtime status detection task process, call the USBD_LL_SetupStage function to process the SETUP stage of the virtual external device, call the USBD_LL_DataInStage function to process the IN stage of the virtual external device, and call the USBD_LL_Data The OutStage function processes the OUT stage of the virtual external device; the baseboard management controller BMC is specifically used to construct the relevant parameters of the HID_GET_REPORT request when it is determined that the virtual external device is configured successfully; the baseboard management controller BMC is specifically used to call the usb_control_msg function to send the HID_GET_REPORT request after completing the construction of the HID_GET_REPORT request; the baseboard management controller BMC is specifically used to call the USBD_HID_GetReport function in the peripheral interrupt processing function to receive the response data of the server operating system based on the target heartbeat request feedback.

[0027] Optionally, the baseboard management controller BMC is specifically used to analyze the target fault information, determine the faulty component, and record the downtime event and the faulty component in the system environment log SEL to prompt the user of the diagnosed faulty component when determining that the server operating system is down.

[0028] In a third aspect, the present application further provides a computer-readable instruction product, comprising computer-readable instructions, which, when executed by a processor, implement the steps of any of the above-mentioned methods for detecting recoverable faults in firmware.

[0029] In a fourth aspect, the present application further provides a server on which is provided the recoverable fault firmware detection system according to any one of the above-mentioned second aspects.

[0030] In a fifth aspect, the present application also provides a non-volatile computer-readable storage medium having computer-readable instructions stored thereon, which, when executed by a processor, implement the steps of any of the recoverable fault firmware detection methods in the first aspect above. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to more clearly illustrate the technical solutions in this application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0032] FIG1 is a schematic diagram of the structure of a recoverable fault firmware detection system provided in one or more embodiments of the present application;

[0033] FIG2 is a flow chart of a method for detecting a recoverable fault in firmware provided in one or more embodiments of the present application;

[0034] FIG3 is a flowchart of a downtime detection task program provided in one or more embodiments of the present application. DETAILED DESCRIPTION

[0035] To make the objectives, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this application.

[0036] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of the same type, and do not limit the number of objects; for example, the first object can be one or more. In addition, the term "and / or" in the specification and claims refers to at least one of the connected objects, and the character " / " generally indicates that the objects connected are in an "or" relationship.

[0037] The following describes the professional terms involved in the embodiments of this application:

[0038] RAS: It is the abbreviation of Reliability, Availability, and Serviceability, which is a requirement for a server to be used reliably. RAS architecture refers to the system architecture designed to meet this requirement. RAS architecture usually includes the following aspects: Reliability (Reliability): The system must run reliably for as long as possible without downtime to reduce system downtime. Availability (Availability): The system must be able to provide output capabilities, even some minor errors can be self-repaired, and errors that cannot be self-repaired should be isolated as much as possible to ensure the normal operation of the rest of the system. Serviceability (Serviceability): The system must provide a hardware detection and reporting mechanism so that the administrator can be notified to replace the hardware in time before the hardware error causes data loss or downtime; provide a hardware error recovery mechanism, and correct errors as much as possible to ensure that the system can run sustainably and reliably. The design goal of the RAS architecture is to improve the reliability, availability, and serviceability of the system, thereby improving the stability and security of the system.

[0039] The Basic Input / Output System (BIOS) is an industry-standard firmware interface. The BIOS is the first software loaded when a computer starts up. Essentially, it's a set of programs hardwired into a read-only memory (ROM) chip on the computer's motherboard. It stores the computer's most important basic input / output (BIO) routines, post-boot self-test routines, and system startup routines. It can read and write detailed system configuration information from the complementary metal oxide semiconductor (CMOS) memory. Its primary function is to provide the lowest-level, most direct hardware configuration and control for the computer. The BIOS also provides system parameters to the operating system.

[0040] The Baseboard Management Controller (BMC) is a core component used for server deployment, diagnosis, and management. The BMC manages the interface between system management software and platform management hardware, providing autonomous monitoring, event logging, and recovery control. The BMC also collects and manages information from all hardware and operating systems on the server, providing this information to upper-level operations and network management software.

[0041] Peripheral Component Interconnect Express (PCIe): A high-speed serial computer expansion bus standard for connecting high-speed components. Every computer motherboard has multiple PCIe slots, which can be used to add graphics processing units (GPUs), RAID (Redundant Arrays of Independent Disks) cards, Wi-Fi cards, or solid-state drive (SSD) expansion cards. These devices are collectively referred to as PCIe devices.

[0042] Platform Controller Hub (PCH): A key component of the motherboard chipset, it is typically located below the motherboard, away from the CPU slot and in front of the PCI slots. Its primary functions include controlling communication between various peripheral devices and the motherboard, such as the PCI bus, USB, Serial Advanced Technology Attachment (SATA), audio controller, keyboard controller, real-time clock controller, and advanced power management. It manages computer input and output interfaces, such as USB, audio, and network cards. It provides hard drive control, storage data transfer, and other functions through the SATA interface. It manages the BIOS chip on the motherboard to ensure the system boots and operates normally.

[0043] The UNCERRSTS register is used to record PCIe bus error information. If a PCIe bus error occurs, such as a data transmission error, protocol error, or data checksum error, the UNCERRSTS register records the corresponding error flag to facilitate error detection and handling.

[0044] The DEVSTS register is a register used to record CPU status. It's a register in the Intel x86 architecture that records device status information. In the Intel x86 architecture, the DEVSTS register typically records device status information, such as whether the device is in an interrupt state or has experienced an exception. If a fault is recorded in the DEVSTS register, it indicates that the device encountered an error while executing instructions, potentially causing a system crash or other issues. Further investigation and repair are required.

[0045] The STATUS register is a register used to record CPU status. It can record various status information, such as whether the CPU is in the interrupt state, whether an exception has occurred, whether there is a carry, whether there is an overflow, whether the result is zero, whether the result is negative, etc. Different CPU architectures may have different STATUS registers, such as RISC-V, ARM, and x86. The contents of the STATUS register can be used to control program flow, such as jumping or branching based on condition codes.

[0046] Human Interface Device (HID): A standard for computer devices typically used by humans to operate and control computer systems. HID devices include keyboards, mice, game controllers, cameras, touchscreens, and more. The HID standard makes these devices compatible with any operating system and application without the need for additional software or drivers. The most common HID standard refers to the USB HID specification, which defines a protocol for transmitting data and commands for HID devices.

[0047] A bmRequest request is a type of request in the USB protocol used to send control commands to a USB device. A bmRequest request is typically specified by the bmRequestType and bmRequest fields in the SETUP packet. The bmRequestType field specifies the type of request, such as the request type, recipient type, and transfer direction; the bmRequest field specifies the specific request type, such as obtaining a device descriptor or setting an endpoint.

[0048] HID_GET_REPORT request: is a request type in the USB protocol, used to obtain reports from HID devices. HID devices are human-computer interaction devices, such as keyboards, mice, game controllers, etc. The HID_GET_REPORT request is usually specified by the REPORT descriptor in the HID descriptor returned by the GET_DESCRIPTOR request. The REPORT descriptor contains information about the input, output, and characteristic reports of the HID device, where input reports are used to send data to the host, output reports are used to receive data from the host, and characteristic reports are used to read or set the status information of the device. If the host sends a HID_GET_REPORT request to the HID device, the HID device will return report data of the specified type.

[0049] To address the aforementioned technical issues in related technologies, embodiments of the present application provide a firmware detection system for recoverable faults that can be detected through firmware. FIG1 shows a firmware detection system for recoverable faults provided by an embodiment of the present application. The system comprises: a basic input / output system (BIOS), a server operating system, a fault register in a central processing unit (CPU), a baseboard management controller (BMC), memory, and a high-speed serial computer expansion bus (PCIe) device.

[0050] Based on the system shown in Figure 1, the firmware detection method for recoverable faults provided in the embodiment of the present application includes: 1. Using the BMC chip for HID communication detection: By using the functions provided by the BMC chip, HID communication detection with the server operating system is realized. 2. Server system fault processing interrupt: When a recoverable fault occurs in the server system, the BIOS enters the fault processing interrupt, collects the fault register information and passes it to the BMC. 3. Parsing the fault information: After receiving the fault register information, the BMC parses the fault type and the faulty component. 4. Downtime status detection task: The BMC starts the downtime status detection task process, and intermittently uses the virtual USB keyboard or mouse device of the BMC chip to send a bmRequest request to the server operating system to detect whether the system can respond normally, and then determine whether the server operating system is down.

[0051] For example, in an embodiment of the present application, the fault information may be generated when a fault occurs in a memory or PCIe device in a server, and the data format of the fault information from different sources is different.

[0052] For example, memory fault types can be divided into three categories: uncorrected NO Action Required (UCNA), software recoverable action required (SRAR), and software recoverable action optional (SRAO).

[0053] For example, in the fault information shown below: the STATUS register of CPU0 Core17 Bank1, i.e., the DCU bank, records a fault. "Bit (61) UC Valid" indicates that an unrecoverable fault has occurred. The address resolution of "MCi_ADDR" points to the memory CPU0_Channel2_Dimm1, i.e., a UCE fault has occurred at this address of the memory. However, since "Bit (56) Signals Valid" and "Bit (55) AR Valid" are both set, it indicates that an SRAR type fault has occurred. If the fault is successfully repaired at the operating system level (including discarding the fault address data, shutting down the program that caused the fault, etc.), the fault will not cause a system downtime. If the operating system fails to repair the fault, the fault will continue to cause a system downtime.

[0054] For example, a recoverable fault in a PCIe device corresponds to a fatal fault and is typically recorded as a non-fatal fault. In the following fault information: the device-side UNCERRSTS register records an unrecoverable fault, "Completion time out error," indicating a data transmission timeout. However, the device-side DEVSTS register also records a "Non-Fatal Error detected," allowing the device to retransmit the timed-out data to the operating system host to repair the fault. If retransmitting the data still fails to repair the fault, the operating system may crash.

[0055] For example, since the firmware layer cannot perceive whether the recoverable faults of PCIe devices and memory can be successfully repaired at the operating system level, the firmware detection method for recoverable faults provided in the embodiment of the present application can be used to determine whether the operating system has crashed after a recoverable fault occurs.

[0056] The following describes in detail the firmware detection method for recoverable faults provided by the embodiments of the present application through specific embodiments and their application scenarios in conjunction with the accompanying drawings.

[0057] Based on FIG1 , as shown in FIG2 , a firmware detection method for recoverable faults provided in an embodiment of the present application includes the following steps 201 to 204:

[0058] Step 201: The basic input / output system (BIOS) obtains target fault information stored in a fault register, and sends the target fault information to a baseboard management controller (BMC).

[0059] The target fault information is generated when a fault occurs in a memory or PCIe device in the server.

[0060] Exemplarily, the above-mentioned fault registers may include: an UNCERRSTS register, a DEVSTS register, and a STATUS register; the functions of each register have been introduced in detail in the above content and will not be repeated here.

[0061] For example, when a memory or PCIe device fails, fault information is generated and stored in the fault register. After a server system failure occurs, the basic input / output system (BIOS) can execute a fault processing interrupt, obtain the fault information from the fault register, and send the obtained fault information to the baseboard management controller (BMC), which then analyzes the fault information.

[0062] Specifically, the above step 201 may include the following steps 201a:

[0063] Step 201a: The basic input / output system BIOS detects the fault information generated in the DEVSTS register and / or the fault information generated in the STATUS register and / or the fault information generated in the UNCERRSTS register according to a preset detection period.

[0064] Exemplarily, the target fault information is the fault information obtained by the basic input and output system BIOS from any of the three registers mentioned above. It should be noted that the fault register in the embodiment of the present application may also include other registers capable of storing fault information.

[0065] Step 202 : Upon receiving target fault information sent by the basic input / output system (BIOS), the baseboard management controller (BMC) parses the target fault information and determines the fault type of the target fault information.

[0066] For example, after receiving the target fault information sent by the basic input and output system BIOS, the baseboard management controller BMC may parse the target fault information and determine the fault type of the target fault information.

[0067] Specifically, if the target fault information is fault information generated when a memory fault occurs, the above step 202 may include the following steps 202a:

[0068] Step 202a: The baseboard management controller (BMC) parses the target fault information and determines the fault type of the target fault information as recoverable if the fields included in the target fault information meet the first preset rule; otherwise, determines the fault type of the target fault information as unrecoverable.

[0069] Exemplarily, the first preset rule is used to determine whether the fault type of the target fault information is a recoverable fault when the target fault information is fault information generated when a memory fault occurs.

[0070] Specifically, the above step 202a may further include the following step 202a1:

[0071] Step 202a1: The baseboard management controller BMC parses the target fault information. If the target fault information indicates that a fault is recorded in the STATUS register and the target fault information contains a preset field, the baseboard management controller BMC determines that the fault type of the target fault information is a recoverable fault.

[0072] For example, the preset fields in the above step 202a1 may be the "Bit (56) Signals Valid" and "Bit (55) AR Valid" fields in the above example.

[0073] It can be understood that based on the judgment of fault information in the above example, the fact that the fault information records a field belonging to an unrepairable fault does not mean that the fault type of the fault information is an unrecoverable fault. A comprehensive judgment is required based on other fields recorded in the fault information.

[0074] Specifically, if the target fault information is fault information generated when a PCIe device fails, the above step 202 may include the following step 202b:

[0075] Step 202b: The baseboard management controller BMC parses the target fault information, and determines that the fault type of the target fault information is a recoverable fault if the fields included in the target fault information meet the second preset rule.

[0076] Exemplarily, similar to the first preset rule, the second preset rule is used to determine whether the fault type of the target fault information is a recoverable fault when the target fault information is fault information generated when a PCIe device fails.

[0077] Specifically, based on the above description of the example of a recoverable fault of a PCIe device, the above step 202b may further include the following step 202b1:

[0078] Step 202b1: When the baseboard management controller BMC records an unrecoverable fault in the UNCERRSTS register and a non-fatal fault in the DEVSTS register, it is determined that the fault type of the target fault information is a recoverable fault.

[0079] Step 203 : When the fault type of the target fault information is a recoverable fault, the baseboard management controller (BMC) controls the virtual peripheral device to send a target heartbeat request to the server operating system according to a preset sending period, and receives response data fed back by the server operating system based on the target heartbeat request.

[0080] The virtual external device is a universal serial bus (USB) device virtualized by the baseboard management controller (BMC).

[0081] For example, when the baseboard management controller BMC determines that the fault type of the target fault information is a recoverable fault, it may start a downtime status detection task process to monitor the status of the server operating system.

[0082] Exemplarily, the target heartbeat request is a HID_GET_REPORT request constructed using the bmRequest field.

[0083] Specifically, the step of the baseboard management controller BMC controlling the virtual peripheral device to send a target heartbeat request to the server operating system according to a preset sending period in step 203 may include the following steps 203a1 to 203a5:

[0084] Step 203a1: The baseboard management controller BMC performs an initialization operation, and after the initialization operation is completed, creates a downtime status detection task process.

[0085] The initialization operation includes: initializing related library functions and configuring the system clock.

[0086] Step 203a2: The baseboard management controller BMC executes the downtime detection task process, calls the USBD_LL_SetupStage function to process the SETUP stage of the virtual external device, calls the USBD_LL_DataInStage function to process the IN stage of the virtual external device, and calls the USBD_LL_DataOutStage function to process the OUT stage of the virtual external device.

[0087] Step 203a3: When the baseboard management controller BMC determines that the virtual peripheral device is configured successfully, it constructs relevant parameters of the HID_GET_REPORT request.

[0088] Step 203a4: After completing the construction of the HID_GET_REPORT request, the baseboard management controller BMC calls the usb_control_msg function to send the HID_GET_REPORT request.

[0089] Step 203a5: The baseboard management controller BMC calls the USBD_HID_GetReport function in the peripheral interrupt processing function to receive the response data fed back by the server operating system based on the target heartbeat request.

[0090] Exemplarily, the virtual external device is a USB device virtualized by a baseboard management controller (BMC).

[0091] It should be noted that in USB communications, HID devices use control transfers to communicate and interact with the host. In control transfers of HID devices, bmRequest is an 8-bit field used to specify the transfer type and request type. It is located in the low byte of the bmRequestType field of the USB request. The upper two bits of bmRequest indicate the transfer type, and its binary values ​​include: 00, 01, and 10. 00 indicates standard transfer, which is used for standard USB requests; 01 indicates class transfer, which is used for requests related to the device class; and 10 indicates vendor transfer, which is used for requests related to a specific vendor. The lower six bits of bmRequest indicate the request type, and the specific value depends on the transfer type. For standard transfers and class transfers, the request type value is defined by the USB specification. For vendor transfers, the request type value is defined by the device manufacturer.

[0092] For example, the code for controlling the virtual peripheral device to send a target heartbeat request to the server operating system according to a preset sending period is as follows:

[0093] VENDOR_ID and PRODUCT_ID can be replaced with the actual vendor and product IDs of the USB device being operated. This code uses the HID bmRequest field to send a HID_GET_REPORT request to the server operating system to receive a data packet from the server operating system. The code then monitors the data packet to determine if the server operating system has crashed.

[0094] For example, the implementation logic of the above code includes: 1. Initializing the USB and HID libraries of the BMC chip and configuring the system clock, etc. 2. Starting the USB device. 3. Creating a process and establishing an infinite loop to handle USB events. 4. Within the process loop, first call the USBD_LL_SetupStage function to handle the USB device's SETUP phase. 5. Then call the USBD_LL_DataInStage function to handle the USB device's IN phase. 6. Next, call the USBD_LL_DataOutStage function to handle the USB device's OUT phase. 7. Check whether the USB device status has been successfully configured. 8. If the USB device has been successfully configured, construct the relevant parameters for the HID_GET_REPORT request. 9. Call the usb_control_msg function to send the HID_GET_REPORT request. 10. In the BMC chip's peripheral interrupt handler, call the USBD_HID_GetReport function to process the system's returned data. 11. Repeat steps 4 to 10 above, continuously sending HID_GET_REPORT requests and checking whether the system has returned any data.

[0095] In a possible implementation, the system further includes a platform path controller PCH. The above step 203 may further include the following step 203b:

[0096] Step 203b: The baseboard management controller BMC controls the virtual peripheral device to send a target heartbeat request to the platform path controller PCH according to a preset sending period, and receives response data fed back by the platform path controller PCH based on the target heartbeat request.

[0097] For example, as shown in FIG1 , the baseboard management controller BMC cannot communicate directly with the server operating system, but needs to rely on the platform path controller PCH to monitor the status of the server operating system.

[0098] It's important to note that while a server operating system reboot can indicate a system downtime to a certain extent, other triggers for a system reboot include not only fatal faults or unrecoverable recoverable faults, but also unstable factors such as physical triggering of the power button, normal in-band OS reboot command execution, and BMC receiving an out-of-band reboot command. Furthermore, virtual machine operating systems can remain down and not reboot after receiving a fault. Therefore, using a trusted channel to detect whether the operating system is operating normally is a more reliable solution.

[0099] Step 204: If the baseboard management controller (BMC) detects an interruption in the response data fed back by the operating system within a preset time period when the baseboard management controller (BMC) sends the target heartbeat request to the server operating system, the baseboard management controller (BMC) determines that the server operating system is down, and determines the faulty component in the server based on the target fault information.

[0100] For example, the downtime detection task will control the virtual USB keyboard or mouse device of the baseboard management controller (BMC) chip to send bmRequest requests to the server operating system according to a preset sending period (for example, an interval of 50ms) to detect whether the system can still respond normally. If the server system experiences a transition from responding to bmRequest to not responding within a preset time (for example, 1 minute) after the recoverable fault occurs, or the baseboard management controller (BMC) detects that the server operating system has been restarted, it can be accurately indicated that the recoverable fault caused the system to shut down, and can be directly recorded in the system environment log (System Event Log, SEL) to prompt the user to replace the diagnosed faulty component.

[0101] In one possible implementation, in step 204, whether the response data fed back by the operating system is interrupted can be determined based on the cumulative number of failures in receiving the response data fed back by the server operating system. Specifically, when the cumulative number of failures exceeds a preset threshold, it is determined whether the response data fed back by the operating system is interrupted.

[0102] Specifically, the above step 204 may further include the following step 204a:

[0103] Step 204a: When the baseboard management controller (BMC) determines that the server operating system is down, it analyzes the target fault information, determines the faulty component, and records the downtime event and the faulty component in the system environment log (SEL) to inform the user of the diagnosed faulty component.

[0104] For example, as shown in Figure 3, which is a flowchart of the downtime status detection task program provided in an embodiment of the present application, when executing a loop to detect the fault status, if a recoverable fault is detected, the timing starts, and a heartbeat is requested from the server operating system HOST at intervals of 50ms, and the heartbeat data fed back by the server operating system is received; if the timing reaches 1 minute and there is no interruption in receiving the heartbeat data fed back by the server operating system, it can be determined that the recoverable fault has been successfully repaired; if there is an interruption in receiving the heartbeat data fed back by the server operating system within 1 minute of the timing, and the cumulative number of heartbeat failures is greater than 10, it can be determined that the recoverable fault has failed to be repaired.

[0105] The present invention provides a recoverable fault firmware detection system. First, the basic input and output system BIOS obtains the target fault information stored in the fault register and sends the target fault information to the baseboard management controller (BMC). When the baseboard management controller (BMC) receives the target fault information sent by the basic input and output system BIOS, it parses the target fault information and determines the fault type of the target fault information. Then, when the fault type of the target fault information is a recoverable fault, the baseboard management controller (BMC) controls the virtual external device to send a target heartbeat request to the server operating system according to a preset sending period, and receives response data fed back by the server operating system based on the target heartbeat request. Finally, within the preset duration of the baseboard management controller sending the target heartbeat request to the server operating system, if the baseboard management controller (BMC) detects an interruption in the response data fed back by the operating system, it determines that the server operating system is down, and determines the faulty component in the server based on the target fault information. The target fault information is generated when the memory and PCIe devices in the server fail. The virtual external device is a universal serial bus (USB) device virtualized by the baseboard management controller (BMC). In this way, it is possible to accurately detect the situation where the operating system is down due to a recoverable fault, and promptly remind the user to handle the situation when the operating system is down.

[0106] It should be noted that the recoverable fault firmware detection method provided in the embodiments of the present application can be executed by a recoverable fault firmware detection system, or by various components of the recoverable fault firmware detection device that are used to execute the recoverable fault firmware detection system. In the embodiments of the present application, the recoverable fault firmware detection method provided in the embodiments of the present application is described using the recoverable fault firmware detection system executing the recoverable fault firmware detection method as an example.

[0107] It should be noted that the recoverable fault firmware detection methods shown in the figures in the above-mentioned methods in the embodiments of the present application are each described by way of example in conjunction with one figure in the embodiments of the present application. In specific implementations, the recoverable fault firmware detection methods shown in the figures in the above-mentioned methods may also be implemented in conjunction with any other combinable figures shown in the above-mentioned embodiments, and will not be further described here.

[0108] The following describes a firmware detection system for recoverable faults provided by the present application. The firmware detection method for recoverable faults described below and described above can refer to each other.

[0109] Based on the structural diagram of the recoverable fault firmware detection system shown in Figure 1, the recoverable fault firmware detection system provided by the embodiment of the present application includes: a basic input and output system BIOS, a server operating system, a fault register set in the central processing unit CPU, a baseboard management controller BMC, a memory and a high-speed serial computer expansion bus standard PCIe device; the basic input and output system BIOS is used to obtain the target fault information stored in the fault register and send the target fault information to the baseboard management controller BMC; the baseboard management controller BMC is used to parse the target fault information when receiving the target fault information sent by the basic input and output system BIOS, and determine the fault type of the target fault information; the baseboard management controller The BMC is also used to control the virtual external device to send a target heartbeat request to the server operating system according to a preset sending cycle and receive response data fed back by the server operating system based on the target heartbeat request when the fault type of the target fault information is a recoverable fault; the baseboard management controller BMC is also used to determine that the server operating system is down if an interruption is detected in the response data fed back by the operating system within the preset time period when the baseboard management controller sends the target heartbeat request to the server operating system, and determine the faulty component in the server based on the target fault information; wherein the target fault information is generated when the memory and PCIe devices in the server fail; the virtual external device is a universal serial bus USB device virtualized by the baseboard management controller BMC.

[0110] Optionally, the fault register includes: an UNCERRSTS register, a DEVSTS register, and a STATUS register; a basic input and output system BIOS, which is specifically used to detect the fault information generated in the DEVSTS register according to a preset detection cycle, and / or the fault information generated in the STATUS register, and / or the fault information generated in the UNCERRSTS register.

[0111] Optionally, the target fault information is fault information generated when a memory fault occurs; the baseboard management controller BMC is specifically used to parse the target fault information, and when the fields contained in the target fault information meet the first preset rule, determine that the fault type of the target fault information is a recoverable fault; otherwise, determine that the fault type of the target fault information is an unrecoverable fault.

[0112] Optionally, the baseboard management controller BMC is specifically configured to parse the target fault information, and determine that the fault type of the target fault information is a recoverable fault when the target fault information indicates that a fault is recorded in the STATUS register and the target fault information includes a preset field.

[0113] Optionally, the target fault information is fault information generated when a PCIe device fails; the baseboard management controller BMC is specifically used to parse the target fault information, and when the fields contained in the target fault information meet the second preset rule, determine that the fault type of the target fault information is a recoverable fault.

[0114] Optionally, the baseboard management controller BMC is specifically configured to determine that the fault type of the target fault information is a recoverable fault when the target fault information indicates that an unrecoverable fault is recorded in the UNCERRSTS register and a non-fatal fault is recorded in the DEVSTS register.

[0115] Optionally, the system also includes: a platform path controller PCH; a baseboard management controller BMC, specifically used to control the virtual external device to send a target heartbeat request to the platform path controller PCH according to a preset sending cycle, and receive response data feedback from the platform path controller PCH based on the target heartbeat request.

[0116] Optionally, the target heartbeat request is a HID_GET_REPORT request constructed using the bmRequest field; the baseboard management controller BMC is specifically used to perform an initialization operation and, after the initialization operation is completed, create a downtime status detection task process; the initialization operation includes: initializing related library functions and configuring the system clock; the baseboard management controller BMC is also specifically used to execute a downtime status detection task process, call the USBD_LL_SetupStage function to process the SETUP stage of the virtual external device, call the USBD_LL_DataInStage function to process the IN stage of the virtual external device, and call the USBD_LL_Data The OutStage function processes the OUT stage of the virtual external device; the baseboard management controller BMC is specifically used to construct the relevant parameters of the HID_GET_REPORT request when it is determined that the virtual external device is configured successfully; the baseboard management controller BMC is specifically used to call the usb_control_msg function to send the HID_GET_REPORT request after completing the construction of the HID_GET_REPORT request; the baseboard management controller BMC is specifically used to call the USBD_HID_GetReport function in the peripheral interrupt processing function to receive the response data of the server operating system based on the target heartbeat request feedback.

[0117] Optionally, the baseboard management controller BMC is specifically used to analyze the target fault information, determine the faulty component, and record the downtime event and the faulty component in the system environment log SEL to prompt the user of the diagnosed faulty component when determining that the server operating system is down.

[0118] The present application provides a firmware detection method for recoverable faults. First, the basic input / output system (BIOS) obtains target fault information stored in a fault register and sends the target fault information to a baseboard management controller (BMC). Upon receiving the target fault information sent by the basic input / output system (BIOS), the baseboard management controller (BMC) parses the target fault information and determines the fault type of the target fault information. Then, if the fault type of the target fault information is a recoverable fault, the baseboard management controller (BMC) controls a virtual peripheral device to send a target heartbeat request to a server operating system according to a preset sending period, and receives response data fed back by the server operating system based on the target heartbeat request. Finally, if the baseboard management controller (BMC) detects an interruption in the response data fed back by the operating system within a preset time period of the baseboard management controller sending the target heartbeat request to the server operating system, the baseboard management controller (BMC) determines that the server operating system is down, and determines the faulty component in the server based on the target fault information. The target fault information is generated when a memory or PCIe device in the server fails, and the virtual peripheral device is a universal serial bus (USB) device virtualized by the baseboard management controller (BMC). In this way, it is possible to accurately detect the situation where the operating system is down due to a recoverable fault, and promptly remind the user to handle the situation when the operating system is down.

[0119] On the other hand, the present application also provides a computer-readable instruction product, which includes computer-readable instructions stored on a non-volatile computer-readable storage medium, and the computer-readable instructions include program instructions. When the program instructions are executed by a computer, the computer can execute the recoverable fault firmware detection method provided by the above methods, the method comprising: the basic input and output system BIOS obtains the target fault information stored in the fault register, and sends the target fault information to the baseboard management controller BMC; when the baseboard management controller BMC receives the target fault information sent by the basic input and output system BIOS, it parses the target fault information and determines the fault type of the target fault information; the baseboard management controller BMC When the fault type of the target fault information is a recoverable fault, the MC controls the virtual external device to send a target heartbeat request to the server operating system according to a preset sending period, and receives response data fed back by the server operating system based on the target heartbeat request; within the preset time period when the baseboard management controller BMC sends the target heartbeat request to the server operating system, if it detects that the response data fed back by the operating system is interrupted, it determines that the server operating system is down, and determines the faulty component in the server based on the target fault information; wherein, the target fault information is generated when the memory and PCIe devices in the server fail; the virtual external device is a universal serial bus USB device virtualized by the baseboard management controller BMC.

[0120] On the other hand, the present application also provides a non-volatile computer-readable storage medium having computer-readable instructions stored thereon, which, when executed by a processor, are implemented to execute the above-mentioned recoverable fault firmware detection methods, the methods comprising: a basic input / output system BIOS obtaining target fault information stored in a fault register and sending the target fault information to a baseboard management controller (BMC); upon receiving the target fault information sent by the basic input / output system BIOS, the baseboard management controller (BMC) parses the target fault information and determines the fault type of the target fault information; when the fault type of the target fault information is a recoverable fault, the baseboard management controller (BMC) controls a virtual external device to send a target heartbeat request to a server operating system according to a preset sending period, and receives response data fed back by the server operating system based on the target heartbeat request; if the baseboard management controller (BMC) detects an interruption in the response data fed back by the operating system within a preset time period of sending the target heartbeat request to the server operating system, the baseboard management controller (BMC) determines that the server operating system is down, and determines the faulty component in the server based on the target fault information; wherein the target fault information is generated when a memory and a PCIe device in the server fail; and the virtual external device is a universal serial bus (USB) device virtualized by the baseboard management controller (BMC).

[0121] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units. That is, they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0122] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the relevant technology, can be embodied in the form of a software product. The computer software product can be stored in a non-volatile computer-readable storage medium, such as a ROM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment or certain parts of the embodiment.

[0123] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A firmware detection system capable of recovering from faults, characterized in that: The system includes: a basic input and output system BIOS, a server operating system, a fault register set in a central processing unit CPU, a baseboard management controller BMC, a memory and a high-speed serial computer expansion bus standard PCIe device; The basic input and output system BIOS is used to obtain the target fault information stored in the fault register and send the target fault information to the baseboard management controller BMC; The baseboard management controller BMC is used to parse the target fault information and determine the fault type of the target fault information when receiving the target fault information sent by the basic input and output system BIOS; The baseboard management controller BMC is further used to control the virtual external device to send a target heartbeat request to the server operating system according to a preset sending cycle when the fault type of the target fault information is a recoverable fault, and receive response data fed back by the server operating system based on the target heartbeat request; The baseboard management controller BMC is further used to determine that the server operating system is down if it is detected that the response data fed back by the operating system is interrupted within a preset time period when the baseboard management controller sends the target heartbeat request to the server operating system, and determine the faulty component in the server according to the target fault information; The target fault information is generated when a memory and a PCIe device in the server fail; and the virtual external device is a universal serial bus (USB) device virtualized by the baseboard management controller (BMC).

2. The system according to claim 1, characterized in that The fault registers include: a UNCERRSTS register, a DEVSTS register and a STATUS register; The basic input and output system BIOS is specifically used to detect the fault information generated in the DEVSTS register and / or the fault information generated in the STATUS register and / or the fault information generated in the UNCERRSTS register according to a preset detection period.

3. The system according to claim 2, characterized in that The target fault information is the fault information generated when a memory fault occurs; The baseboard management controller BMC is specifically used to parse the target fault information, and determine that the fault type of the target fault information is a recoverable fault if the field included in the target fault information meets the first preset rule, otherwise determine that the fault type of the target fault information is an unrecoverable fault; The first preset rule is used to determine whether the target fault information contains a preset field.

4. The system according to claim 3, characterized in that The baseboard management controller BMC is specifically used to parse the target fault information, and determine that the fault type of the target fault information is a recoverable fault when the target fault information indicates that a fault is recorded in the STATUS register and the target fault information contains a preset field.

5. The system according to claim 2, characterized in that The target fault information is the fault information generated when a PCIe device fails; The baseboard management controller BMC is specifically used to parse the target fault information, and when the field included in the target fault information satisfies the second preset rule, determine that the fault type of the target fault information is a recoverable fault. barrier.

6. The system according to claim 5, characterized in that The baseboard management controller BMC is specifically used to determine that the fault type of the target fault information is a recoverable fault when the target fault information indicates that an unrecoverable fault is recorded in the UNCERRSTS register and a non-fatal fault is recorded in the DEVSTS register.

7. The system according to claim 1, characterized in that The system further comprises: a platform path controller PCH; The baseboard management controller BMC is specifically used to control the virtual external device to send the target heartbeat request to the platform path controller PCH according to the preset sending period, and receive response data fed back by the platform path controller PCH based on the target heartbeat request.

8. The system according to claim 1 or 7, characterized in that: The target heartbeat request is a HID_GET_REPORT request constructed using the bmRequest field; The baseboard management controller BMC is specifically used to perform an initialization operation and, after the initialization operation is completed, create a downtime status detection task process; the initialization operation includes: initializing related library functions and configuring a system clock; The baseboard management controller BMC is further used to execute the downtime detection task process, call USBD_LL_SetupStage function to process the SETUP stage of the virtual external device, call USBD_LL_DataInStage function to process the IN stage of the virtual external device, and call USBD_LL_DataOutStage function to process the OUT stage of the virtual external device; The baseboard management controller BMC is further configured to construct relevant parameters of a HID_GET_REPORT request when it is determined that the virtual external device is configured successfully; The baseboard management controller BMC is further configured to call the usb_control_msg function to send the HID_GET_REPORT request after completing the construction of the HID_GET_REPORT request; The baseboard management controller BMC is further specifically used to call the USBD_HID_GetReport function in the peripheral interrupt processing function to receive the response data fed back by the server operating system based on the target heartbeat request.

9. The system according to claim 1, characterized in that The baseboard management controller BMC is specifically used to analyze the target fault information and determine the faulty component when determining that the server operating system is down, and record the downtime event and the faulty component in the system environment log SEL to prompt the user of the diagnosed faulty component.

10. A method for detecting a recoverable fault in firmware, characterized in that: A firmware detection system for recoverable faults, the system comprising: a basic input and output system BIOS, a server operating system, a fault register set in a central processing unit CPU, a baseboard management controller BMC, a memory, and a high-speed serial computer expansion bus standard PCIe device; The method comprises: The basic input and output system BIOS obtains the target fault information stored in the fault register, and sends the target fault information to the baseboard management controller BMC; The baseboard management controller BMC, when receiving the target fault information sent by the basic input and output system BIOS, parses the target fault information and determines the fault type of the target fault information; The baseboard management controller BMC controls the virtual external device to send a target heartbeat request to the server operating system according to a preset sending cycle, and receives response data fed back by the server operating system based on the target heartbeat request, when the fault type of the target fault information is a recoverable fault; and The baseboard management controller BMC determines that the server operating system is down if it detects that the response data fed back by the operating system is interrupted within a preset time period when the baseboard management controller sends the target heartbeat request to the server operating system, and determines the faulty component in the server according to the target fault information; The target fault information is generated when a memory and a PCIe device in the server fail; and the virtual external device is a universal serial bus (USB) device virtualized by the baseboard management controller (BMC).

11. The method according to claim 10, characterized in that The fault registers include: a UNCERRSTS register, a DEVSTS register and a STATUS register; The basic input and output system BIOS obtains the target fault information stored in the fault register, including: The basic input and output system BIOS detects the fault information generated in the DEVSTS register and / or the fault information generated in the STATUS register and / or the fault information generated in the UNCERRSTS register according to a preset detection period.

12. The method according to claim 11, characterized in that The target fault information is the fault information generated when a memory fault occurs; When receiving the target fault information sent by the basic input and output system BIOS, the baseboard management controller BMC parses the target fault information and determines the fault type of the target fault information, including: The baseboard management controller BMC parses the target fault information, and determines that the fault type of the target fault information is a recoverable fault if the field included in the target fault information satisfies a first preset rule; otherwise, determines that the fault type of the target fault information is an unrecoverable fault.

13. The method according to claim 12, characterized in that The baseboard management controller BMC parses the target fault information, and determines that the fault type of the target fault information is a recoverable fault when a field included in the target fault information satisfies a first preset rule, including: The baseboard management controller BMC parses the target fault information, and determines that the fault type of the target fault information is a recoverable fault if the target fault information indicates that a fault is recorded in the STATUS register and the target fault information includes a preset field.

14. The method according to claim 11, characterized in that The target fault information is the fault information generated when a PCIe device fails; When receiving the target fault information sent by the basic input and output system BIOS, the baseboard management controller BMC parses the target fault information and determines the fault type of the target fault information, including: The baseboard management controller BMC parses the target fault information, and determines that the fault type of the target fault information is a recoverable fault if a field included in the target fault information satisfies a second preset rule.

15. The method according to claim 14, characterized in that The determining that the fault type of the target fault information is a recoverable fault when the field included in the target fault information satisfies a second preset rule includes: The baseboard management controller BMC records the unrecoverable fault in the UNCERRSTS register, and the When a non-fatal fault is recorded in the DEVSTS register, it is determined that the fault type of the target fault information is a recoverable fault.

16. The method according to claim 10, characterized in that The system further comprises: a platform path controller PCH; The baseboard management controller BMC controls the virtual external device to send a target heartbeat request to the server operating system according to a preset sending cycle, and receives response data fed back by the server operating system based on the target heartbeat request, including: The baseboard management controller BMC controls the virtual external device to send the target heartbeat request to the platform path controller PCH according to the preset sending period, and receives response data fed back by the platform path controller PCH based on the target heartbeat request.

17. The method according to claim 10 or 16, characterized in that The target heartbeat request is a HID_GET_REPORT request constructed using the bmRequest field; The baseboard management controller BMC controls the virtual external device to send a target heartbeat request to the server operating system according to a preset sending cycle, including: The baseboard management controller BMC performs an initialization operation and creates a downtime status detection task process after the initialization operation is completed; the initialization operation includes: initializing related library functions and configuring the system clock; The baseboard management controller BMC executes the downtime status detection task process, calls USBD_LL_SetupStage function to process the SETUP stage of the virtual external device, calls USBD_LL_DataInStage function to process the IN stage of the virtual external device, and calls USBD_LL_DataOutStage function to process the OUT stage of the virtual external device; When determining that the configuration of the virtual external device is successful, the baseboard management controller BMC constructs relevant parameters of a HID_GET_REPORT request; After completing the construction of the HID_GET_REPORT request, the baseboard management controller BMC calls the usb_control_msg function to send the HID_GET_REPORT request; and The baseboard management controller BMC calls the USBD_HID_GetReport function in the peripheral interrupt processing function to receive the response data fed back by the server operating system based on the target heartbeat request.

18. The method according to claim 10, characterized in that The baseboard management controller BMC determines the faulty component in the server according to the target fault information, including: When determining that the server operating system is down, the baseboard management controller BMC analyzes the target fault information, determines the faulty component, and records the downtime event and the faulty component in the system environment log SEL to prompt the user of the diagnosed faulty component.

19. A server, characterized in that: A recoverable fault firmware detection system as claimed in any one of claims 1 to 9 is arranged thereon.

20. A non-volatile computer-readable storage medium, characterized in that: Computer-readable instructions are stored thereon, and when the computer-readable instructions are executed by a processor, the steps of the firmware detection method for recoverable faults as claimed in any one of claims 10 to 18 are implemented.

Citation Information

Patent Citations

  • Fault processing method and device and server

    CN111414268A

  • Fault diagnosis method and device, electronic equipment and storage medium

    CN111767184A

  • Memory fault information recording method and equipment

    CN113742123A

  • Process monitoring based on memory scanning

    CN114625600A

  • Heterogeneous computing system and server system

    CN116932274A

Cited By

  • Baseboard management controller platform identification method, product, equipment and medium

    CN120561757A

  • Equipment fault processing method, management firmware, electronic equipment and storage medium

    CN121233386A