Server startup firmware exception processing method, device, equipment and storage medium

By acquiring and switching to the backup flash memory through the baseboard management controller, and loading the normal firmware after a power outage and restart, the problem of inaccurate firmware handling during server startup is solved, and the server startup stability is improved.

CN120950309BActive Publication Date: 2026-01-27INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511463136.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-14
Publication Date
2026-01-27
Estimated Expiration
2045-10-14

AI Technical Summary

Technical Problem

Existing technologies have low accuracy in handling server boot firmware anomalies, resulting in low server startup stability.

Method used

The system obtains firmware anomaly information from the flash memory through the baseboard management controller, switches to the backup flash memory, and loads the normal firmware, including platform firmware, microcode firmware, and input/output system firmware, by power-off restart.

Benefits of technology

It improves the efficiency and accuracy of handling server startup anomalies, ensures normal server startup, and enhances startup stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120950309B_ABST
    Figure CN120950309B_ABST
Patent Text Reader

Abstract

The application discloses an abnormality processing method, device and equipment of server startup firmware and a storage medium, and relates to the technical field of servers. The method comprises the following steps: in response to a startup operation, obtaining firmware abnormality information of a first flash memory currently used by a baseboard management controller; if the firmware abnormality information comprises server platform firmware abnormality and / or microcode firmware abnormality, sending a channel switching instruction to a programmable control device by the baseboard management controller; the channel switching instruction is used for switching the first flash memory to a second flash memory for backup; sending a power-off instruction to a power supply unit by the baseboard management controller and triggering a server startup instruction; after power supply of the power supply unit is restored, loading server platform firmware, microcode firmware and input / output system firmware from the second flash memory to complete startup of the server. The method can improve stability of the server.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of server technology, and in particular to a method, apparatus, device and storage medium for handling abnormalities in server boot firmware. Background Technology

[0002] With the rapid development of server technology, users have increasingly higher requirements for server stability. Among these requirements, startup stability is a crucial indicator of server stability.

[0003] In related technologies, a series of boot firmware programs need to run when a server powers on. These boot firmware programs are stored in flash memory. When the boot firmware malfunctions, the server cannot boot normally. Therefore, error handling of the boot firmware is necessary during the server boot process to ensure normal startup. However, the accuracy of error handling in these technologies is relatively low, thus reducing the server's startup stability. Summary of the Invention

[0004] This application provides a method, apparatus, device, and storage medium for handling abnormalities in server boot firmware, in order to at least solve the problem that the accuracy of abnormal handling of boot firmware in related technologies is low, resulting in low server stability.

[0005] On the one hand, this application provides a method for handling exceptions in server boot firmware, including:

[0006] In response to the power-on operation, the firmware anomaly information of the currently used first flash memory is obtained through the server's baseboard management controller;

[0007] If the firmware error information includes server platform firmware error and / or microcode firmware error, a channel switching instruction is sent to the server's programmable controller through the baseboard management controller; the channel switching instruction is used to switch the first flash memory to the backup second flash memory.

[0008] The baseboard management controller sends a power-off command to the server's power supply unit and triggers the server's power-on command; the power-off command is used to power off and restart the power supply unit.

[0009] After the power supply unit restores power, the server platform firmware, microcode firmware, and input / output system firmware are loaded from the second flash memory to complete the server boot process.

[0010] On the other hand, this application provides an anomaly handling device for server boot firmware, including:

[0011] The acquisition unit is used to acquire firmware abnormality information of the currently used first flash memory through the baseboard management controller of the server in response to the power-on operation.

[0012] The switching unit is used to send a channel switching instruction to the programmable controller of the server through the baseboard management controller if the firmware abnormality information includes server platform firmware abnormality and / or microcode firmware abnormality; the channel switching instruction is used to switch the first flash memory to the backup second flash memory.

[0013] The instruction sending unit is used to send a power-off instruction to the power supply unit of the server through the baseboard management controller and trigger the power-on instruction of the server; the power-off instruction is used to power off and restart the power supply unit.

[0014] The processing unit is used to load the server platform firmware, microcode firmware, and input / output system firmware from the second flash memory after the power supply unit restores power, so as to complete the server boot-up.

[0015] This application also provides an electronic device, including: a memory for storing a computer program; and a processor for implementing the steps of any of the above-described server boot firmware exception handling methods when executing the computer program.

[0016] This application also provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, it implements the steps of any of the above-described server boot firmware exception handling methods.

[0017] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above-described server boot firmware exception handling methods.

[0018] This application provides a method, apparatus, device, and storage medium for handling firmware anomalies during server startup. The method includes: in response to a power-on operation, obtaining firmware anomaly information of a currently used first flash memory through the server's baseboard management controller; if the firmware anomaly information includes server platform firmware anomalies and / or microcode firmware anomalies, sending a channel switching instruction to the server's programmable controller through the baseboard management controller; the channel switching instruction is used to switch the first flash memory to a backup second flash memory; sending a power-off instruction to the server's power supply unit through the baseboard management controller and triggering a server startup instruction; the power-off instruction is used to power off and restart the power supply unit; after power is restored to the power supply unit, loading the server platform firmware, microcode firmware, and input / output system firmware from the second flash memory to complete the server startup. In this embodiment, by obtaining firmware anomaly information of the currently used first flash memory through the server's baseboard management controller and executing anomaly handling steps corresponding to the type of firmware anomaly information, the efficiency and accuracy of anomaly handling are improved, thus enhancing the server's startup stability. Furthermore, since server platform firmware and microcode firmware anomalies cannot be eliminated by powering on and restarting the server, the power supply unit can be powered off and restarted by power-off command in this embodiment of the application. This allows the server platform firmware, microcode firmware, and input / output system firmware to be reloaded from the second flash memory, thereby eliminating server platform firmware and microcode firmware anomalies in the first flash memory and enabling the server to start normally. Therefore, the startup stability of the server is further improved. Attached Figure Description

[0019] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 A schematic diagram illustrating an application scenario of the server boot firmware exception handling method provided in this application embodiment;

[0021] Figure 2 The flowchart of the server boot firmware exception handling method provided in the embodiments of this application Figure 1 ;

[0022] Figure 3 A schematic diagram of the abnormal handling method for server boot firmware provided in this application embodiment. Figure 1 ;

[0023] Figure 4 The flowchart of the server boot firmware exception handling method provided in the embodiments of this application Figure 2 ;

[0024] Figure 5 Schematic diagram of the abnormal handling device for server boot firmware provided in the embodiments of this application Figure 1 ;

[0025] Figure 6 A schematic diagram of the structure of the electronic device provided in this application. Detailed Implementation

[0026] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0027] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0028] With the rapid development of server technology, users have increasingly higher requirements for server stability. Among these requirements, startup stability is a crucial indicator of server stability.

[0029] In related technologies, a series of boot firmware programs need to run when a server powers on. These boot firmware programs are stored in flash memory, which is non-volatile. Different boot firmware programs are stored in different areas of the flash memory. When the boot firmware malfunctions, the server will fail to boot normally. Therefore, abnormal handling of the boot firmware is necessary during the server boot process to ensure normal startup. Currently, this is generally determined by checking the boot time to ascertain whether the boot firmware is functioning correctly.

[0030] Optionally, the boot firmware may include SPS (System Platform Services), microcode firmware, and BIOS (Basic Input / Output System firmware), all stored in a flash memory. Different firmware versions are stored in different areas of the FLASH memory.

[0031] The SPS firmware, executed by the ME (Management Engine), is primarily responsible for managing and monitoring the server's core hardware (CPU, memory, I / O controller, storage controller, etc.). It handles hardware wake-up, detection, and configuration after the server powers on, and provides standardized interfaces for upper-layer software (operating system, management tools) to access the hardware. The microcode firmware, executed by the PCU (Platform Controller Unit) within the CPU, decomposes complex instructions from upper-layer software (such as the operating system and applications) into micro-operations that the chip hardware can directly recognize. It also fixes hardware vulnerabilities present at the chip's manufacturing stage and optimizes instruction execution efficiency, forming the underlying foundation for the chip to correctly and efficiently execute upper-layer instructions. The BIOS firmware, executed by the CPU, is used for the server's system-wide power-on self-test initialization. It manages the hardware functions of I / O devices, performs protocol conversion between devices and the host (e.g., converting host read / write instructions into device-recognizable operations), and ensures stable device operation. This is a prerequisite for I / O devices to be recognized and accessed by the server.

[0032] However, because server boot time is affected by various factors, the accuracy of determining whether the firmware is abnormal by boot time is low, thus reducing the server's startup stability.

[0033] Therefore, accurately detecting firmware execution anomalies to improve server startup stability is a pressing technical problem that needs to be solved.

[0034] To address the aforementioned technical problems, this application proposes a method for handling server boot firmware anomalies. The specific steps include: First, in response to a boot operation, obtaining firmware anomaly information of the currently used first flash memory through the server's baseboard management controller; Second, if the firmware anomaly information includes server platform firmware anomalies and / or microcode firmware anomalies, sending a channel switching command to the server's programmable controller through the baseboard management controller; the channel switching command is used to switch the first flash memory to a backup second flash memory; sending a power-off command to the server's power supply unit through the baseboard management controller and triggering the server's boot command; the power-off command is used to power off and restart the power supply unit; Finally, after power is restored to the power supply unit, loading the server platform firmware, microcode firmware, and input / output system firmware from the second flash memory to complete the server boot process.

[0035] In this embodiment, by acquiring firmware anomaly information of the currently used first flash memory through the server's baseboard management controller and executing anomaly handling steps matching the anomaly type, the efficiency and accuracy of anomaly handling are improved, thus enhancing the server's startup stability. Furthermore, since server platform firmware and microcode firmware anomalies cannot be eliminated by simply restarting the server, this embodiment uses a power-off command to power off and restart the power supply unit, enabling the reloading of the server platform firmware, microcode firmware, and I / O system firmware from the second flash memory. This eliminates server platform firmware and microcode firmware anomalies in the first flash memory, allowing the server to start normally, further improving server startup stability.

[0036] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0037] The specific application environment architecture or specific hardware architecture on which the execution of the exception handling method of the server boot firmware depends is described here.

[0038] Figure 1 This is a schematic diagram illustrating an application scenario of the server boot firmware exception handling method provided in this application embodiment. For example... Figure 1 As shown, the server includes a BMC (Baseboard Management Controller) and a CPLD (Complex Programmable Logic Device). The BMC and CPLD work together to detect firmware anomalies. After identifying the firmware anomaly, the server performs anomaly handling through anomaly handling steps that match the firmware anomaly information, thereby improving the startup stability of the server.

[0039] The BMC (Browser Control Center) is a dedicated microcontroller on the server motherboard, primarily used for monitoring and managing the status of system hardware. It collects information through integrated sensors and interfaces and provides remote management capabilities, ensuring the server can be managed and monitored even when the operating system is not running. Key functions include: Health monitoring and alarms: The BMC can monitor key parameters such as server temperature, voltage, and fan speed in real time, issuing warnings or taking action when anomalies occur. Remote management: Supports the IPMI (Intelligent Platform Management Interface) protocol, allowing administrators to remotely access the server over the network for diagnostics, restarts, or shutdowns. Event logging: Records hardware-related events and alarms for subsequent analysis and troubleshooting. Firmware updates: Supports updating server component firmware over the network. Power management: Controls the server's power status, including power-on, power-off, and restart functions. Application scenarios: Data center operations personnel can remotely manage a large number of server nodes using the BMC, improving efficiency and reducing maintenance costs. In the event of a server hardware failure, the BMC can quickly locate the problem and help restore service rapidly.

[0040] A CPLD (Programmable Logic Device) is a programmable logic device used to implement complex digital logic circuits. Its primary function is to coordinate communication and control signal transmission between different hardware components. Key functions include signal processing and routing. A CPLD can process input signals according to preset logic rules and send the processed signals to the correct output ports. For example, in a multiprocessor system, a CPLD can arbitrate access requests from multiple CPUs to shared resources. Initialization and configuration: During system startup, a CPLD may perform initialization tasks such as setting the bus frequency and configuring peripheral devices. Fault detection and isolation: In some cases, a CPLD can detect hardware faults and attempt to isolate the faulty component to prevent it from affecting the stability of the entire system. Flexible design and adjustment: Because a CPLD is programmable, its internal logic can be reconfigured according to specific needs to adapt to different hardware architectures or upgrade requirements. Application scenarios: In high-performance computing environments, a CPLD can optimize data paths and improve overall system performance. For highly customized server platforms, a CPLD provides a convenient way to achieve specific functional requirements without changing the physical circuit design.

[0041] In this embodiment, by combining the baseboard management controller with the programmable controller program detection, fault information is mapped to a custom fault register flag. By establishing various firmware startup fault identifiers on the BMC side, the repair of various firmware faults and automatic power-on after a fault occurs are effectively achieved. The following specific embodiments describe the handling methods for different firmware anomalies.

[0042] Figure 2 The flowchart of the server boot firmware exception handling method provided in the embodiments of this application Figure 1 The exception handling method of the server boot firmware can be executed by the server itself, on which a basic input / output system can be deployed. For example... Figure 2 As shown, the method includes:

[0043] S201. In response to the power-on operation, obtain firmware abnormality information of the currently used first flash memory through the server's baseboard management controller.

[0044] In this embodiment of the disclosure, the boot firmware in the first flash memory may include: server platform firmware, microcode firmware, and input / output system firmware.

[0045] In some embodiments, firmware exception information of the currently used first flash memory is obtained from the fault register of the programmable controller via the server's baseboard management controller. Optionally, the firmware exception information is stored via an identifier bit in the fault register. Accordingly, this step may include:

[0046] The server's baseboard management controller obtains the setting information of the first and second flag bits of the power-on fault register from the programmable controller; the first flag bit is used to indicate whether the server platform firmware is abnormal, and the second flag bit is used to indicate whether the microcode firmware is abnormal.

[0047] If the first flag is set, the firmware error information is determined to be: server platform firmware error; and / or, if the second flag is set, the firmware error information is determined to be: microcode firmware error.

[0048] It should be noted that the type of programmable controller is not specifically limited in the disclosed embodiments. Optionally, the programmable controller may be a programmable logic device, an MCU (Microcontroller Unit), or an FPGA (Field-Programmable Gate Array).

[0049] For example, the first flag bit can be represented as bit0, and the second flag bit can be represented as bit1. If bit0 is set, the server platform firmware is determined to be abnormal; if bit0 is not set, the server platform firmware is determined to be normal. If bit1 is set, the microcode firmware is determined to be abnormal; if bit1 is not set, the microcode firmware is determined to be normal.

[0050] Optionally, when the firmware error information includes server platform firmware errors and / or microcode firmware errors, a first fault identification information can be recorded. For example, the first fault identification information is an "SPS_microcode fault" identifier.

[0051] In some embodiments, the execution order of the boot firmware is matched with the power-on timing of the power signal pins of each key component in the server system platform. The stage at which a boot failure occurs is determined by whether the power signals of each component are normal. Accordingly, before obtaining the setting information of the first and second flag bits of the boot failure register from the programmable controller via the server's baseboard management controller, the method further includes: in response to the server's boot command, monitoring the first power signal of the platform control hub (PCH) and the second power signal of the processor (CPU) via the server's programmable controller; if the first power signal is not detected within a first preset time period, a server platform firmware failure is determined, and the first flag bit in the boot failure register is set; if the second power signal is not detected within the first preset time period, a microcode firmware failure is determined, and the second flag bit in the boot failure register is set.

[0052] In this embodiment of the disclosure, the value of the first preset duration is not specifically limited. For example, the first preset duration may be 10 seconds, 15 seconds, or 20 seconds.

[0053] For example, such as Figure 1 and Figure 3 As shown, when the power button signal is triggered, PCH Power OK (PCH power signal) is triggered, followed by SYS Power OK (system power signal), and then CPU Power GOOD (CPU power signal). Subsequently, the CPU begins to load the BIOS and execute the BIOS firmware program.

[0054] In this example, the first power signal can be the PCH (Platform Controller Hub) power signal, and the second power signal can be the CPU power signal. A CPLD detection program can be designed to start timing after the power button is pressed. If either PCH Power OK (PCH power signal) or CPU PowerGOOD (CPU power signal) fails to trigger within a first preset time period (e.g., 10 seconds), a fault is recorded and updated to a custom power-on fault register (which can be represented as Power On error). Specifically, if the PCH power signal fails to trigger, it indicates an SPS firmware execution fault, and the first flag bit (bit 0) of the power-on fault register is set. If the CPU power signal fails to trigger, it indicates a microcode firmware execution fault, and the first flag bit (bit 1) of the power-on fault register is set.

[0055] In this embodiment of the disclosure, the hardware signals of the server are used to characterize the boot execution anomalies of various firmware on the server. This can effectively and reliably detect firmware execution anomalies, ensuring that firmware anomaly detection will not result in false detection or false triggering, thus improving the accuracy of firmware anomaly detection.

[0056] In other embodiments, such as Figure 3 As shown, the server's baseboard management controller obtains firmware error information corresponding to the input / output system firmware of the currently used first flash memory based on the power-on self-test (POST) completion signal of the input / output system firmware. If the POST completion signal is received within a second preset time period, the output system firmware is determined to be normal; otherwise, if the POST completion signal is not received within the second preset time period, the output system firmware is determined to be abnormal. The POST completion signal can be represented as the POST (Power On Self Test) Complete signal.

[0057] In this embodiment of the disclosure, the value of the second preset duration is not specifically limited. For example, the second preset duration may be 15 seconds, 20 seconds, or 25 seconds.

[0058] Optionally, when the firmware error information includes an output system firmware error, a second fault identification information can be recorded. For example, the second fault identification information is a "BIOS_boot failure" identifier.

[0059] S202. If the firmware error information includes server platform firmware error and / or microcode firmware error, a channel switching instruction is sent to the server's programmable controller through the baseboard management controller; the channel switching instruction is used to switch the first flash memory to the backup second flash memory.

[0060] Optionally, such as Figure 1 As shown, the first flash memory can be the primary flash memory, and the second flash memory can be a backup flash memory. After the programmable controller receives a channel switching command, it can switch the first flash memory to the backup second flash memory via a serial peripheral interface switch (SPI switch). For example, if the serial peripheral interface corresponding to the first flash memory is "SPI1" and the serial peripheral interface corresponding to the second flash memory is "SPI2", the SPISwitch can enable SPI2 and disable SPI1, thus switching the first flash memory to the backup second flash memory.

[0061] In some embodiments, before sending the channel switching command, it can be determined whether the boot firmware of the second flash memory is normal. Accordingly, before sending the channel switching command to the programmable controller of the server through the baseboard management controller, the method further includes: querying the set information of the fault status register of the second flash memory; if it is 0, then sending the channel switching command to the programmable controller of the server through the baseboard management controller; if it is 1, then not sending the channel switching command, and recording the fault information of the first flash memory and / or the fault information of the second flash memory in the system event log.

[0062] If the fault status register is set to 0, it indicates that the boot firmware of the second flash memory is normal; if the fault status register is set to 1, it indicates that the boot firmware of the second flash memory is abnormal.

[0063] Optionally, the method further includes: initializing the fault status register of the target flash memory to 0, wherein the target flash memory is a first flash memory and / or a second flash memory; if the firmware exception information of the target flash memory includes one or more of server platform firmware exception, microcode firmware exception, and input / output system firmware exception, then determining the serial peripheral interface used by the target flash memory; and adjusting the fault status register of the target flash memory to 1 through the serial peripheral interface.

[0064] For example, a fault status register for the first flash memory and a fault status register for the second flash memory can be designed, with a default value of 0, indicating that the flash memory is fault-free. When the baseboard management controller queries the "SPS_microcode fault" flag or the "BIOS_boot fault" flag, it queries the CPLD to see whether the currently used SPI channel is SPI1 or SPI2, and sets the fault status register of the flash memory corresponding to the currently used SPI channel to 1.

[0065] In this embodiment of the disclosure, before sending the channel switching command, it is first determined whether the boot firmware of the second flash memory is normal; if it is normal, the switching is performed. This avoids the situation where both the first and second flash memories fail, thus preventing the server from getting stuck in a flash memory switching loop and further improving the server's boot stability.

[0066] In some embodiments, a test result register for the flash memory can also be added. If the boot firmware of the flash memory is normal, the test result register is set to 1, and the channel switching instruction is executed normally. In response to the completion of switching the first flash memory to the backup second flash memory, it is determined whether the set information of the test result register is 1. If it is 1, it indicates that the boot firmware in the newly switched second flash memory is normal and usable. At this time, the board management controller can copy the data in the second flash memory to the first flash memory. In response to the completion of data copying, the test result register of the second flash memory is cleared, and the fault status register of the first flash memory is cleared, so that the first flash memory can be used as a backup flash memory for the second flash memory.

[0067] In this embodiment of the disclosure, after the firmware fault of the first flash memory is repaired, the data in the second flash memory is copied to the first flash memory, and the fault status register of the first flash memory is cleared, so that the first flash memory can be used as a backup flash memory for the second flash memory. This facilitates repair through the first flash memory when the second flash memory fails, thereby further improving the boot stability of the server.

[0068] S203. Send a power-off command to the power supply unit of the server through the baseboard management controller and trigger the power-on command of the server; the power-off command is used to power off and restart the power supply unit.

[0069] In this embodiment of the disclosure, when the server platform firmware and / or microcode firmware are abnormal, the power supply unit can be temporarily powered off and restarted by a temporary power-off command, thereby enabling the server to reload the server platform firmware and microcode firmware.

[0070] For example, such as Figure 1 As shown, a power supply unit can be represented as a PSU (full name: Power Supply Unit).

[0071] S204. After the power supply unit restores power, load the server platform firmware, microcode firmware, and input / output system firmware from the second flash memory to complete the server boot.

[0072] In some embodiments, this step may include: after the power supply unit restores power, restarting the platform control hub and the baseboard management controller; loading the server platform firmware from the second flash memory through the restarted platform control hub; determining whether first fault identification information is recorded through the restarted baseboard management controller; if so, triggering a server power-on command, wherein the first fault identification information is recorded when the firmware abnormality information includes server platform firmware abnormality and / or microcode firmware abnormality; and loading the microcode firmware and input / output system firmware from the second flash memory through the server's processor to complete the server power-on.

[0073] Optionally, after the baseboard management controller triggers the server's power-on command, the first fault identification information can also be cleared.

[0074] In this embodiment of the disclosure, after the power supply unit is temporarily powered off and restarted, the platform control hub and the baseboard management controller restart. At this time, the platform control hub can load the server platform firmware from the second flash memory, and the processor can load the microcode firmware and input / output system firmware from the second flash memory. The boot firmware in the second flash memory is normal, so the server can be booted normally.

[0075] This application proposes a method for handling firmware anomalies during server startup: First, in response to a power-on operation, firmware anomaly information of the currently used first flash memory is obtained through the server's baseboard management controller; second, if the firmware anomaly information includes server platform firmware anomalies and / or microcode firmware anomalies, a channel switching command is sent to the server's programmable controller through the baseboard management controller; the channel switching command is used to switch the first flash memory to a backup second flash memory; a power-off command is sent to the server's power supply unit through the baseboard management controller, triggering a server startup command; the power-off command is used to power off and restart the power supply unit; finally, after power is restored to the power supply unit, the server platform firmware, microcode firmware, and input / output system firmware are loaded from the second flash memory to complete the server startup. In this embodiment, because the firmware anomaly information of the currently used first flash memory is obtained through the server's baseboard management controller, and anomaly handling steps matching the type of firmware anomaly information are executed, the efficiency and accuracy of anomaly handling are improved, thus enhancing the server's startup stability. Furthermore, since server platform firmware and microcode firmware anomalies cannot be eliminated by powering on and restarting the server, the power supply unit can be powered off and restarted by power-off command in this embodiment of the application. This allows the server platform firmware, microcode firmware, and input / output system firmware to be reloaded from the second flash memory, thereby eliminating server platform firmware and microcode firmware anomalies in the first flash memory and enabling the server to start normally. Therefore, the startup stability of the server is further improved.

[0076] It should be noted that if the firmware error message includes an input / output system firmware error, a power-down command is not required. Accordingly, such as... Figure 4 As shown, the method also includes:

[0077] S205. If the firmware error information includes an input / output system firmware error, a channel switching instruction is sent to the programmable controller of the server through the baseboard management controller; the channel switching instruction is used to switch the first flash memory to the backup second flash memory.

[0078] The method for switching the first flash memory to the backup second flash memory in step S205 is the same as the method for switching the first flash memory to the backup second flash memory in step S202, and will not be described again here.

[0079] S206. In response to the completion of switching to the second flash memory, a server restart command is triggered to load the input / output system firmware from the second flash memory to complete the server boot.

[0080] Optionally, this step may include: in response to the completion of switching to the second flash memory, determining whether a second fault identification information is recorded by the baseboard management controller; if so, triggering a power-on command for the server, wherein the second fault identification information is recorded in the case of firmware abnormality information including input / output system firmware abnormality.

[0081] Optionally, after the baseboard management controller triggers the server's power-on command, the second fault identification information can also be cleared.

[0082] In this embodiment, when the firmware error information includes server platform firmware and microcode firmware errors, the power supply unit is powered off and restarted via a power-off command to eliminate the server platform firmware error and microcode firmware error. When the firmware error information includes input / output system firmware errors, the input / output system firmware error is eliminated by sending a channel switching command, thereby realizing the repair of various firmware faults and automatic power-on after a fault occurs.

[0083] Figure 5 This is a schematic diagram of the abnormal handling device for server boot firmware provided in an embodiment of this application. Figure 5 As shown, the device includes:

[0084] The acquisition unit 501 is used to acquire firmware abnormality information of the currently used first flash memory through the baseboard management controller of the server in response to the power-on operation.

[0085] The switching unit 502 is used to send a channel switching instruction to the programmable controller of the server through the baseboard management controller if the firmware abnormality information includes server platform firmware abnormality and / or microcode firmware abnormality; the channel switching instruction is used to switch the first flash memory to the backup second flash memory.

[0086] The instruction sending unit 503 is used to send a power-off instruction to the power supply unit of the server through the baseboard management controller and trigger the power-on instruction of the server; the power-off instruction is used to power off and restart the power supply unit.

[0087] The processing unit 504 is used to load the server platform firmware, microcode firmware, and input / output system firmware from the second flash memory after the power supply unit restores power, so as to complete the server boot-up.

[0088] In some embodiments, the acquisition unit 501 acquires firmware abnormality information of the currently used first flash memory through the server's baseboard management controller, including: acquiring the setting information of the first flag bit and the setting information of the second flag bit of the power-on fault register from the programmable controller through the server's baseboard management controller; the first flag bit is used to identify whether the server platform firmware is abnormal, and the second flag bit is used to identify whether the microcode firmware is abnormal; if the setting information of the first flag bit is set, the firmware abnormality information is determined to be: server platform firmware abnormal; and / or, if the setting information of the second flag bit is set, the firmware abnormality information is determined to be: microcode firmware abnormal.

[0089] In some embodiments, the device further includes: a monitoring unit; the monitoring unit is configured to, in response to a power-on command from the server, monitor a first power signal of the control platform hub and a second power signal of the processor via the server's programmable controller; if the first power signal is not detected within a first preset time period, determine that the server platform firmware is abnormal and set a first flag bit in the power-on fault register; if the second power signal is not detected within the first preset time period, determine that the microcode firmware is abnormal and set a second flag bit in the power-on fault register.

[0090] In some embodiments, after the power supply unit restores power, the processing unit 504 loads the server platform firmware, microcode firmware, and input / output system firmware from the second flash memory to complete the server power-on. This includes: restarting the platform control hub and the baseboard management controller after the power supply unit restores power; loading the server platform firmware from the second flash memory through the restarted platform control hub; determining whether first fault identification information is recorded through the restarted baseboard management controller; if so, triggering a server power-on command, wherein the first fault identification information is recorded when the firmware abnormality information includes server platform firmware abnormality and / or microcode firmware abnormality; and loading the microcode firmware and input / output system firmware from the second flash memory through the server's processor to complete the server power-on.

[0091] In some embodiments, the processing unit 504 is further configured to, if the firmware error information includes an input / output system firmware error, send a channel switching instruction to the programmable controller of the server via the baseboard management controller; the channel switching instruction is configured to switch the first flash memory to a spare second flash memory; in response to the completion of the switch to the second flash memory, trigger a server restart instruction to load the input / output system firmware from the second flash memory to complete the server power-on.

[0092] In some embodiments, the processing unit 504 triggers a server restart command in response to the completion of switching to the second flash memory, including: in response to the completion of switching to the second flash memory, determining through the baseboard management controller whether second fault identification information is recorded, and if so, triggering a server power-on command; wherein the second fault identification information is recorded when firmware abnormality information includes an input / output system firmware abnormality.

[0093] In some embodiments, the device further includes: a query unit; the query unit is configured to query the set information of the fault status register of the second flash memory; if the value is 0, a channel switching instruction is sent to the programmable controller of the server through the baseboard management controller; if the value is 1, no channel switching instruction is sent, and the fault information of the first flash memory and / or the fault information of the second flash memory is recorded in the system event log.

[0094] In some embodiments, the apparatus further includes: an initialization unit; the initialization unit is configured to initialize the set information of the fault status register of the target flash memory to 0, wherein the target flash memory is a first flash memory and / or a second flash memory; if the firmware abnormality information of the target flash memory includes one or more of server platform firmware abnormality, microcode firmware abnormality, and input / output system firmware abnormality, then determine the serial peripheral interface used by the target flash memory; and adjust the set information of the fault status register of the target flash memory to 1 through the serial peripheral interface.

[0095] This application provides an anomaly handling device for server boot firmware. By acquiring firmware anomaly information of the currently used first flash memory through the server's baseboard management controller, and executing anomaly handling steps corresponding to the type of anomaly information, the efficiency and accuracy of anomaly handling are improved, thus enhancing server boot stability. Furthermore, since server platform firmware and microcode firmware anomalies cannot be eliminated by simply restarting the server, this embodiment uses a power-off command to power off and restart the power supply unit, enabling the reloading of the server platform firmware, microcode firmware, and I / O system firmware from the second flash memory. This eliminates server platform firmware and microcode firmware anomalies in the first flash memory, achieving normal server boot and further improving server boot stability.

[0096] For a description of the features of the server boot firmware exception handling device provided in this application, please refer to the relevant description of the server boot firmware exception handling method, which will not be repeated here.

[0097] Figure 6 A schematic diagram of the structure of the electronic device provided in this application. Figure 6As shown, the electronic device 60 provided in this embodiment includes at least one processor 601 and a memory 602. Optionally, the electronic device 60 further includes a communication component 603. The processor 601, memory 602, and communication component 603 are connected via a bus.

[0098] In the specific implementation process, at least one processor 601 executes computer execution instructions stored in memory 602, causing at least one processor 601 to execute the above-described embodiment of the abnormal handling method for server boot firmware.

[0099] The specific implementation process of processor 601 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.

[0100] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the application can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.

[0101] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.

[0102] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.

[0103] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above embodiments of the exception handling method for server boot firmware.

[0104] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0105] The embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above embodiments of the exception handling method for server boot firmware.

[0106] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described embodiments of the exception handling method for server boot firmware.

[0107] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0108] The above provides a detailed description of the abnormal handling method, apparatus, device, and storage medium for server boot firmware provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and core ideas of this application. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A method for handling abnormalities in server boot firmware, characterized in that, include: In response to the power-on operation, the first power signal of the hub and the second power signal of the processor are controlled by the programmable controller monitoring platform of the server. Firmware anomaly information of the currently used first flash memory is obtained from the programmable controller via the baseboard management controller of the server; Based on the exception type of the firmware exception information, execute exception handling steps that match the exception type. If the firmware error information includes server platform firmware error and / or microcode firmware error, then the baseboard management controller sends a channel switching instruction to the programmable controller of the server. The abnormal information of the server platform firmware is determined based on the monitored first power signal; the abnormal information of the microcode firmware is determined based on the monitored second power signal. The channel switching command is used to switch the first flash memory to the backup second flash memory; The baseboard management controller sends a power-off command to the power supply unit of the server and triggers a power-on command for the server; the power-off command is used to power off and restart the power supply unit. After the power supply unit restores power, the server platform firmware, microcode firmware, and input / output system firmware are loaded from the second flash memory to complete the power-on of the server. If the firmware error information includes an input / output system firmware error, a channel switching instruction is sent to the programmable controller of the server via the baseboard management controller; the channel switching instruction is used to switch the first flash memory to a backup second flash memory. In response to the completion of switching to the second flash memory, a restart command is triggered on the server to load the input / output system firmware from the second flash memory to complete the power-on of the server.

2. The anomaly handling method according to claim 1, characterized in that, The step of obtaining firmware anomaly information of the currently used first flash memory from the programmable controller via the server's baseboard management controller includes: The server's baseboard management controller obtains the setting information of the first flag bit and the second flag bit of the power-on fault register from the programmable controller; the first flag bit is used to indicate whether the server platform firmware is abnormal, and the second flag bit is used to indicate whether the microcode firmware is abnormal. If the setting information of the first identifier bit is set, then the firmware abnormality information is determined to be: server platform firmware abnormality; and / or, if the setting information of the second identifier bit is set, then the firmware abnormality information is determined to be: microcode firmware abnormality.

3. The anomaly handling method according to claim 2, characterized in that, The method further includes: If the first power signal is not detected within the first preset time period, it is determined that the server platform firmware is abnormal, and the first flag bit in the power-on fault register is set. If the second power signal is not detected within the first preset time period, it is determined that the microcode firmware is abnormal, and the second flag bit in the power-on fault register is set.

4. The anomaly handling method according to claim 1, characterized in that, After the power supply unit restores power, the server platform firmware, microcode firmware, and input / output system firmware are loaded from the second flash memory to power on the server, including: After the power supply unit restores power, the platform control hub and baseboard management controller are restarted, and the server platform firmware is loaded from the second flash memory through the restarted platform control hub. The restarted baseboard management controller determines whether the first fault identification information is recorded. If so, the server power-on command is triggered. The first fault identification information is recorded when the firmware abnormality information includes server platform firmware abnormality and / or microcode firmware abnormality. The server's processor loads the microcode firmware and input / output system firmware from the second flash memory to power on the server.

5. The anomaly handling method according to claim 1, characterized in that, In response to the completion of the switch to the second flash memory, a restart command is triggered for the server, including: In response to the completion of switching to the second flash memory, the baseboard management controller determines whether a second fault identification information is recorded. If so, the power-on command of the server is triggered. The second fault identification information is recorded when the firmware anomaly information includes an input / output system firmware anomaly.

6. The anomaly handling method according to any one of claims 1-5, characterized in that, Before sending the channel switching command to the programmable controller of the server via the baseboard management controller, the method further includes: The fault status register of the second flash memory is checked. If it is 0, a channel switching command is sent to the programmable controller of the server through the baseboard management controller. If the value is 1, no channel switching command will be sent, and the fault information of the first flash memory and / or the fault information of the second flash memory will be recorded in the system event log.

7. The anomaly handling method according to claim 6, characterized in that, Also includes: The fault status register of the initialization target flash memory is set to 0, wherein the target flash memory is a first flash memory and / or a second flash memory; If the firmware anomaly information of the target flash memory includes one or more of server platform firmware anomalies, microcode firmware anomalies, and input / output system firmware anomalies, then the serial peripheral interface used by the target flash memory is determined. The fault status register of the target flash memory is set to 1 via the serial peripheral interface.

8. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the exception handling method of the server boot firmware as described in any one of claims 1 to 7 when executing the computer program.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, it implements the steps of the exception handling method of the server boot firmware as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method and device for online updating target area of server platform service firmware

    CN117289963A

  • Operating system entering method and device, equipment and storage medium

    CN117608674A