A method and computing device for determining that an OS failed to boot
By using a collaborative marking mechanism between the BIOS and the OS, and by utilizing periodic SMI monitoring timers and CMOS memory marker values, the system can accurately determine OS boot failures and automatically repair faults. This solves the problem of inaccurate boot failure detection in existing technologies and improves system reliability and automatic repair capabilities.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NEW H3C TECH CO LTD
- Filing Date
- 2026-02-05
- Publication Date
- 2026-06-12
AI Technical Summary
Existing technologies cannot accurately determine the cause of operating system startup failure, resulting in low efficiency and inability to provide effective fault information in unattended environments, often leading to a vicious cycle of startup-failure-reset.
By using a collaborative marking mechanism between the BIOS and the OS, and by periodically monitoring the SMI timer and CMOS memory marking values, the system can determine whether the OS boot process is successful or not, and identify the fault type by comparing differences, thus enabling automatic repair.
Accurately determine whether the OS boots successfully or not, automatically analyze the cause of the failure and perform targeted repairs, avoid misjudgments, break dead loops, and improve system reliability and automatic repair capabilities.
Smart Images

Figure CN122195592A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of communication technology, and in particular to a method and computing device for determining OS startup failure. Background Technology
[0002] BIOS: Basic Input / Output System, responsible for hardware initialization and booting the operating system during the initial system startup.
[0003] OS: Operating system. Such as Linux, Windows, etc.
[0004] SMI: System Management Interrupt. A high-priority CPU interrupt used to trigger System Management Mode (SMM) to perform low-level hardware management tasks in an environment isolated from the OS.
[0005] SMM: System Management Mode. A separate, high-priority x86 CPU operating mode used to handle SMIs and execute firmware-level hardware control functions.
[0006] CMOS: Complementary Metal-Oxide-Semiconductor. A type of read / write RAM chip used to store BIOS hardware configuration information; in this invention, it specifically refers to its storage space.
[0007] BMC: Baseboard Management Controller. A separate, dedicated service processor used to monitor the physical status of the server (such as temperature and voltage) and provide out-of-band management functions.
[0008] In modern computing devices, especially server devices, a stable operating system startup is the foundation of business continuity. However, the OS startup process is complex, involving multiple stages such as hardware initialization, driver loading, and file system mounting. Startup failures often occur due to the following reasons: (1) corruption of critical data in the file system; (2) abnormality or alteration of external hardware devices (such as PCIe cards and hard drives); (3) incorrect modification of BIOS / UEFI configuration parameters; (4) errors in the BIOS firmware itself due to abnormal flashing or corruption.
[0009] Currently, the monitoring mechanism for OS startup failures can only be achieved by observing the OS dmesg output information. Administrators still need to manually troubleshoot the problem by connecting to serial port logs and entering recovery mode, which is inefficient and almost unusable in unattended environments. Summary of the Invention
[0010] To overcome the problems existing in related technologies, this specification provides a method and computing device for determining OS startup failure.
[0011] According to a first aspect of the embodiments of this specification, a method for determining OS startup failure is provided, the method comprising: The BIOS enables monitoring commands, instructing the CPU to execute a Handler; Retrieve the value from the preset address using a Handler; When this value is the first threshold, the OS is confirmed to have started successfully; when this value is the second threshold, the OS is confirmed to have started unsuccessfully.
[0012] The method further includes: Configure a periodic monitoring timer for the BIOS; The BIOS enable monitoring instructions include: When the periodic monitoring timer set in the BIOS times out, the monitoring command is enabled.
[0013] The method further includes: During the OS environment preparation phase of the BIOS boot process, a ReadyToBoot event is created.
[0014] The method further includes: Write a flag value to a preset address in the CMOS memory.
[0015] The step of obtaining the value from the preset address through the Handler includes: The fault detection handler is executed to read the flag value from a preset address in the CMOS memory; When the flag value is the first threshold, the OS is confirmed to have started successfully. When the flag value is the second threshold, a failure counter is started and the number of failures is accumulated. If the accumulated value reaches the target value, the OS startup is determined to have failed.
[0016] The method further includes: The BIOS stores the current configuration data in the first Config file and the previous configuration data in the second Config file; Once it is determined that the OS has failed to boot, the configurations in the first Config file and the second Config file are compared, and the differences are obtained. The fault type is determined based on the difference value, and repairs are carried out according to the fault type.
[0017] As can be seen from the above embodiments, by adding a value to the preset address, a collaborative marking mechanism between the OS and BIOS is implemented, which can accurately determine whether the OS has successfully started, avoiding the possibility of misjudgment by the traditional watchdog, and enabling the identification and repair of fault types.
[0018] According to a second aspect of the embodiments of this specification, a computing device is provided, the computing device comprising: The indicator module is used to instruct the BIOS to enable monitoring instructions and instruct the CPU to execute the Handler; The acquisition module is used to retrieve values from a preset address via a Handler; The judgment module is used to confirm that the OS has started successfully when the value is the first threshold, or to confirm that the OS has started unsuccessfully when the value is the second threshold.
[0019] The computing device further includes: a configuration module. The configuration module is used to set a periodic monitoring timer for the BIOS; The indicator module is used to enable the monitoring command when the periodic monitoring timer set by the BIOS times out.
[0020] The configuration module is also used to create a ReadyToBoot event during the OS environment preparation phase of the BIOS boot process.
[0021] The configuration module is also used to write a flag value to a preset address of the CMOS memory.
[0022] Specifically, the acquisition module is used to execute a fault detection handler to read a flag value from a preset address in the CMOS memory; when the flag value is a first threshold, the OS is confirmed to have started successfully. When the flag value is the second threshold, a failure counter is started and the number of failures is accumulated. If the accumulated value reaches the target value, the OS startup is determined to have failed.
[0023] The configuration module is also used to instruct the BIOS to store the current configuration data in the first Config file and store the previous configuration data in the second Config file. The judgment module is also used to compare the configurations in the first Config file and the second Config file and obtain the difference value when it is determined that the OS has failed to start. The fault type is determined based on the difference value, and repairs are carried out according to the fault type.
[0024] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this specification. Attached Figure Description
[0025] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this specification and, together with the description, serve to explain the principles of this specification.
[0026] Figure 1This is a flowchart illustrating a method for determining OS startup failure according to an exemplary embodiment.
[0027] Figure 2 This is a flowchart illustrating a method for determining OS startup failure according to an exemplary embodiment. Detailed Implementation
[0028] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this specification. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this specification as detailed in the appended claims.
[0029] The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of this specification. The singular forms “a,” “the,” and “the” as used in this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.
[0030] It should be understood that although the terms first, second, third, etc., may be used in this specification to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this specification, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0031] Currently, the primary system reset solution for detecting and resetting OS boot failures is a hardware watchdog timer. Specifically, a hardware watchdog timer is initialized during the BIOS phase. After the BIOS transfers control to the OS, the watchdog timer starts timing. Once the OS boots successfully, its kernel takes over the watchdog service and periodically "feeds" the watchdog. If the OS boot process fails at any stage (kernel panic, driver error, file system check failure, etc.), the "feeding" behavior stops. When the watchdog timer times out, a hardware-level system reset is triggered, causing the device to restart.
[0032] However, the above methods can only determine if the startup timed out, but cannot determine why it failed. They cannot distinguish between BIOS configuration errors and hardware changes that caused the failure, and therefore cannot provide any effective fault information for maintenance, testing, and engineering development personnel. Furthermore, the repair method described above is a global reset, which is only a temporary fix for faults caused by configuration errors or firmware corruption; after restarting, the device will again fall into a vicious cycle of startup-failure-reset.
[0033] To address the aforementioned technical problems, this disclosure provides a method for determining OS boot failure, such as... Figure 1 As shown, the method includes: The S101 BIOS enables monitoring instructions, which instruct the CPU to execute a Handler. S102 obtains the value from the preset address through Handler; S103 If the value is the first threshold, the OS is confirmed to have started successfully; if the value is the second threshold, the OS is confirmed to have started unsuccessfully.
[0034] In this embodiment, the OS can be identified as having started normally by implementing a collaborative marking mechanism between the OS and BIOS, which involves setting a periodic SMI monitoring timer on the BIOS side and collaboratively marking the CMOS with a success marking script on the OS side.
[0035] Specifically, the BIOS sets a monitoring timer. During the BIOS boot process, before handing over control to the OS loader, a ReadyToBoot event is created, and a periodic monitoring timer is enabled, for example, set to 5 seconds.
[0036] Meanwhile, during OS startup, typically after the graphical user interface is fully ready (i.e., when the system is fully usable), a specific success flag value (e.g., 0xFE) is written to a preset address in the CMOS memory (e.g., offset address 0xF0) via a high-privilege script.
[0037] In one implementation, such as in Linux, this is achieved by creating a systemd service; in another implementation, such as in Windows, it is achieved by creating a scheduled task.
[0038] With the above settings, after the periodic monitoring timer created in the BIOS times out, the CPU is enabled to enter SMM mode and execute the fault detection handler in the fault handling main program.
[0039] The processing program obtains the value of the preset address, such as the preset address (0xF0) in the CMOS mentioned above, and performs the following judgment.
[0040] It should be noted that in this example, the first threshold can be set to 0xFE, while the second threshold is a value other than 0xFE. In other embodiments, the first and second thresholds can also be other values.
[0041] In step S103, when the read value is 0xFE (i.e., the first threshold), the OS is considered to have started successfully. At this time, the counter can be cleared and the SMI timer can be disabled, and the monitoring cycle ends.
[0042] If the read value is not 0xFE, such as 0xFF (i.e., the second threshold), then the OS startup (temporary) is considered to have failed.
[0043] At this point, the failure counter Count can be started and incremented by 1.
[0044] When the counter Count value reaches the threshold (e.g., 60, corresponding to 5 minutes), it is determined that the OS startup has failed, triggering the complete fault handling main program.
[0045] In this example, a counter is set to prevent false positives for OS startup failure. If the OS can start before the value recorded by the counter reaches the threshold, the OS is considered to have started successfully; otherwise, the OS startup is considered to have failed.
[0046] As can be seen from the above embodiments, by setting a preset address value and setting a periodic monitoring timer in the BIOS, a collaborative marking mechanism between the OS and BIOS is achieved, which can accurately determine whether the OS has successfully started up, avoiding the possibility of misjudgment by traditional watchdog timers.
[0047] In this embodiment, the BIOS stores the current configuration data in the first Config file and stores the previous configuration data in the second Config file.
[0048] Specifically, during each boot process, the BIOS transfers and saves configuration data (NVRAM options, peripheral list, firmware checksum) to two files in the BMC memory: Current Config and Last Boot Config.
[0049] The Current Config file can be understood as the first Config file, used to store the configuration data for this configuration. Last Boot Config can be understood as a second Config file, used to store the configuration data from the previous boot.
[0050] When an OS boot failure is detected, the startup analysis capability is activated. Specifically, the fault handling program reads the first and second Config files through the BMC interface to obtain the current and previous configuration data and perform a differential comparison.
[0051] Specifically, in one example, NVRAM option comparison can be performed: Compare the UEFI variable set to filter out changed options. External device configuration comparison: Compare the device list (Vendor ID, Device ID, BDF number) obtained through PCIe enumeration to detect device changes. BIOS firmware comparison: Calculate the checksum (e.g., CRC32) of the main firmware area in the current SPI Flash and compare it with the previous value to determine if the firmware has changed or is corrupted.
[0052] The system can categorize the differences into "Fault Type 1 (NVRAM option configuration change)," "Fault Type 2 (external device change)," or "Fault Type 3 (BIOS firmware anomaly)," and report the detailed data to the log or management platform through the BMC out-of-band management interface so that the administrator can take further action based on the fault type.
[0053] In this embodiment, corresponding repair measures can be implemented based on the fault type obtained above.
[0054] For example, for fault type 1: Write the NVRAM data from the last normal boot stored in the BMC back to the NVRAM area of the SPIFlash and roll back the BIOS configuration.
[0055] For fault type 2: Access via PCIe configuration space, disable newly connected or failed-to-identify devices so that they are not enumerated on the next boot.
[0056] For fault type 3: Use the out-of-band flashing function of the BMC or the SpiFlash interface inside the BIOS to restore the main BIOS area using the last normal firmware backup.
[0057] After the repair operation is completed, the system will automatically perform a reset and restart the device, thus completing a complete "monitoring-diagnosis-repair" closed loop.
[0058] As can be seen from the above embodiments, after identifying OS faults through the collaborative marking mechanism between the OS and BIOS, the potential causes of boot failure can be automatically analyzed, categorized, and reported, greatly shortening the time required for manual troubleshooting and achieving rapid fault location. Storing suspected erroneous configurations and data in the BMC ensures they are not lost during system resets, providing valuable data support for fault analysis.
[0059] Furthermore, based on different diagnostic results, targeted repair operations (such as rolling back configurations and disabling devices) are performed to try to solve the problem at its root, breaking the vicious cycle of failure, reset, and repeated failures. This significantly improves the system's reliability and automatic repair capabilities, making it particularly suitable for unattended remote devices and data center servers.
[0060] Based on the above embodiments, this disclosure also provides an embodiment, such as... Figure 2 As shown, after the periodic SMI timer created by the BIOS is triggered, the CPU enters SMM mode and executes the fault detection Handler in the fault handling main program.
[0061] The Handler reads the value at a preset address (0xF0) in the CMOS and makes a judgment: If the read value is 0xFE, the OS is considered to have started successfully. The counter is cleared and the SMI timer is disabled, and the monitoring cycle ends.
[0062] If the read value is not 0xFE (the default value is usually 0x00), the startup failure counter Count value is incremented by 1.
[0063] When the counter Count value reaches the threshold (e.g., 60, corresponding to 5 minutes), it is determined that the OS startup has failed, triggering the complete fault handling main program.
[0064] Data preparation: During each boot process, the BIOS will transfer and save the current configuration data (NVRAM options, peripheral list, firmware checksum) to two files in the BMC memory: Current Config and Last Boot Config.
[0065] Configuration Comparison: The fault handling main program reads the current and previous configuration data through the BMC interface and performs a difference comparison, which mainly consists of the following parts: NVRAM option comparison: Compare the UEFI variable set and filter out the changing options.
[0066] External device configuration comparison: By comparing the device list (Vendor ID, DeviceID, BDF number) obtained through PCIe enumeration, device changes can be detected.
[0067] BIOS firmware comparison: Calculate the checksum (such as CRC32) of the main firmware area in the current SPI Flash, compare it with the previous value, and determine whether the firmware has changed or is damaged.
[0068] Fault reporting: The differences found will be classified as "Fault Type 1 (NVRAM option configuration change)", "Fault Type 2 (external device change)" or "Fault Type 3 (BIOS firmware abnormality)" and the detailed data will be reported to the log or management platform through the BMC out-of-band management interface.
[0069] Based on preset strategies and the diagnosed fault types, the main program can automatically perform repair operations: For fault type 1: Write the NVRAM data from the last normal boot stored in the BMC back to the NVRAM area of the SPI Flash, and roll back the BIOS configuration.
[0070] For fault type 2: Access via PCIe configuration space, disable newly connected or failed-to-identify devices so that they are not enumerated on the next boot.
[0071] For fault type 3: Use the out-of-band flashing function of the BMC or the SpiFlash interface inside the BIOS to restore the main BIOS area using the last normal firmware backup.
[0072] After the repair operation is completed, the system will automatically perform a reset and restart the device, thus completing a complete "monitoring-diagnosis-repair" closed loop.
[0073] Based on the above-described method embodiments, this disclosure also provides a computing device, the computing device comprising: The indicator module is used to instruct the BIOS to enable monitoring instructions and instruct the CPU to execute the Handler; The acquisition module is used to retrieve values from a preset address via a Handler; The judgment module is used to confirm that the OS has started successfully when the value is the first threshold, or to confirm that the OS has started unsuccessfully when the value is the second threshold.
[0074] The computing device further includes: a configuration module. The configuration module is used to set a periodic monitoring timer for the BIOS; The indicator module is used to enable the monitoring command when the periodic monitoring timer set by the BIOS times out.
[0075] The configuration module is also used to create a ReadyToBoot event during the OS environment preparation phase of the BIOS boot process.
[0076] The configuration module is also used to write a flag value to a preset address of the CMOS memory.
[0077] Specifically, the acquisition module is used to execute a fault detection handler to read a flag value from a preset address in the CMOS memory; when the flag value is a first threshold, the OS is confirmed to have started successfully. When the flag value is the second threshold, a failure counter is started and the number of failures is accumulated. If the accumulated value reaches the target value, the OS startup is determined to have failed.
[0078] The configuration module is also used to instruct the BIOS to store the current configuration data in the first Config file and store the previous configuration data in the second Config file. The judgment module is also used to compare the configurations in the first Config file and the second Config file and obtain the difference value when it is determined that the OS has failed to start. The fault type is determined based on the difference value, and repairs are carried out according to the fault type.
[0079] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of the solution in this specification according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0080] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0081] Other embodiments of this specification will readily occur to those skilled in the art upon consideration of the specification and practice of the invention claimed herein. This specification is intended to cover any variations, uses, or adaptations that follow the general principles of this specification and include common knowledge or customary techniques in the art not claimed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this specification are indicated by the following claims.
[0082] It should be understood that this specification is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this specification is limited only by the appended claims.
[0083] The above description is merely a preferred embodiment of this specification and is not intended to limit this specification. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of protection of this specification.
Claims
1. A method for determining OS boot failure, characterized in that, The method includes: The BIOS enables monitoring commands, instructing the CPU to execute a Handler; Retrieve the value from the preset address using a Handler; When this value is the first threshold, the OS is confirmed to have started successfully; when this value is the second threshold, the OS is confirmed to have started unsuccessfully.
2. The method according to claim 1, characterized in that, The method further includes: Configure a periodic monitoring timer for the BIOS; The BIOS enable monitoring instructions include: When the periodic monitoring timer set in the BIOS times out, the monitoring command is enabled.
3. The method according to claim 1, characterized in that, The method further includes: During the OS environment preparation phase of the BIOS boot process, a ReadyToBoot event is created.
4. The method according to claim 1, characterized in that, The method further includes: Write a flag value to a preset address in the CMOS memory.
5. The method according to claim 1, characterized in that, The step of obtaining the value from the preset address through the Handler includes: The fault detection handler is executed to read the flag value from a preset address in the CMOS memory; When the flag value is the first threshold, the OS is confirmed to have started successfully. When the flag value is the second threshold, a failure counter is started and the number of failures is accumulated. If the accumulated value reaches the target value, the OS startup is determined to have failed.
6. The method according to claim 1, characterized in that, The method further includes: The BIOS stores the current configuration data in the first Config file and the previous configuration data in the second Config file; Once it is determined that the OS has failed to boot, the configurations in the first Config file and the second Config file are compared, and the differences are obtained. The fault type is determined based on the difference value, and repairs are carried out according to the fault type.
7. A computing device, characterized in that, The computing device includes: The indicator module is used to instruct the BIOS to enable monitoring instructions and instruct the CPU to execute the Handler; The acquisition module is used to retrieve values from a preset address via a Handler; The judgment module is used to confirm that the OS has started successfully when the value is the first threshold, or to confirm that the OS has started unsuccessfully when the value is the second threshold.
8. The computing device according to claim 7, characterized in that, The computing device further includes: a configuration module, The configuration module is used to set a periodic monitoring timer for the BIOS; The indicator module is used to enable the monitoring command when the periodic monitoring timer set by the BIOS times out.
9. The computing device according to claim 8, characterized in that, The configuration module is also used to create a ReadyToBoot event during the OS environment preparation phase of the BIOS boot process.
10. The computing device according to claim 7, characterized in that, The configuration module is also used to write a flag value to a preset address of the CMOS memory.
11. The computing device according to claim 7, characterized in that, The acquisition module is specifically used to execute a fault detection handler to read a flag value from a preset address in the CMOS memory; when the flag value is a first threshold, the OS is confirmed to have started successfully. When the flag value is the second threshold, a failure counter is started and the number of failures is accumulated. If the accumulated value reaches the target value, the OS startup is determined to have failed.
12. The computing device according to claim 8, characterized in that, The configuration module is also used to instruct the BIOS to store the current configuration data in the first Config file and store the previous configuration data in the second Config file; The judgment module is also used to compare the configurations in the first Config file and the second Config file and obtain the difference value when it is determined that the OS has failed to start. The fault type is determined based on the difference value, and repairs are carried out according to the fault type.