Baseboard management controller starting method, computer equipment, medium and product

By reading the register value after BMC starts and detecting abnormal processes, storing the reason for reset, BMC's intelligent partition switching is realized, solving the problem of difficulty in determining partitions after BMC restarts, and improving the system availability and fault recovery capabilities.

CN120255970AActive Publication Date: 2025-07-04INSPUR SUZHOU INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202510705561.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-07-04
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

After the baseboard management controller (BMC) is restarted and reset, it is difficult to determine whether the partition needs to be switched based on the actual situation, which makes it difficult for the system to realize intelligent switching of the redundant backup system.

Method used

By reading the value of the first register, determining the target running partition and environment variables, and storing the reset cause and exception process when an exception is detected, restarting the BMC, using the hardware watchdog timer and the application layer daemon to coordinate monitoring, to achieve fault detection and automatic recovery.

Benefits of technology

It realizes intelligent partition switching in the case of BMC abnormalities, avoids long-term downtime, ensures system availability and service continuity, and reduces the impact of single point of failure on the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120255970A_ABST
    Figure CN120255970A_ABST
Patent Text Reader

Abstract

The invention discloses a substrate management controller starting method, computer equipment, a medium and a product, and relates to the technical field of server management.The substrate management controller starting method comprises the steps that after a substrate management controller is started, a value of a first register is read, and a target running partition and an environment variable are determined based on a specific field of the first register; detecting whether the substrate management controller is abnormal or not in the operation stage; and if the abnormality exists, storing a reset reason in a specific field of the first register, and restarting the substrate management controller. After the substrate management controller is started, whether the operation stage is abnormal or not is detected, after the operation stage is abnormal, the reset reason and the abnormal operation process are stored in the first register, restarting is carried out, the first register is a register which does not lose data after restarting, and after starting, the value in the first register is read, so that the reset reason and the abnormal operation process are stored in the first register. And determining whether to partition or not in combination with a reset reason, thereby realizing intelligent switching of the redundant backup system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of server management, and particularly to a method for starting a baseboard management controller, a computer device, a medium, and a product. Background Art

[0002] A baseboard management controller (BMC) is a baseboard management controller integrated on a server motherboard. As a dedicated microcontroller independent of the main system, it is used to implement functions such as hardware monitoring, remote power management, and out-of-band operation and maintenance. Usually, in order to improve the reliability and maintainability of the BMC system, two BMC partitions are usually designed, including a main partition and a backup partition, to ensure that the system can be quickly restored in case of firmware update failure or damage to the main partition. However, during the operation of the BMC, if kernel exceptions, application layer exceptions, etc. occur, it is difficult to distinguish the specific restart reason after the BMC restarts and resets, so it is difficult to determine whether partitioning is required. Summary of the Invention

[0003] This application provides a method for starting a baseboard management controller, a computer device, a medium, and a product, so as to at least solve the problem in the related art that it is difficult to switch partitions according to the actual situation after the baseboard management controller starts.

[0004] This application provides a method for starting a baseboard management controller, including: After the baseboard management controller starts, read the value of the first register, and determine the target operating partition and environment variables based on a specific field of the first register. The environment variables are used to represent the operating partition when the baseboard management controller starts. The specific field of the first register is at least used to represent the operating state of the baseboard management controller, and the operating state at least includes the reset reason and the running process where an exception occurs; Detect whether there is an exception during the operation stage of the baseboard management controller; If there is an exception, store the reset reason and the running process where an exception occurs in the specific field of the first register, and restart the baseboard management controller.

[0005] This application also provides a computer device, including: a memory for storing a computer program; a processor for implementing the steps of any of the above baseboard management controller starting methods when executing the computer program.

[0006] This application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above baseboard management controller starting methods are implemented.

[0007] The present application also provides a computer program product, including a computer program which, when executed by a processor, implements the steps of any of the above-mentioned substrate management controller startup methods.

[0008] The substrate management controller startup method provided by the embodiments of the present invention includes, after the substrate management controller is started, reading the value of a first register, and determining a target operating partition and environment variables based on specific fields of the first register; detecting whether there is an abnormality during the operation stage of the substrate management controller; if there is an abnormality, storing the reset reason in a specific field of the first register, and restarting the substrate management controller. The present invention detects whether there is an abnormality during the operation stage after the substrate management controller is started, stores the reset reason and the running process where the abnormality occurs in the first register after an abnormality occurs, and performs a restart. The first register is a register whose data is not lost during a restart. After startup, by reading the value in the first register and combining the reset reason, it is determined whether to perform partitioning, so as to realize the intelligent switching of a redundant backup system. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0010] Figure 1 It is a schematic flowchart of a substrate management controller startup method provided by an embodiment of the present application; Figure 2 It is a schematic flowchart of partition switching provided by an embodiment of the present application; Figure 3 It is a schematic diagram of the operation stage of a substrate management controller provided by an embodiment of the present application; Figure 4 It is a schematic structural diagram of a substrate management controller startup device provided by an embodiment of the present application; Figure 5 It is a schematic hardware structure diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0011] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present application.

[0012] It should be noted that in the description of this application, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0013] In order to enable those skilled in the art of this technology to better understand the solution of this application, the following further describes this application in detail with reference to the accompanying drawings and specific embodiments.

[0014] Combined with the specific application environment architecture or specific hardware architecture on which the execution of the baseboard management controller startup method depends, the specific application environment architecture or specific hardware architecture is described herein.

[0015] The baseboard management controller is an independent microcontroller embedded in server or high-end computer hardware, responsible for monitoring and managing the physical state of the device. It is integrated on the motherboard and provides remote management functions through an independent out-of-band management interface, including power control, temperature monitoring, fan speed regulation, logging, and fault diagnosis, etc. The BMC can still operate even when the host operating system crashes, ensuring that administrators can access and maintain the device through the network. It is a key component for achieving efficient operation and maintenance in data centers and cloud computing environments. In the BMC system, designing a Flash chip with two independent partitions (main partition and backup partition) is a typical redundant architecture, aiming to enhance system reliability and fault tolerance. The main partition runs the current BMC firmware, while the backup partition stores a verified stable version or a copy of the firmware from the last successful update. When the main partition fails due to abnormal firmware upgrade, data corruption, or hardware failure, the BMC's boot loader can automatically detect the error and switch to the backup partition to start, ensuring the continuous availability of the system. However, during the operation of the BMC, when a kernel crash or hang occurs, or when a key process in the BMC application layer is abnormal during operation, it will restart. After the restart, it is difficult to distinguish whether a partition switch is required because the specific reason for the restart reset is not clear. Based on this, the present invention provides a baseboard management controller startup method.

[0016] The embodiments of this application provide a baseboard management controller startup method, and the method is described in detail in combination with the execution process of the baseboard management controller startup method.

[0017] The embodiments of the present invention provide a baseboard management controller startup method, as Figure 1 shown in the schematic flowchart of the baseboard management controller startup method, the method includes the following steps: Step S101: After the baseboard management controller is started, read the value of the first register, and determine the target operating partition and the environment variable based on specific fields of the first register.

[0018] Among them, the environment variable is used to characterize the operating partition when the baseboard management controller is started. The specific fields of the first register are at least used to characterize the operating state of the baseboard management controller, and the operating state at least includes the reset reason and the running process where an exception occurs.

[0019] A register is a storage unit in a computer's hardware for temporarily storing data. In this embodiment, the first register is a register whose data is not lost after a restart, which is a specific register whose value is not lost after the baseboard management controller is restarted and reset. Different bits (BITs) in the first register represent different meanings, that is, the specific fields represent the operating state of the baseboard management controller.

[0020] Taking a 32-bit register as an example, among them, [0] (the 1st bit) is a Reserved reserved bit with a default value of 0; [2:1] (the 2nd - 3rd bits) use two BITs to mark the reset reason, with a default value of 00, where 00 means normal, 01 means kernel exception, and 10 means BMC exception (baseboard management controller exception); [7:3] (the 4th - 8th bits) use five BITs to represent the running process during the running stage, with a default value of 00000, where 00000 means no exception occurred in the running process. The running processes include kernel decompression and initialization, device driver loading, root file system mounting, user space initialization, and BMC process monitoring service startup in sequence. Different running processes correspond to different bits among the five BITs. Exemplarily, if it is 00001, it means there is an exception in kernel decompression and initialization; [8] (the 9th bit) uses one BIT to mark the image refresh status, 1 means the image is being refreshed, and 0 means the image refresh is completed or not refreshed; [31:9] (the 10th - 32nd bits) are Reserved reserved bits.

[0021] After the baseboard management controller is started, read the value stored in the first register. Different fields in the first register represent different meanings. By reading the values of the specific fields in the first register, it is determined whether the current startup of the baseboard management controller is a restart. If it is a restart, what is the reset reason before this restart, and what is the running process where an exception occurs, so as to determine whether to switch partitions for the situation where an exception occurred last time, including switching to another partition and maintaining the current partition. If switching to another partition, then the other partition is the target operating partition; if maintaining the current partition, then the current operating partition is used as the target operating partition.

[0022] The value of the environment variable is 0 or 1, corresponding to different partitions respectively. After the baseboard management controller is started, the value of the environment variable is read first to start from the corresponding partition. Exemplarily, if the value of the environment variable is 0, it starts from partition 0.

[0023] After reading the value of the first register, if it is determined that a partition switch is needed, the value of the environment variable is modified to achieve the partition switch. If there is no need to switch partitions, the original value of the environment variable remains unchanged.

[0024] Specifically, after the baseboard management controller is started, the value of the first register is read at the bootloader (e.g., u-boot, a kind of bootloader) stage, so as to judge whether to switch partitions according to the value of the first register and set the corresponding environment variable. The bootloader stage is mainly used to initialize the hardware and boot the operating system, including hardware initialization, environment variable setting, etc. In the bootloader stage, the hardware watchdog function is enabled for process detection in subsequent stages.

[0025] Step S102, detect whether there is an abnormality in the baseboard management controller during the running stage.

[0026] The running stage involves the Kernel stage and the BMC user layer stage. Specifically, the Kernel stage includes processes such as kernel decompression and initialization, device driver loading, root file system mounting, user space initialization stage, etc. The BMC user layer stage includes the BMC process monitoring service self-start process. Each process is detected in the Kernel stage and the BMC user layer stage, and the watchdog mechanism can be started during detection. There is a hardware watchdog timer in the baseboard management controller, which is implemented by a dedicated WDT (Watchdog Timer) hardware module.

[0027] As an example, a timeout threshold can be set. After the baseboard management controller is started, the processes in the user layer stage regularly write a specific value to the WDT register to reset the watchdog counter for feeding the dog. If the process fails to feed the dog in time due to a fault, the watchdog counter will continue to accumulate until the preset timeout threshold is reached, then a hardware reset signal socreset will be triggered to restart.

[0028] In the Kernel stage, a notification function can be registered through the registered callback function mechanism. If an abnormal process occurs, the abnormal event is captured.

[0029] Step S103, if there is an abnormality, store the reset reason and the running process where the abnormality occurs in a specific field of the first register, and restart the baseboard management controller.

[0030] When an abnormal process is detected during the running stage, write the abnormal running process and the corresponding reset reason into a specific field of the first register. The reset reason includes the stage where the abnormal running process occurs, including kernel stage exceptions and user layer stage exceptions; the running processes include processes such as kernel decompression and initialization, device driver loading, root file system mounting, and user space initialization in the kernel stage, as well as the BMC process monitoring service startup process in the user layer stage.

[0031] As an example, mark the reset reason as 01, indicating a kernel stage exception, and mark the running process as 00001, indicating that there is an exception in the kernel decompression and initialization process.

[0032] The method for starting a baseboard management controller provided by an embodiment of the present invention includes, after the baseboard management controller is started, reading the value of the first register, and determining the target running partition and environment variables based on the specific field of the first register; detecting whether there is an abnormality in the baseboard management controller during the running stage; if there is an abnormality, store the reset reason in the specific field of the first register, and restart the baseboard management controller. The present invention detects whether there is an abnormality during the running stage after the baseboard management controller is started, stores the reset reason and the running process where the abnormality occurs in the first register after there is an abnormality, and performs a restart. The first register is a register that does not lose data during restart. After startup, by reading the value in the first register and combining the reset reason, it is determined whether to perform partitioning, so as to realize the intelligent switching of the redundant backup system.

[0033] In some optional embodiments, before reading the value of the first register in step S101, it further includes: Step S201, reading the environment variables; Step S202, determining the current running partition based on the environment variables.

[0034] The environment variables are used to represent the running partition when the baseboard management controller is started, and different environment variables correspond to different running partitions. As an example, the environment variable values are 0 and 1. After the baseboard management controller is started, first read the environment variables, and determine the current running partition according to the environment variables. If the value of the environment variable is 1, start from partition 1, and partition 1 is the current running partition.

[0035] In some optional embodiments, after restarting the baseboard management controller, the method further includes: Reset the first register, and the specific field of the reset first register represents that the baseboard management controller has no abnormality.

[0036] Reset the first register to clear the previously stored abnormal information, so that the numerical values of each specific field in the first register are default values, representing that the baseboard management controller has no abnormality.

[0037] In some alternative embodiments, the first specific field of the first register characterizes the running process in which an exception occurs. Determining the target running partition and the environment variable based on the specific field of the first register in step S101 includes: Step S301, if the first specific field indicates that there is a running process in which an exception occurs, then modify the environment variable from the first value to the second value, and switch the current running partition to the target running partition. Wherein, the first value corresponds to the current running partition, and the second value corresponds to the target running partition.

[0038] Taking a 32-bit register as an example, the [7:3] (bits 4-8) of the first register represent the running process in the running stage with five BIT bits. The first specific field corresponds to the BIT bits of bits 4-8, characterizing the running process in which an exception occurs, and the default value is 00000. When the value of the first specific field is 00000, it indicates that there is no running process in which an exception occurs. The running processes in the running stage sequentially include kernel decompression and initialization, device driver loading, root file system mounting, user space initialization, and BMC process monitoring service startup in the order of occurrence. Different running processes correspond to different bits among the five BIT bits. Exemplarily, if it is 00001, it indicates that there is an exception in kernel decompression and initialization.

[0039] First, read the first specific field in the first register. If the value of the first specific field is not the default value, that is, 00000, indicating that there is a running process in which an exception occurs, then the partition needs to be switched. The environment variable corresponding to the current running partition is the first value. By modifying the environment variable from the first value to the second value, the current running partition is switched to the target running partition corresponding to the second value.

[0040] Furthermore, determining the target running partition and the environment variable based on the specific field of the first register in step S101 further includes: Step S302, if the first specific field of the first register indicates that there is no running process in which an exception occurs, then read the second specific field and the third specific field of the first register, wherein the second specific field characterizes the reset reason, and the third specific field characterizes the mirror refresh state; Step S303, determine the target running partition and the environment variable based on the reading result.

[0041] If the first specific field of the first register is the default value, indicating that there is no running process in which an exception occurs in the previous startup before the current startup, that is, the restart of the baseboard management controller is not caused by the existence of an abnormal running process, then further read the second specific field and the third specific field of the first register.

[0042] Taking a 32-bit register as an example, [2:1] (bits 2-3) of the first register are used to mark the reset reason with two BITs. The second specific field corresponds to the BITs of bits 2-3. The default value of the second specific field is 00, where 00 indicates normal, 01 indicates a kernel exception, and 10 indicates a BMC exception (Baseboard Management Controller exception).

[0043] [8] (bit 9) of the first register is used to mark the mirror refresh status with one BIT. The third specific field corresponds to the BIT of bit 1. The default value of the third specific field is 0, where 0 indicates that the mirror refresh is completed or not refreshed, and 1 indicates that the mirror refresh is in progress or there is an exception in the mirror refresh.

[0044] Based on the read results of the second specific field and the third specific field, it is determined whether to switch the current running partition and modify the environment variables.

[0045] Further, the above step S303 includes: Step S3031, if the second specific field indicates the existence of a reset reason, then read the third specific field of the first register, and the third specific field indicates the mirror refresh status.

[0046] If the value of the second specific field is non-zero, including 01 or 10, it indicates the existence of a reset reason, specifically including a kernel exception and a BMC exception, and read the third specific field.

[0047] If the value of the second specific field is 0, it indicates that there is no reset reason, that is, there is no abnormal reset situation during the previous startup and operation process. Therefore, there is no need to switch partitions during this startup.

[0048] Step S3032, if the third specific field indicates an exception in the mirror refresh, then determine the current running partition as the target running partition.

[0049] If the value of the third specific field is non-zero, it indicates that there is an exception in the mirror refresh during the previous startup and operation process. Therefore, there is no need to switch partitions and modify the environment variables during this startup and operation process, and the current running partition is used as the target running partition.

[0050] Further, after reading the third specific field of the first register in step S3031, the method: if the third specific field indicates that there is no exception during the mirror refresh of the controller, then modify the environment variable from the first value to the second value and switch the current running partition to the target running partition.

[0051] If the value of the third specific field is 0, it indicates that there is no abnormality in the mirror refresh during the previous startup operation. Then, it is necessary to switch partitions. Modify the environment variable, change the current first value to the second value, so as to switch the current running partition corresponding to the first value to the target running partition corresponding to the second value. At the next startup, start from the target running partition by reading the modified environment variable.

[0052] In this embodiment, the partition switching logic is described in detail. Figure 2 It is a schematic diagram of the partition switching process. After the baseboard management controller starts up, read the value of the first register. First, read the first specific field and determine whether the register value in the running stage is non-zero. If the judgment result is yes, modify the environment variable from the first value to the second value; if the judgment result is no, further determine whether there is a reset reason. If the value corresponding to the reset reason is 0, it indicates that the system is normal and there is no need to switch partitions. If the value corresponding to the reset reason is non-zero, it indicates an abnormal reset. Then, further determine whether the mirror refresh is abnormal. If the corresponding value is 1, it indicates that the mirror refresh is abnormal, that is, there is an abnormality during the mirror refresh operation during the previous startup operation, and the mirror of the other partition may be damaged. Therefore, do not switch partitions and continue to use the current partition. If the mirror refresh is normal, switch partitions. Optionally, add a rollback mechanism. If it is designed for dual-mirror backup, mark the damaged target partition as "invalid", pull the previous version of the complete mirror from the remote server or local backup storage, overwrite and write it to the damaged partition, and verify the integrity of the mirror (such as hash verification). After passing, clear the mirror refresh flag bit. Through this method, it can be quickly restored when the mirror refresh fails, minimizing the risk of service interruption to the greatest extent.

[0053] In some alternative embodiments, the running stage includes the kernel stage, and step S102 includes: Step S401, detect whether there is an abnormal process in the process of the baseboard management controller in the kernel stage.

[0054] When running in the kernel stage, monitor the running status of the processes in the kernel stage. Specifically, through the notification function, capture the panic exception event in the callback. If an exception event is captured, it indicates that the corresponding process is abnormal.

[0055] Step S402, if there is an abnormal process, determine that there is an abnormality in the baseboard management controller in the kernel stage.

[0056] If there is an abnormal process in this stage, it is determined that there is an abnormality in the kernel stage, and the reset reason is a kernel exception.

[0057] Further, if it is determined that there is an abnormality in the baseboard management controller during the kernel stage, in step S103, storing the reset reason and the running process where the abnormality occurs in a specific field of the first register includes: modifying the first specific field and the second specific field of the first register, where the modified first specific field indicates that there is an abnormality in the process during the kernel stage, and the modified second specific field indicates that the reset reason of the baseboard management controller is a kernel abnormality.

[0058] If an abnormal process is detected during the kernel stage, the value of the first register is correspondingly modified. Specifically, taking a 32-bit register as an example, [7:3] (bits 4 - 8) of the first register represent the running process during the running stage with five BITs, and the first specific field corresponds to the BITs of bits 4 - 8, representing the running process where the abnormality occurs, with a default value of 00000. When the value of the first specific field is 00000, it indicates that there is no running process where an abnormality occurs. Among them, the running processes during the kernel stage sequentially include kernel decompression and initialization, device driver loading, root file system mounting, and user space initialization in the order of occurrence. Exemplarily, if it is 00001, it indicates that there is an abnormality in kernel decompression and initialization. [2:1] (bits 2 - 3) of the first register mark the reset reason with two BITs, and the second specific field corresponds to the BITs of bits 2 - 3. The default value of the second specific field is 00, where 00 indicates normal, 01 indicates a kernel abnormality, and 10 indicates a BMC abnormality (baseboard management controller abnormality).

[0059] In some alternative embodiments, the running stage includes the user layer stage, and step S102 includes: Step S403, detecting whether there is an abnormal process in the user layer stage.

[0060] The user layer stage is the environment where applications run. In the baseboard management controller, the user layer stage refers to the stage when various applications running on the baseboard management controller are executing. Detect whether there is an abnormality in the processes in the user layer stage. Specifically, the CPU usage rate, memory occupancy, disk I / O, etc. of the processes can be monitored. If a certain process occupies excessive resources for a long time, it may indicate that there is an abnormality in this process, such as memory leakage, infinite loop, etc. Detect abnormalities by analyzing the behavior patterns of the processes. For example, a normal management process should execute tasks according to a predetermined logic and frequency. If the behavior pattern of a certain process does not match the expectation, it may indicate that there is an abnormality in this process. Check the log files and error reports of the processes. Abnormal processes usually leave error messages or warnings in the logs.

[0061] Specifically, step S403 includes: detecting whether the processes in the user layer stage time out based on a preset counter; if it times out, it is determined that there is an abnormal process.

[0062] Step S404: If there is an abnormal process in the user layer stage, it is determined that there is an abnormality in the baseboard management controller in the user layer stage.

[0063] The pre-designed counter is the WDT counter, and the timeout threshold is set. When the baseboard management controller runs to the user layer stage, the user layer daemon process periodically writes a specific value to the WDT register to reset the timer and feed the dog. If there is a process abnormality and the dog cannot be fed in time, a timeout situation will occur, and the value of the pre-designed counter will continue to accumulate until the preset time threshold is reached. After the timeout, a hardware reset signal will be triggered for a restart operation.

[0064] If it is detected that there is an abnormal process in the user layer stage, the criterion determines that there is an abnormality in the user layer stage, that is, the reset reason is the baseboard management controller abnormality.

[0065] Furthermore, if it is determined that there is an abnormality in the baseboard management controller in the user layer stage, storing the reset reason and the running process where the abnormality occurs in the specific field of the first register in step S103 includes: modifying the first specific field and the second specific field of the first register. The modified first specific field indicates that there is an abnormality in the process in the user layer stage, and the modified second specific field indicates that the reset reason of the baseboard management controller is the baseboard management controller abnormality.

[0066] If it is detected that there is an abnormal process in the user layer stage, the value of the first register is correspondingly modified. Specifically, taking a 32-bit register as an example, [7:3] (bits 4-8) of the first register represent the running process in the running stage with five BITs. The first specific field corresponds to the BITs of bits 4-8 and represents the running process where the abnormality occurs, with a default value of 00000. When the value of the first specific field is 00000, it indicates that there is no running process where the abnormality occurs. Among them, the running process in the user layer stage includes the start of the BMC process monitoring service. If there is an abnormality, the corresponding value is 10000. [2:1] (bits 2-3) of the first register mark the reset reason with two BITs. The second specific field corresponds to the BITs of bits 2-3, and the default value of the second specific field is 00. 00 indicates normal, 01 indicates kernel abnormality, and 10 indicates BMC abnormality (baseboard management controller abnormality).

[0067] An embodiment of the present invention provides a method for starting a baseboard management controller. Figure 3 It is a schematic diagram of the running stage of the baseboard management controller. Described in chronological order, after the baseboard management controller is started, the running stage includes: During the period from system startup to the completion of its initialization tasks and readiness to load the kernel (Boot phase), in this phase, the values of environment variables are first read, and the current running partition is determined based on the values of the environment variables. For example, if the current value of the environment variable is 0, then start from partition 0. Then the value of the first register is read, and based on the read result, it is judged whether a partition switch is needed. If a partition switch is needed, the corresponding environment variable is set. The specific partition switch logic includes: first read the first specific field, and judge whether the register value in the running phase is non-zero. If the judgment result is yes, then modify the environment variable from the first value to the second value; if the judgment result is no, then further judge whether there is a reset reason. If the value corresponding to the reset reason is 0, it indicates that the system is normal and there is no need to switch partitions. If the value corresponding to the reset reason is non-zero, it indicates an abnormal reset, then further judge whether the image refresh is abnormal. If the corresponding value is 1, it indicates that the image refresh is abnormal, that is, an abnormality occurred during the image refresh operation during the previous startup and operation, and the mirror of the other partition may be damaged. Therefore, do not switch partitions and continue to use the current partition. If the image refresh is normal, then switch partitions.

[0068] In the kernel phase, it includes processes such as kernel decompression and initialization, device driver loading, root file system mounting, and user space initialization phase. Through the notifier mechanism, a panic notification function is registered (in a program or the operating system kernel, when an unrecoverable error or abnormal situation is encountered, a function or mechanism is triggered. This function will perform a series of operations, such as recording error information, attempting to clean up resources, notifying users or administrators, and ultimately causing the program or system to stop running). The Panic event is captured in the callback to detect whether there is an abnormal process. If there is, the reset reason and the abnormal process in the first register are modified, and the baseboard management controller is restarted. After restarting, the running partition is modified accordingly by reading the value of the first register. In the callback function, the value of the register is set.

[0069] In the BMC user layer phase, configure the BMC process monitoring service to start automatically, regularly write a specific value to the WDT count restart register to restart the WDT counter, thereby detecting critical processes. If an abnormal process is detected, save the reset reason as BMC exception to the first register and restart the baseboard management controller. After restarting, the running partition is modified accordingly by reading the value of the first register.

[0070] The method provided by the embodiment of the present invention realizes fault detection and automatic recovery through the collaborative monitoring of the hardware watchdog timer (WDT) and the application layer daemon process, avoiding long-term downtime of the system due to kernel / BMC exceptions; when the watchdog timeout triggers a hardware reset (socreset), the system can be forced to restart and resume operation based on a preset policy, ensuring service continuity. By not losing the BIT bit mark of the register (the first register) during the BMC restart, the reset reason (such as kernel exception, BMC exception) and the fault occurrence stage are accurately recorded, providing key data for subsequent fault diagnosis; combining the reset reason and environmental variables to automatically select the startup partition, realizing the intelligent switching of the redundant backup system and improving system availability. The hardware watchdog is deeply integrated with the BootLoader (boot loader), starting monitoring at the initial stage of system startup, covering the entire life cycle; dynamically controlling the startup partition through environmental variables, combined with the dual-partition redundancy design, reducing the impact of single-point failures on the system.

[0071] This method divides BIT bits in the first register to store corresponding information, including the reset reason, running process, etc. Since the first register is a register that is not lost during the BMC restart, persistent storage of hardware-level fault information can be achieved. Start the watchdog timer when the BMC just starts and set the corresponding timeout threshold, extending the monitoring scope to the initial stage of system startup, breaking through the limitation that traditional monitoring only acts on the application layer. The collaborative mechanism of the application layer daemon process and the hardware WDT ensures real-time status feedback when the system is running normally.

[0072] Through the dynamic partition switching mechanism of environmental variables, set environmental variables to realize dynamic control of dual-partition startup, combined with the BMC reset reason judgment logic, to achieve automatic isolation of the faulty partition and seamless switching of the backup partition.

[0073] From the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.

[0074] The embodiment of the present application also provides a substrate management controller startup device, as Figure 4 shown, including: A partition determination module, configured to, after the substrate management controller starts, read the value of the first register, and determine the target running partition and environmental variables based on a specific field of the first register. The environmental variables are used to represent the running partition when the substrate management controller starts. The specific field of the first register is at least used to represent the running state of the substrate management controller, and the running state at least includes the reset reason and the running process where an exception occurs; An operation detection module, configured to detect whether there is an abnormality in the baseboard management controller during the operation phase; A restart module, configured to, if an abnormality exists, store a reset reason and a running process with an abnormality in a specific field of the first register, and restart the baseboard management controller.

[0075] In some alternative embodiments, the apparatus further includes: A variable reading module, configured to read environment variables; A current partition determination module, configured to determine a current running partition based on the environment variables.

[0076] In some alternative embodiments, a first specific field of the first register represents a running process with an abnormality, and the partition determination module includes: A first switching unit, configured to, if the first specific field represents a running process with an abnormality, modify the environment variable from a first value to a second value, and switch the current running partition to a target running partition, where the first value corresponds to the current running partition and the second value corresponds to the target running partition.

[0077] In some alternative embodiments, the partition determination module further includes: A first reading unit, configured to, if the first specific field of the first register represents that there is no running process with an abnormality, read a second specific field and a third specific field of the first register, where the second specific field represents a reset reason and the third specific field represents a mirror refresh status; A partition determination unit, configured to determine a target running partition and environment variables based on the reading result.

[0078] In some alternative embodiments, the partition determination unit includes: A second reading subunit, configured to, if the second specific field represents that there is a reset reason, read the third specific field of the first register, where the third specific field represents a mirror refresh status; A partition switching subunit, configured to, if the third specific field represents a mirror refresh abnormality, determine the current running partition as the target running partition.

[0079] In some alternative embodiments, the partition determination unit includes: A value modification subunit, configured to, if the third specific field represents that there is no abnormality in the controller during mirror refresh, modify the environment variable from a first value to a second value, and switch the current running partition to a target running partition.

[0080] In some alternative embodiments, the operation phase includes a kernel phase, and the operation detection module includes: A kernel process detection unit for detecting whether there is an abnormal process in the process of the detection baseboard management controller in the kernel stage; A kernel exception unit for determining that there is an exception in the baseboard management controller in the kernel stage if there is an abnormal process.

[0081] In some alternative embodiments, if it is determined that there is an exception in the baseboard management controller in the kernel stage, the restart module includes: A first modification unit for modifying a first specific field and a second specific field of the first register, where the modified first specific field indicates that there is an abnormal process in the kernel stage, and the modified second specific field indicates that the reset reason of the baseboard management controller is a kernel exception.

[0082] In some alternative embodiments, the running stage includes a user layer stage, and the running detection module includes: A user layer detection unit for detecting whether there is an abnormal process in the user layer stage; A user layer exception unit for determining that there is an exception in the baseboard management controller in the user layer stage if there is an abnormal process in the user layer stage.

[0083] In some alternative embodiments, if it is determined that there is an exception in the baseboard management controller in the user layer stage, the restart module includes: A second modification unit for modifying a first specific field and a second specific field of the first register, where the modified first specific field indicates that there is an abnormal process in the user layer stage, and the modified second specific field indicates that the reset reason of the baseboard management controller is a baseboard management controller exception.

[0084] In some alternative embodiments, the user layer detection unit includes: A timeout detection subunit for detecting whether the process in the user layer stage times out based on a preset counter; A timeout determination subunit for determining that there is an abnormal process if it times out.

[0085] In some alternative embodiments, the device further includes: A reset module for resetting the first register, where the specific field of the reset first register indicates that the baseboard management controller has no exception.

[0086] For the description of the features in the corresponding embodiments of the baseboard management controller startup device, reference can be made to the relevant description of the corresponding embodiments of the baseboard management controller startup method, which will not be elaborated here one by one.

[0087] Embodiments of the present application further provide a computer device, such asFigure 5 As shown in Figure 5 , it includes a memory 10 and a processor 20. A computer program is stored in the memory 10, and the processor 20 is configured to run the computer program to execute the steps in any of the above-described embodiments of the method for starting a baseboard management controller.

[0088] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above-described embodiments of the method for starting a baseboard management controller when running.

[0089] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (ROMs), random access memories (RAMs), external hard drives, magnetic disks, or optical discs, etc., various media that can store computer programs.

[0090] An embodiment of the present application further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described embodiments of the method for starting a baseboard management controller.

[0091] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described embodiments of the method for starting a baseboard management controller.

[0092] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0093] The above has introduced in detail a method for starting a baseboard management controller provided by the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

Claims

1. A method for starting a baseboard management controller, characterized in that Including: After the baseboard management controller is started, read the value of the first register, and determine the target operating partition and the environment variable based on a specific field of the first register. The environment variable is used to characterize the operating partition when the baseboard management controller is started. The specific field of the first register is at least used to characterize the operating state of the baseboard management controller, and the operating state at least includes the reset reason and the running process where an exception occurs; Detect whether there is an exception during the running stage of the baseboard management controller; If there is an exception, store the reset reason and the running process where the exception occurs in the specific field of the first register, and restart the baseboard management controller.

2. The controller startup method according to claim 1, wherein Before reading the value of the first register, the method further includes: Read the environment variable; Determine the current operating partition based on the environment variable.

3. The controller startup method according to claim 2, wherein The first specific field of the first register characterizes the running process where an exception occurs. Determining the target operating partition and the environment variable based on the specific field of the first register includes: If the first specific field characterizes that there is a running process where an exception occurs, modify the environment variable from a first value to a second value, and switch the current operating partition to the target operating partition. The first value corresponds to the current operating partition, and the second value corresponds to the target operating partition.

4. The controller startup method according to claim 3, characterized in that, Determining the target operating partition and the environment variable based on the specific field of the first register further includes: If the first specific field of the first register characterizes that there is no running process where an exception occurs, read the second specific field and the third specific field of the first register. The second specific field characterizes the reset reason, and the third specific field characterizes the mirror refresh status; Determine the target operating partition and the environment variable based on the read result.

5. The controller startup method according to claim 4, wherein, Determining the target operating partition and the environment variable based on the read result includes: If the second specific field characterizes that there is a reset reason, read the third specific field of the first register. The third specific field characterizes the mirror refresh status; If the third specific field characterizes that there is an exception in mirror refresh, determine the current operating partition as the target operating partition.

6. The controller startup method according to claim 5, characterized in that, After reading the third specific field of the first register, the method further includes: If the third specific field characterizes that there is no exception when the controller performs mirror refresh, modify the environment variable from a first value to a second value, and switch the current operating partition to the target operating partition.

7. The controller startup method according to claim 1, wherein The running stage includes the kernel stage. Detecting whether there is an exception during the running stage of the baseboard management controller includes: Detect whether there is an abnormal process in the process of the baseboard management controller in the kernel stage; If there is an abnormal process, determine that there is an exception in the baseboard management controller in the kernel stage.

8. The controller startup method according to claim 7, wherein, If it is determined that there is an exception in the baseboard management controller in the kernel stage, storing the reset reason and the running process where the exception occurs in the specific field of the first register includes: Modify the first specific field and the second specific field of the first register, where the modified first specific field indicates that there is an exception in the process during the kernel stage, and the modified second specific field indicates that the reset reason of the baseboard management controller is a kernel exception.

9. The controller startup method according to claim 1, characterized in that, The running stage includes the user layer stage. Detecting whether there is an exception in the running stage of the baseboard management controller includes: Detecting whether there is an abnormal process in the user layer stage; If there is an abnormal process in the user layer stage, it is determined that there is an exception in the baseboard management controller in the user layer stage.

10. The method for starting a controller according to claim 9, wherein, If it is determined that there is an exception in the baseboard management controller in the user layer stage, storing the reset reason and the running process where the exception occurs in the specific field of the first register includes: Modify the first specific field and the second specific field of the first register, where the modified first specific field indicates that there is an exception in the process during the user layer stage, and the modified second specific field indicates that the reset reason of the baseboard management controller is a baseboard management controller exception.

11. The method for starting a controller according to claim 9, wherein, Detecting whether there is an abnormal process in the user layer stage includes: Detecting whether the process in the user layer stage times out based on a preset counter; If it times out, it is determined that there is an abnormal process.

12. The controller startup method according to claim 1, wherein After restarting the baseboard management controller, the method further includes: Reset the first register, and the specific field of the reset first register indicates that the baseboard management controller has no exception.

13. A computer device, characterized in that, Includes: A memory for storing a computer program; A processor for implementing the steps of the baseboard management controller startup method according to any one of claims 1 to 12 when executing the computer program.

14. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, where the computer program, when executed by a processor, implements the steps of the baseboard management controller startup method according to any one of claims 1 to 12.

15. A computer program product, comprising a computer program, characterized in that, The computer program, when executed by a processor, implements the steps of the baseboard management controller startup method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • BMC (Baseboard Management Controller) starting method and equipment based on double BMC FLASH chips

    CN111158764A

  • Firmware loading method, system and equipment of BMC and medium

    CN113867739A

  • Double-BMC Flash optimization upgrading method and device, equipment and medium

    CN115167903A

  • System starting method, system, equipment and medium

    CN117034296A

  • Embedded system restarting method and device, electronic equipment and readable storage medium

    CN117667258A

Cited By

  • Management controller, management controller starting method and device, and storage medium

    CN120429026A

  • Fan management and control method and device in BMC starting process, equipment and medium

    CN121478358A

  • Mirror image switching method and electronic equipment

    CN122220148A

  • Mirror switching method and electronic device

    CN122220148B