Fault identification method and device
By receiving LPC signals in the BMC and querying Post Code, identifying the failures before and after the ABL phase during the BIOS startup process, the problem of the lack of fault diagnosis capabilities in the existing BIOS during the startup process is solved, and rapid fault identification and business recovery are achieved.
Patent Information
- Application Number
- CN202411223166.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-02
- Publication Date
- 2025-06-03
AI Technical Summary
During the startup process of the existing BIOS, the ABL phase lacks fault diagnosis capabilities, and only supports memory read and write fault detection, and lacks the ability to detect CPU-related faults.
By receiving the low-pin number LPC signal sent by the complex convertible logic device CPLD in the substrate management controller BMC, obtaining the current power supply status, and querying the power-on self-test coded Post Code within the preset time to identify the BIOS's pre-start fault and the post-start fault.
It realizes the identification of faults before and after the ABL stage during the BIOS startup process, solves the problem of fault-free information reporting, helps users to understand fault information in a timely manner and quickly recover business, and enhances the brand influence and reputation of the server.
Smart Images

Figure CN120086074A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technologies, and in particular, to a fault identification method and device. Background Art
[0002] In the IT industry, during the startup process of a server's Basic Input Output System (BIOS), the BIOS only has the ability to diagnose faults when it loads the program after the Agent-Based Boot Loader (ABL).
[0003] However, in actual applications, CPU or memory faults often occur, resulting in the BIOS being unable to load after the ABL. At this time, the server is in a black screen state, and it cannot diagnose the source of the fault itself, nor can the user perceive any fault information. In this case, the user troubleshoots the fault by unplugging and plugging components one by one, and the fault repair time is relatively long, which poses a great challenge to business recovery.
[0004] Meanwhile, after the BIOS loads the ABL, the BIOS program starts to run, and fault detection basically depends entirely on the BIOS. However, the types of faults that the BIOS supports detecting are very limited. It can detect memory-related faults (e.g., memory read / write faults), but lacks the ability to detect CPU-related faults.
[0005] In summary, the following defects are exposed during the existing BIOS startup process: 1) Before the ABL stage, the BIOS does not have the ability to diagnose faults; 2) After the ABL stage, it only supports the detection of memory read / write faults, the type of fault detection is relatively single, and it lacks the ability to detect CPU-related faults. Summary of the Invention
[0006] In view of this, this application provides a fault identification method and device to solve the problems that the BIOS does not have the ability to diagnose faults and the type of fault detection is relatively single before and after the ABL stage during the existing BIOS startup process.
[0007] In a first aspect, this application provides a fault identification method, which is applied to a Baseboard Management Controller (BMC). The BMC is inside a server, and the server also includes a Complex Programmable Logic Device (CPLD). The method includes:
[0008] Receiving a Low Pin Count (LPC) signal sent by the CPLD;
[0009] Obtaining the current power supply state from the CPLD register according to the LPC signal;
[0010] If the current power supply state indicates that the server is in the powered-on state, query the Power-On Self-Test Code (Post Code) within a preset continuous time period.
[0011] If the Post Code is empty or in an unupdated state within the preset continuous time period, determine a pre-boot failure of the Basic Input / Output System (BIOS) that the server is running.
[0012] In a second aspect, the present application provides a fault identification device. The device is applied to a Baseboard Management Controller (BMC), the BMC is within a server, and the server further includes a Complex Programmable Logic Device (CPLD). The device includes:
[0013] A receiving unit, configured to receive a Low Pin Count (LPC) signal sent by the CPLD.
[0014] An obtaining unit, configured to obtain the current power supply state from a CPLD register according to the LPC signal.
[0015] A querying unit, configured to query the Power-On Self-Test Code (Post Code) within a preset continuous time period if the current power supply state indicates that the server is in the powered-on state.
[0016] A determining unit, configured to determine a pre-boot failure of the Basic Input / Output System (BIOS) that the server is running if the Post Code is empty or in an unupdated state within the preset continuous time period.
[0017] In a third aspect, the present application provides a network device, including a processor and a machine-readable storage medium. The machine-readable storage medium stores machine-executable instructions that can be executed by the processor, and the processor is prompted by the machine-executable instructions to execute the method provided in the first aspect of the present application.
[0018] Therefore, by applying the fault identification method and device provided in the present application, the BMC receives the Low Pin Count (LPC) signal sent by the CPLD; according to the LPC signal, the BMC obtains the current power supply state from the CPLD register; if the current power supply state indicates that the server is in the powered-on state, the BMC queries the Power-On Self-Test Code (Post Code) within a preset continuous time period; if the Post Code is empty or in an unupdated state within the preset continuous time period, the BMC determines a pre-boot failure of the Basic Input / Output System (BIOS) that the server is running.
[0019] In this way, by combining the existing LPC signals and Post Codes, the BMC identifies BIOS failures that occur before the ABL stage. Also, by judging the status values of the registers in the register set, the BMC identifies BIOS failures that occur after the ABL stage. This solves the problem of no fault information reporting such as system crashes and black screens during the BIOS startup process, helps users or administrators promptly understand the fault information, guides them to replace faulty components in a timely manner, quickly resume operations, and enhances the brand influence and reputation of the server. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 is a flowchart of the fault identification method provided by an embodiment of the present application;
[0021] Figure 2 is a structural diagram of the fault identification device provided by an embodiment of the present application;
[0022] Figure 3 is the hardware structure of the network device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0024] The terms used in the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. The singular forms "a", "the", and "said" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the corresponding listed items.
[0025] It should be understood that although the terms first, second, third, etc. may be used in the present application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to a determination".
[0026] The following provides a detailed description of the fault identification method provided by an embodiment of the present application. Refer to Figure 1, Figure 1 It is a flowchart of the fault identification method provided by the embodiment of the present application. This method is applied to a Baseboard Management Controller (BMC for short), which is inside a server. The server also includes a Complex Programmable Logic Device (CPLD). The server is configured with a Basic Input / Output System (BIOS). The fault identification method provided by the embodiment of the present application may include the following steps.
[0027] Step 110: Receive the Low Pin Count (LPC) signal sent by the CPLD;
[0028] Specifically, after the server is powered on, its power-on and power-off states can be sensed by the CPLD. The CPLD is a programmable logic device on the motherboard, responsible for the power-on timing and logic control of the motherboard. When the power state of the server changes, the CPLD can notify the BMC of the change in the power state of the server through the LPC signal.
[0029] In the embodiment of the present application, after the power state of the server changes, the CPLD generates an LPC signal and sends it to the BMC. The BMC receives the above LPC signal. For example, the change in the power state of the server can be specifically from power-on to power-off; from power-off to power-on.
[0030] Step 120: Obtain the current power supply state from the CPLD register according to the LPC signal;
[0031] Specifically, according to the description in Step 110, after the BMC receives the LPC signal, it determines the change in the power state of the server. The BMC accesses the CPLD register and obtains the current power supply state from it.
[0032] It can be understood that after the CPLD senses the change in the power state of the server, it writes the current power supply state into the CPLD register.
[0033] Step 130: If the current power supply state indicates that the server is in the power-on state, query the Power-On Self-Test Code (Post Code) within a preset continuous time;
[0034] Specifically, according to the description in Step 120, after the BMC obtains the current power supply state from the CPLD register, it identifies the current power supply state.
[0035] If the current power supply state indicates that the server is in the power-on state, the BMC queries the Power-On Self-Test Code (Post Code) within a preset continuous time.
[0036] Among them, the above preset continuous time can be specifically 5 consecutive minutes.
[0037] It should be noted that after the server is powered on, the BIOS starts immediately. The main sequence of the BIOS startup process is as follows: Power-On Self Test (POST), initialize BIOS settings, identify and configure hardware, determine the boot device order, load the boot program, and transfer control to the operating system.
[0038] If each of the above parts works properly during the BIOS startup process, after the status of each part or each sub-part included in each part changes, the BIOS generates and sends a Post Code to the BMC to feedback the startup status of the BIOS. After the BIOS is fully started, the BIOS stops sending the Post Code to the BMC.
[0039] For example, during the Power-On Self Test, after POST checks each hardware component, if the hardware component can work properly, its status has changed from the non-working state to the working state. At this time, the BIOS generates and sends a Post Code to the BMC to feedback the status change of the hardware component. If the entire POST part is started successfully, its status has changed from the non-working state to the working state. At this time, the BIOS generates and sends a Post Code to the BMC to feedback the status change of the POST part.
[0040] The following briefly describes the BIOS startup process:
[0041] a. Power-On Self Test: When the server power is turned on, the BIOS first executes POST. POST will check the hardware components of the server. For example, memory, keyboard, monitor, hard disk, etc., to ensure that they can work properly.
[0042] b. Initialize BIOS settings: If POST passes successfully, the BIOS will start to initialize the settings of various hardware components. For example, initialize the system clock, set the memory controller, check and set peripheral devices (such as keyboard and monitor).
[0043] c. Identify and configure hardware: The BIOS will then identify and configure all hardware devices, including CPU, memory, hard disk, optical drive, USB devices, etc.
[0044] d. Determine the boot device order (Boot Sequence): The BIOS will search for bootable devices according to the pre-set boot device order (usually can be adjusted through the BIOS setup interface). The boot order usually includes: floppy drive, optical drive, hard disk, network boot, etc.
[0045] e. Loading the bootloader: The BIOS loads the bootloader (such as GRUB, Windows Boot Manager, etc.) from the first bootable device (e.g., hard disk). The common bootloaders include GRUB, Windows Boot Manager, etc.
[0046] The bootloader is usually located in the Master Boot Record (MBR) or Partition Boot Record (PBR) of the hard disk.
[0047] f. Transferring control to the operating system: If the bootloader is successfully loaded and run, the task of the BIOS is basically completed. Then, the bootloader loads the core files of the operating system and transfers control to the operating system to complete the remaining startup process.
[0048] Step 140: If the Post Code is empty or in an unupdated state within the preset continuous time, determine the pre-startup failure of the BIOS running on the server.
[0049] Specifically, according to the description of Step 130, during the process of the BMC querying the Post Code within 5 consecutive minutes, if within 5 consecutive minutes, the queried Post Code is empty or in an unupdated state (the Post Code status has not changed within 5 consecutive minutes), the BMC determines the pre-startup failure of the BIOS running on the server.
[0050] Optionally, if the queried Post Code is non-empty or in an updated state within 5 consecutive minutes, the BMC determines that the pre-startup of the BIOS running on the server is normal.
[0051] Optionally, after the BMC determines the BIOS failure of the server, it generates a first warning message. The BMC displays the first warning message through a display device to prompt the user that a pre-startup failure has occurred during the BIOS startup process. Among them, the first warning message can be specifically in the form of a System Event Log (SEL).
[0052] It can be understood that the first warning message can include the cause of the failure, the location of the failure, the time of the failure, etc., so that the user can replace the faulty component through the warning message and quickly resume business operation.
[0053] It should be noted that during the BIOS startup process, the BIOS startup process is divided into the early BIOS startup and the late BIOS startup. Among them, steps a - d are called the early BIOS startup, that is, before the ABL stage; steps e - f are called the late BIOS startup, that is, after the ABL stage. In the embodiments of the present application, the BIOS failures that occur before the ABL stage can be identified through the process described in steps 110 - 140; the BIOS failures that occur after the ABL stage can be identified through the process described in the subsequent embodiments.
[0054] Therefore, by applying the fault identification method provided in the present application, the BMC receives the low - pin - count LPC signal sent by the CPLD; according to the LPC signal, the BMC obtains the current power supply state from the CPLD register; if the current power supply state indicates that the server is in the powered - on state, the BMC queries the power - on self - test code Post Code within a preset time; if the Post Code is empty or in an unupdated state within the preset time, the BMC determines a pre - startup fault of the basic input / output system BIOS of the server.
[0055] In this way, by combining the existing LPC signal and Post Code, the BMC identifies the BIOS failures that occur before the ABL stage. And, by judging the status values of each register in the register group, the BMC identifies the BIOS failures that occur after the ABL stage. It solves the problem of no fault information reporting such as system crash and black screen during the BIOS startup process, helps users or administrators to timely understand the fault information, guides users or administrators to timely replace the faulty components, quickly resume business, and enhances the brand influence and reputation of the server.
[0056] Optionally, the foregoing embodiments describe the process of how BIOS failures are identified when they occur before the ABL stage during the BIOS startup process. In the embodiments of the present application, it also includes the process of how BIOS failures are identified after they occur during the BIOS startup process and after the ABL stage.
[0057] Specifically, when the BIOS startup process enters step e, the CPU and memory included in the server have both been powered on. At this time, the synchronous flood (SYNCFLOOD) register group inside the CPU will record the CPU - related status information.
[0058] Every preset period (for example, 30 seconds), the BMC reads the status value of each register included in the SYNCFLOOD register group through the Advanced Platform Management Link (APML) channel.
[0059] Based on the status value of each register, the BMC determines whether the CPUs or memories included in the server are faulty.
[0060] Optionally, the above SYNCFLOOD register set includes a NorthBridge Input / Output (NBIO)_Syncflood register, an Input / Output Die (IOD) register, a Cluster-on-Die (COD) register, a Core Complex (CXX) register, a Graphics Output Protocol (GOP) register, and a Round-Robin Write (RRW) register;
[0061] The specific process of the BMC determining whether the CPUs or memories included in the server are faulty according to the status value of each register is as follows:
[0062] The BMC identifies whether the status value of the NBIO_Syncflood register is 0. If the status value of the NBIO_Syncflood register is 0, the BMC determines that the BIOS running on the server is normal, that is, the current CPUs and memories are all normal. If the status value of the NBIO_Syncflood register is non-0, the BMC continues to identify whether the status values of the IOD register, the COD register, the CXX register, and the GOP register are all 0.
[0063] If any one of the status values of the IOD register, the COD register, the CXX register, and the GOP register is non-0, the BMC determines that the CPU is faulty.
[0064] Under the condition that the status value of the NBIO_Syncflood register is non-0, if the status values of the IOD register, the COD register, the CXX register, and the GOP register are all 0, the BMC continues to determine whether the status value of the RRW register is 0.
[0065] If the status value of the RRW register is non-0, the BMC determines that the memory is faulty.
[0066] If the status value of the RRW register is 0, the BMC determines that the BIOS running on the server is normal, that is, the current CPUs and memories are all normal.
[0067] The following gives a brief description of each register in the SYNCFLOOD register set.
[0068] NBIO_SYNCFLOOD Register: Responsible for handling data transfer tasks connecting the CPU and high-speed I / O devices. Read instruction: GET_NBIO_SYNCFLOOD_STATUS(0x1A); Function description: Read the status value of NBIO's Syncflood; Return value: 0 indicates that no SYNCFLOOD is generated by NBIOS, and non-0 indicates that SYNCFLOOD is generated.
[0069] IOD Register: Responsible for the input / output tasks of the processor and memory control. Read instruction: READ_IOD_BIST_RESULT(0x13); Function description: Read the status value of the IOD register; Return value: 0 indicates Pass, and non-0 indicates Fail.
[0070] COD Register: A strategy that divides the cores of a multi-core processor into multiple smaller clusters to optimize performance and memory latency. Read instruction: READ_COD_BIST_RESULT(0x14); Function description: Read the status value of the COD register; Return value: 0 indicates Pass, and non-0 indicates Fail.
[0071] CXX Register: Represents a core complex. It is part of the AMD Zen processor architecture. Each CCX contains multiple CPU cores and shared cache resources (e.g., L3 cache). Read instruction: READ_CXX_BIST_RESULT(0x15); Function description: Read the status value of the CCX register; Return value: 0 indicates Pass, and non-0 indicates Fail.
[0072] GOP Register: GPO is a standard protocol used in the UEFI (Unified Extensible Firmware Interface) firmware environment, specifically for the initialization and management of graphics output devices. Read instruction: GET_GOP_LINK_FAILURE(0x19); Function description: Read whether the status value of the GOP register is abnormal; Return value: 0 indicates Pass, and non-0 indicates Fail.
[0073] RRW Register: A scheduling strategy or data writing strategy used to allocate resources or balance the load. Read instruction: GET_RRW_FAILURE(0x1B); Function description: Read whether there are failure problems with the RRW register; Return value: 0 indicates no RRW failure problems, and non-0 indicates there are RRW failure problems.
[0074] Optionally, in the embodiments of the present application, after the BMC determines CPU and memory failures, it also generates alarm messages. For example, a second alarm message indicating a CPU failure and a third alarm message indicating a memory failure. The BMC displays the second alarm message and the third alarm message through a display device to prompt the user that the CPU and memory have failed during the BIOS startup process, that is, a late startup failure has occurred during the BIOS startup process. Among them, both the second alarm message and the third alarm message can be specifically in the form of SEL.
[0075] It can be understood that the second alarm message and the third alarm message may include the cause of the failure, the location of the failure, the time of the failure, etc., so that the user can replace the faulty component through the alarm message and quickly resume business operation.
[0076] Based on the same inventive concept, the embodiments of the present application also provide a fault identification device corresponding to the fault identification method. See Figure 2 , Figure 2 This is the fault identification device provided by the embodiments of the present application. The device is applied to a baseboard management controller BMC, and the BMC is inside a server. The server also includes a complex programmable logic device CPLD. The device includes:
[0077] A receiving unit 210, configured to receive a low pin count LPC signal sent by the CPLD;
[0078] An obtaining unit 220, configured to obtain the current power supply state from a CPLD register according to the LPC signal;
[0079] A query unit 230, configured to query a power-on self-test code Post Code within a preset continuous time if the current power supply state indicates that the server is in a powered-on state;
[0080] A determining unit 240, configured to determine a pre-startup failure of the basic input / output system BIOS running on the server if the Post Code is empty or in an unupdated state within the preset continuous time.
[0081] Optionally, the determining unit 240 is further configured to determine that the pre-startup of the BIOS running on the server is normal if the Post Code is non-empty or in an updated state within the preset continuous time.
[0082] Optionally, the device further includes:
[0083] A reading unit (not shown in the figure), configured to read the status value of each register included in a synchronous flood SYNCFLOOD register group through an advanced platform management link APML channel every preset period;
[0084] A judgment unit (not shown in the figure) is configured to judge whether the CPU or the memory included in the server is faulty according to the status value of each register.
[0085] Optionally, the SYNCFLOOD register group includes a North Bridge Input / Output NBIO_Syncflood register, an Input / Output Chip IOD register, an On-chip Cluster COD register, a Core Complex CXX register, a Graphics Output Protocol GOP register, and a Polled Write RRW register;
[0086] The judgment unit (not shown in the figure) is specifically configured to determine that the CPU is faulty if the status value of the NBIO_Syncflood register is non-zero and the status value of any one of the IOD register, the COD register, the CXX register, and the GOP register is non-zero;
[0087] If the status value of the NBIO_Syncflood register is non-zero and the status values of the IOD register, the COD register, the CXX register, and the GOP register are all zero, then judge whether the status value of the RRW register is zero;
[0088] If the status value of the RRW register is non-zero, then determine that the memory is faulty.
[0089] Optionally, the device further includes:
[0090] A display unit (not shown in the figure) is configured to display a first warning message for the BIOS fault; and / or; display a second warning message for the CPU fault; and / or; display a third warning message for the memory fault.
[0091] Therefore, by applying the fault identification device provided in this application, the BMC receives a Low Pin Count LPC signal sent by the CPLD; according to the LPC signal, the BMC obtains the current power supply state from the CPLD register; if the current power supply state indicates that the server is in the power-on state, the BMC queries the Power-On Self-Test Code Post Code within a preset continuous time; if the Post Code is empty or in an unupdated state within the preset continuous time, the BMC determines a pre-boot fault of the Basic Input / Output System BIOS of the server.
[0092] In this way, by combining the existing LPC signals and Post Codes, the BMC identifies BIOS failures that occur before the ABL stage. Moreover, by judging the status values of the registers in the register bank, the BMC identifies BIOS failures that occur after the ABL stage. This solves the problem of no fault information reporting such as system halting and black screen during the BIOS startup process, helping users or administrators to timely understand the fault information, guiding them to replace the faulty components in time, quickly resume operations, and enhancing the brand influence and reputation of the server.
[0093] Based on the same inventive concept, an embodiment of the present application further provides a network device, such as Figure 3 shown, including a processor 310, a transceiver 320, and a machine-readable storage medium 330. The machine-readable storage medium 330 stores machine-executable instructions that can be executed by the processor 310, and the processor 310 is prompted by the machine-executable instructions to execute the fault identification method provided by the embodiment of the present application. The foregoing Figure 2 shown fault identification device can be implemented by using the hardware structure of the network device as shown in Figure 2 shown.
[0094] The above computer-readable storage medium 330 may include a random access memory (English: Random Access Memory, abbreviated as: RAM), and may also include a non-volatile memory (English: Non-volatile Memory, abbreviated as: NVM), such as at least one disk memory. Optionally, the computer-readable storage medium 330 may also be at least one storage device located far from the foregoing processor 310.
[0095] The above processor 310 may be a general-purpose processor, including a central processing unit (English: Central Processing Unit, abbreviated as: CPU), a network processor (English: Network Processor, abbreviated as: NP), etc.; it may also be a digital signal processor (English: Digital Signal Processor, abbreviated as: DSP), an application-specific integrated circuit (English: Application Specific Integrated Circuit, abbreviated as: ASIC), a field-programmable gate array (English: Field-Programmable Gate Array, abbreviated as: FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0096] In the embodiments of the present application, the processor 310 reads machine-executable instructions stored in the machine-readable storage medium 330, and is prompted by the machine-executable instructions to enable the processor 310 itself and to call the transceiver 320 to execute the fault identification method described in the foregoing embodiments of the present application.
[0097] In addition, the embodiments of the present application provide a machine-readable storage medium 330, which stores machine-executable instructions. When the machine-executable instructions are called and executed by the processor 310, the machine-executable instructions prompt the processor 310 itself and to call the transceiver 320 to execute the fault identification method described in the foregoing embodiments of the present application.
[0098] For the implementation processes of the functions and roles of each unit in the above device, please refer to the implementation processes of the corresponding steps in the above method for details, which will not be elaborated here.
[0099] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial descriptions of the method embodiments. The device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present application. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0100] For the embodiments of the fault identification device and the machine-readable storage medium, since the method content involved is basically similar to the foregoing method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial descriptions of the method embodiments.
[0101] The above are only the preferred embodiments of the present application, and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A fault identification method, characterized in that: The method is applied to a baseboard management controller BMC, the BMC is in a server, and the server further includes a complex programmable logic device CPLD, and the method includes: Receiving a low pin count LPC signal sent by the CPLD; According to the LPC signal, obtaining the current power supply status from the CPLD register; If the current power supply status indicates that the server is powered on, querying a power-on self-test code Post Code within a preset continuous time; If the Post Code is empty or not updated within the preset continuous time, it is determined that an early startup failure of the basic input and output system BIOS running on the server occurs.
2. The method according to claim 1, characterized in that The method further comprises: If the Post Code is not empty or is in an updated state within the preset continuous time, it is determined that the early startup of the BIOS running on the server is normal.
3. The method according to claim 1, characterized in that The method further comprises: At every preset cycle, the state value of each register included in the SYNCFLOOD register group is read through the advanced platform management link APML channel; According to the status value of each register, it is determined whether the CPU or memory included in the server is faulty.
4. The method according to claim 3, characterized in that The SYNCFLOOD register group includes a north bridge input / output NBIO_Syncflood register, an input / output chip IOD register, an on-chip cluster COD register, a core complex CXX register, a graphics output protocol GOP register, and a polling write RRW register; The determining, according to the status value of each register, whether a CPU or a memory included in the server is faulty specifically includes: If the status value of the NBIO_Syncflood register is non-0 and any one of the status values of the IOD register, the COD register, the CXX register, and the GOP register is non-0, it is determined that the CPU is faulty; If the status value of the NBIO_Syncflood register is non-0 and the status values of the IOD register, the COD register, the CXX register, and the GOP register are all 0, then determine whether the status value of the RRW register is 0; If the status value of the RRW register is non-zero, the memory fault is determined.
5. The method according to any one of claims 1 to 4, characterized in that: The method further comprises: Displaying the first warning information of the BIOS failure; and / or; Displaying the second alarm information of the CPU failure; and / or; The third alarm information of the memory failure is displayed.
6. A fault identification device, characterized in that: The device is applied to a baseboard management controller BMC, the BMC is in a server, the server also includes a complex programmable logic device CPLD, and the device includes: A receiving unit, used for receiving a low pin count LPC signal sent by the CPLD; An acquisition unit, used for acquiring a current power supply state from a CPLD register according to the LPC signal; A query unit, configured to query a power-on self-check code Post Code within a preset continuous time if the current power supply status indicates that the server is in a powered-on state; The determining unit is configured to determine an early startup failure of a basic input / output system BIOS running on the server if the Post Code is empty or in an unupdated state within the preset continuous time.
7. The device according to claim 6, characterized in that The determining unit is further configured to determine that the early startup of the BIOS running on the server is normal if the Post Code is not empty or is in an updated state within the preset continuous time.
8. The device according to claim 6, characterized in that The device also includes: A reading unit, used for reading the status value of each register included in the SYNCFLOOD register group through the advanced platform management link APML channel at every preset cycle; The judgment unit is used to judge whether the CPU or memory included in the server is faulty according to the status value of each register.
9. The device according to claim 8, characterized in that The SYNCFLOOD register group includes a north bridge input / output NBIO_Syncflood register, an input / output chip IOD register, an on-chip cluster COD register, a core complex CXX register, a graphics output protocol GOP register, and a polling write RRW register; The judgment unit is specifically configured to determine that the CPU is faulty if the status value of the NBIO_Syncflood register is non-0 and any one of the status values of the IOD register, the COD register, the CXX register, and the GOP register is non-0; If the status value of the NBIO_Syncflood register is non-0 and the status values of the IOD register, the COD register, the CXX register, and the GOP register are all 0, then determine whether the status value of the RRW register is 0; If the status value of the RRW register is non-zero, the memory fault is determined.
10. The device according to any one of claims 6 to 9, characterized in that: The device also includes: A display unit is used to display the first alarm information of the BIOS failure; and / or; display the second alarm information of the CPU failure; and / or; display the third alarm information of the memory failure.