Error processing method, computing apparatus, server, device, medium, and product
Patent Information
- Application Number
- PCT/CN2025/139278
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-26
- Filing Date
- 2025-12-02
- Publication Date
- 2026-10-01
Smart Images

Figure CN2025139278_01102026_PF_FP_ABST
Abstract
Description
Error handling methods, computing devices, servers, equipment, media, and products
[0001] This disclosure claims priority to Chinese Patent Application No. 202510373333.7, filed on March 26, 2025, entitled "Error Handling Method, Computing Device, Server, Equipment, Media and Product", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This disclosure relates to the field of computer technology, and more particularly to an error handling method, computing device, server, equipment, medium, and product. Background Technology
[0003] With the widespread adoption of Artificial Intelligence (AI) and cloud computing technologies, cloud servers have become increasingly prevalent. Against this backdrop, customers are placing higher demands on server hardware failure handling mechanisms, requiring not only timely fault detection but also low-latency, high-efficiency fault resolution. Currently, the number of Central Processing Unit (CPU) cores in mainstream servers is increasing, and server architectures are becoming more complex, leading to a relatively higher probability of Correctable Errors (CEs) occurring in CPU cores. Although CEs can be corrected promptly, the server triggers a System Management Interrupt (SMI) upon detecting such faults, suspending the execution of the server's operating system and applications until exiting the SMI. This fault handling method can cause system jitter and latency. Summary of the Invention
[0004] This disclosure provides an error handling method, a server, an electronic device, a computer-readable storage medium, and a computer program product to alleviate or solve one or more technical problems existing in the prior art.
[0005] In a first aspect, embodiments of this disclosure provide a computing device, comprising: at least one processor for running an operating system of the computing device; a bridge chip connected to the processor, the bridge chip integrating a basic input / output system, the basic input / output system being configured to: trigger a set signal of a target pin in response to the presence of a target processor experiencing a core correctable error among the at least one processor, the set signal of the target pin indicating the presence of the target processor; a baseboard management controller connected to the bridge chip via the target pin, the baseboard management controller being configured to: determine the target processor from the at least one processor in response to the set signal of the target pin, record information about the core correctable error occurring in the target processor, and send a system control interrupt to the operating system; the operating system being configured to: process the core correctable error of the target processor in response to the system control interrupt.
[0006] Secondly, embodiments of this disclosure provide an error handling method applied to a baseboard management controller of a computing device, the computing device further including at least one processor for running an operating system of the computing device, the method comprising: determining, in response to a set signal of a target pin, a target processor from the at least one processor for which a core correctable error has occurred, wherein the set signal of the target pin is used to indicate the presence of the target processor; recording information about the core correctable error occurring in the target processor; and sending a system control interrupt to the operating system, the system control interrupt being used to notify the operating system to handle the core correctable error of the target processor.
[0007] Thirdly, embodiments of this disclosure provide an error handling method applied to the basic input / output system of a computing device, the computing device further including a baseboard management controller and at least one processor, the at least one processor being used to run an operating system of the computing device, the method comprising: in response to the presence of a target processor among the at least one processors that has experienced a core correctable error, triggering a set signal of a target pin, wherein the set signal of the target pin is used to indicate the presence of the target processor, so that the baseboard management controller records information of the core correctable error occurring in the target processor, and sending a system control interrupt to the operating system, the system control interrupt being used to notify the operating system to handle the core correctable error of the target processor.
[0008] Fourthly, embodiments of this disclosure provide a server, including the computing device and memory described in any one of the embodiments of this disclosure.
[0009] Fifthly, embodiments of this disclosure provide an electronic device, including a memory, a processor, and a computer program stored in the memory, wherein the processor implements the method of any one of the embodiments of this disclosure when executing the computer program.
[0010] Sixthly, embodiments of this disclosure provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the method of any one of the embodiments of this disclosure.
[0011] In a seventh aspect, embodiments of this disclosure provide a computer program product, including a computer program that, when executed by a processor, implements the method of any one of the embodiments of this disclosure.
[0012] According to the technical solution of this disclosure embodiment, after the BIOS detects the target processor (i.e., the processor that has experienced a core correctable error), it triggers a set signal on the target pin to indicate the presence of a processor with a core correctable error. After detecting the set signal on the target pin, the BMC records the information about the core correctable error of the target processor and triggers the SCI (System Change Component) to enable the operating system to handle the core correctable error of the target processor. Since the SCI does not interrupt the normal execution of the operating system and applications, it can reduce the impact on the operating system, reduce system jitter or latency. At the same time, the BMC can accurately collect and record core correctable errors, realize fault monitoring, facilitate fault detection and maintenance of computing devices, and improve the overall performance, stability and availability of the server.
[0013] The above description is only an overview of the technical solution of this disclosure. In order to better understand the technical means of this disclosure, it can be implemented in accordance with the contents of the specification. In order to make the above and other objects, features and advantages of this disclosure more obvious and understandable, specific embodiments of this disclosure are given below. Attached Figure Description
[0014] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the various drawings denote the same or similar parts or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings depict only some embodiments according to this disclosure and should not be construed as limiting the scope of this disclosure.
[0015] Figure 1 shows a system architecture diagram of the server provided in an embodiment of this disclosure;
[0016] Figure 2 shows a flowchart of an error handling method according to an embodiment of the present disclosure;
[0017] Figure 3 shows a flowchart of an error handling method according to an embodiment of the present disclosure;
[0018] Figure 4 shows a flowchart of an error handling method according to an embodiment of the present disclosure;
[0019] Figure 5 shows an application example of the error handling method according to an embodiment of the present disclosure;
[0020] Figure 6 shows a block diagram of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0021] In the following description, only certain exemplary embodiments are briefly described. As those skilled in the art will recognize, the described embodiments can be modified in various ways without departing from the spirit or scope of this disclosure. Therefore, the drawings and description are considered to be exemplary in nature and not restrictive.
[0022] To facilitate understanding of the technical solutions of the embodiments of this disclosure, the related technologies of the embodiments of this disclosure are described below. The following related technologies are optional solutions and can be combined with the technical solutions of the embodiments of this disclosure in any way, and all of them fall within the protection scope of the embodiments of this disclosure.
[0023] Hardware errors may occur during the operation of a server system, such as memory errors (including memory errors and main memory errors) and bus errors. These peripheral hardware errors of the CPU can be recovered through peripheral hardware error handling mechanisms.
[0024] Correctable core errors occur within the CPU core, such as those occurring in the Arithmetic Logic Unit (ALU), Floating Point Unit (FPU), Cache, Instruction Fetch Unit (IFU), and Instruction Decode Unit (IDU). Currently, servers primarily employ two mechanisms for handling correctable CPU core errors.
[0025] One handling mechanism is based on System Management Interrupt (SMI): During the Basic Input / Output System (BIOS) initialization phase, the kernel correctable error handling mode is set to SMI. When a kernel correctable error occurs, SMI is triggered to the BIOS. The BIOS collects a large amount of fault logs and records them in the Baseboard Management Controller (BMC). Subsequently, the BIOS triggers an interrupt to notify the server's operating system (OS), and the OS records the fault logs in the operating system. Since SMI processing is executed by System Management Mode (SMM) code in the BIOS, it runs outside the server's operating system and applications. Therefore, when SMI is triggered, the server's operating system and applications, even if they can execute normally, will be suspended until SMI processing is complete. The SMI processing flow is relatively long, and the system jitter or service latency caused by a single SMI is typically on the order of hundreds of milliseconds. As the amount of data to be processed increases, system jitter or service latency will further increase, thereby affecting the overall performance of the server and the stability of the operating system.
[0026] Another handling mechanism is based on Corrected Machine Check Interrupt (CMCI): During the BIOS initialization phase, the handling mode for kernel correctable errors is set to CMCI. When a kernel correctable error occurs, the BIOS triggers CMCI, and the operating system directly handles the kernel correctable error and records the fault log in the operating system. In this mechanism, the normal execution of the operating system and applications is not interrupted, thus reducing the impact on the operating system and applications (the latency impact on the operating system is approximately on the order of hundreds of microseconds). However, since the handling of kernel correctable errors does not go through the BIOS and BMC, fault monitoring cannot be performed through the BMC, resulting in the inability to detect kernel correctable errors and affecting fault detection and corresponding maintenance of the server.
[0027] Based on this, the embodiments of this disclosure aim to provide a processing solution for core correctable errors, which can balance low system latency and high fault monitoring capabilities. The technical solution of this disclosure and how it solves the aforementioned technical problems are described in detail below with specific embodiments. The listed specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0028] Figure 1 shows a system architecture diagram of a server provided in an embodiment of this disclosure. As shown in Figure 1, the server includes a computing device 110 and a memory 120. The computing device 110 includes at least one processor 111, a bridge chip 112, and a baseboard management controller (BMC) 113.
[0029] Memory 120 may include volatile memory or non-volatile memory, or both. Processor 111 serves as the core processor of computing device 110, such as a CPU. This disclosure does not specifically limit the number of processors 111, but may include, for example, CPU0 and CPU1 as shown in FIG. 1. One or more processors 111 are used to run the operating system of computing device 110. Exemplarily, when computing device 110 is started, processor 111 reads a boot program from memory 120, thereby loading the operating system.
[0030] Bridge chip 112 is a chip integrated on the motherboard of computing device 110. It is mainly responsible for managing and coordinating various peripheral devices and input / output (I / O) interfaces. Bridge chip 112 may be, for example, a southbridge chip or a platform controller hub (PCH). The basic input / output system (BIOS) is integrated on bridge chip 112.
[0031] The BMC113 can be deployed with its own processor, memory, and firmware, and can run independently of the operating system of the computing device 110. The BMC113 communicates with other components of the computing device 110 (such as sensors and power management units) through a dedicated interface, and interacts with remote management tools through a network interface (such as Ethernet), providing remote access to sensor data, logs, console output, and other functions of the computing device 110.
[0032] BMC113 is connected to bridge chip 112 via the programmable input / output (GPIO) pins of bridge chip 112. In this embodiment, an unused pin can be selected from the GPIO pins as the target pin. During the BIOS initialization phase, the enable and reset signals of the target pin can be configured and defined. For example, the set signal can be enabled when the target pin is high, and the reset signal can be enabled when the target pin is low. The set signal of the target pin indicates the presence of a target processor, i.e., a processor that has experienced a core correctable error; the reset signal of the target pin indicates that no core correctable error has occurred.
[0033] The BIOS integrated on the bridge chip 112 is configured during the initialization phase to trigger a set signal on a target pin in response to the presence of a target processor (e.g., CPU0) experiencing a core correctable error in at least one processor 111. Exemplarily, processor 111 internally includes a status register for detecting and reporting hardware faults occurring throughout the server 100 system, including memory errors (e.g., memory correctable errors), cache errors, core errors (e.g., core correctable errors), bus errors, etc. The BIOS identifies whether a core correctable error has occurred by reading the status register. If a core correctable error is detected, it triggers a set signal on the target pin, for example, by pulling the target pin high (pulling the target pin high), thereby indicating the presence of a processor experiencing a core correctable error.
[0034] BMC113 is configured to: in response to a set signal on a target pin, determine a target processor from at least one processor 111 and record information about a core correctable error occurring in the target processor. In one implementation, BMC113 can obtain information about the target processor's core correctable error through the BIOS, thereby obtaining information about the target processor and the core correctable error it has encountered. In another implementation, BMC113 can read the core error register of each processor 111 to determine the target processor and obtain information about the target processor's core correctable error, thereby reducing the BIOS resource consumption of the error handling process. The core error register is an internal configuration of the processor 111 used to record correctable errors occurring within its core. Each processor 111 has an internal core error register for recording correctable errors occurring in that processor's core. Therefore, BMC113 can read the core error register of each processor 111, thereby identifying the processor with the recorded core correctable error as the target processor and obtaining information about the target processor's core correctable error.
[0035] For example, the BMC113 can record information about a core correctable error occurring in the target processor in the memory of the BMC113 (e.g., non-volatile memory). The recorded information includes, but is not limited to, timestamps, error types (e.g., codes that can identify core correctable errors), error locations (e.g., CPU0, CPU1, etc.), and error counts (e.g., the number of times a core correctable error has occurred).
[0036] For example, the BMC113 can record information about a core correctable error occurring in the target processor based on the System Event Log. For instance, the information about a core correctable error can be recorded as a new SEL entry in the SEL. Users can view the SEL through the interface provided by the BMC113 to obtain information about the core correctable error in the target processor. Alternatively, the BMC113 can send the SEL to the user via the network to alert or notify the user that a core correctable error has occurred in the target processor.
[0037] Furthermore, BMC113 is also configured to send a System Control Interrupt (SCI) to the operating system of computing device 110 in response to a set signal on the target pin. The SCI is used to transmit system control events to the operating system, which then processes these events; therefore, the SCI does not suspend the operation of the operating system. In this embodiment, the SCI is used to notify the operating system to handle a core correctable error in the target processor. Exemplarily, the operating system reads the core error register of each processor 111 to obtain information about a core correctable error occurring in the target processor and processes the core correctable error.
[0038] In related technologies, the CMCI interrupt method is used, and a masking parameter is added to the OS kernel boot parameters to disable CMCI responses. The BIOS is configured to handle Core Correctable Errors (Core CEs) using CMCI, and the OS disables CMCI to prevent the Core CE count in the registers from being cleared by the OS. Combined with the BMC polling the Core register values via PECI, fault information is recorded in the BMC log, ultimately achieving out-of-band monitoring of CPU CE faults without affecting the OS. While this method avoids the latency of SMI interrupts, practical verification shows that the masking parameter cannot permanently disable CMCI interrupts. After a period of time, the response to CMCI resumes, and the register information is cleared again, preventing the BMC from reading fault information and hindering out-of-band monitoring. Furthermore, the masking parameter also masks Memory Correctable Errors (CEs), resulting in no OS record when memory CEs occur, affecting memory CE monitoring.
[0039] According to the computing device provided in this disclosure embodiment, after the BIOS detects a target processor (i.e., a processor that has experienced a core correctable error), it triggers a set signal on a target pin to indicate the presence of a processor experiencing a core correctable error. Upon detecting the set signal on the target pin, the BMC records the information about the core correctable error in the target processor and triggers the SCI (System Change Component) to enable the operating system to handle the core correctable error. Since the SCI does not interrupt the normal execution of the operating system and applications, it reduces the impact on the operating system, lowers system jitter or latency. Simultaneously, the BMC can accurately collect and record core correctable errors, enabling fault monitoring and facilitating fault detection and maintenance of the computing device.
[0040] The server including the above-described computing device provided according to the embodiments of this disclosure can reduce the impact of error handling processes on the server operating system and applications, reduce system jitter or service latency, and at the same time, both the BMC and the operating system can accurately collect and record the core correctable errors of the server, realize fault monitoring and corresponding maintenance, thereby improving the overall performance, stability and availability of the server.
[0041] As described above, processor 111 includes a core error register for recording correctable core errors in the event of such errors. In one embodiment, BMC 113 has at least one Platform Environment Control Interface (PECI), with each PECI corresponding to at least one processor 111. Each PECI typically consists of two signal lines: a clock line (CLK) and a data line (DATA), which are directly connected between processor 111 and BMC 113, providing a fast and reliable communication channel between them.
[0042] Based on PECI, the BMC113 can send PECI commands to each processor 111 (such as CPU0 and CPU1). PECI commands are used to read the corresponding processor's core error register. For example, the BMC113 sends a PECI command to CPU0 via its PECI with CPU0 to read CPU0's core error register; similarly, the BMC113 sends a PECI command to CPU1 via its PECI with CPU1 to read CPU1's core error register. By reading the core error register of each processor 111 (such as CPU0 and CPU1), a count value for each core error register can be obtained. For example, a count value of "0" indicates that no core correctable error has occurred in that core; a count value other than "0" indicates that a core correctable error has occurred in that core. Based on these count values, it is possible to determine which processor(s) are the target processors for the core correctable error. For example, if the core error register value of CPU0 is read as non-"0", then CPU0 can be identified as the target processor, indicating that a core correctable error exists on CPU0.
[0043] Because the PECI channel can provide a low-latency and high-reliability communication mechanism, using the PECI channel between the BMC and the processor to read the processor's core error register can not only achieve efficient and accurate fault monitoring and management, but also enhance scalability and compatibility, making it particularly suitable for data center and server environments that require high reliability and high performance.
[0044] In one implementation, the operating system is further configured to: in response to SCI, read the kernel error register of the target processor to record information about a kernel correctable error occurring in the target processor, and clear the kernel correctable error recorded in the kernel error register of the target processor after recording the information.
[0045] For example, the operating system can respond to the SCI by initiating a corresponding error handler. The error handler reads the kernel error register via register instructions to obtain detailed error information and records this information in the system log, such as the timestamp of the kernel correctable error, error type, error location, and error count. After recording this error information, the operating system clears the kernel correctable error recorded in the target processor's kernel error register by writing a specific value via register instructions, preventing the same error from being processed repeatedly by the operating system or BMC.
[0046] By utilizing the SCI interrupt mechanism, the operating system can efficiently respond to the BMC's SCI, read the target processor's kernel error register to record information about correctable kernel errors, and then clear this error information. This approach allows the operating system to promptly detect and collect correctable kernel errors, while also preventing duplicate reporting and processing of errors, thus enhancing the system's automated management and reliability.
[0047] In one implementation, the BIOS is specifically configured to: perform an SMI in response to the presence of a target processor in at least one processor 111, and exit the SMI after a set signal on the target pin is triggered.
[0048] In this embodiment of the disclosure, SMI is used to trigger the set signal of the target pin and interrupt the operating system. As mentioned above, SMI is a high-priority interrupt mechanism; when the BIOS enters SMI, it will interrupt the execution of the operating system and applications. In the technical solution of this embodiment of the disclosure, after the BIOS detects a correctable kernel error, it will automatically enter SMI, and after the set signal of the target pin is triggered, the BIOS will automatically exit SMI.
[0049] Therefore, this implementation provides a fast-in / fast-out SMI approach. On the one hand, it reduces the impact on the operating system and applications by decreasing the processing time of SMI. On the other hand, the protected mode of SMI allows the BIOS to perform necessary error handling, such as initial recovery measures, to prevent malicious setting of target pins, thereby performing error detection in a safer and more controlled environment. Furthermore, SMI is often used as a standard response mechanism in server or computing device specifications; therefore, BIOS fast-in / fast-out SMI also improves the compatibility of computing devices and servers.
[0050] In one implementation, the BMC113 is also configured to trigger a reset signal on a target pin after recording information about a core correctable error occurring in the target processor. Exemplarily, the BMC can send a reset command to the target pin via a hardware or firmware interface thereon, thereby triggering a reset signal on the target pin, for example, by pulling the target pin low.
[0051] Based on this, the target pin can be reset in a timely manner, avoiding missing any core correctable errors that occurred before the reset.
[0052] It should be noted that other components of the computing device and server in the embodiments of this disclosure can adopt various technical solutions that are now and will be known to those skilled in the art, and will not be described in detail here.
[0053] Figure 2 shows a flowchart of an error handling method according to an embodiment of the present disclosure. This method can be applied to the BMC of a computing device, for example, executed by the BMC113 shown in Figure 1. As shown in Figure 2, the method may include:
[0054] Step S201: In response to the set signal of the target pin, determine the target processor from at least one processor where the core correctable error has occurred, wherein the set signal of the target pin is used to indicate the presence of the target processor;
[0055] Step S202: Record information about a core correctable error occurring in the target processor;
[0056] Step S203: Send a system control interrupt to the operating system. The system control interrupt is used to notify the operating system to handle the core correctable error of the target processor.
[0057] According to the error handling method provided in Figure 2 of this embodiment, after detecting the set signal of the target pin, the BMC records the information of the core correctable error occurring in the target processor, and triggers the SCI to enable the operating system to handle the core correctable error of the target processor. Since the SCI does not interrupt the normal execution of the operating system and applications, it can reduce the impact on the operating system, reduce system jitter or latency. At the same time, the BMC can accurately collect and record core correctable errors, realize fault monitoring, facilitate fault detection and maintenance of computing devices and servers, thereby improving the overall performance, stability and availability of the server.
[0058] In one implementation, in step S201, determining the target processor from at least one processor includes: sending a PECI command to each processor, the PECI command being used to read the corresponding processor's core error register, the core error register being used to record a core correctable error in the event that a core correctable error occurs in its corresponding processor; and determining the processor corresponding to the core error register that records the core correctable error as the target processor.
[0059] In other words, the PECI channel provides a communication channel between the processor and the BMC. The BMC can send PECI commands to each processor, thereby reading the core error register of each processor and obtaining the count value of each core error register. The count value can indicate whether a core-correctable error has occurred in that processor core. Based on these count values, the BMC can determine the processor that has experienced a core-correctable error, i.e., the target processor.
[0060] Because the PECI channel can provide a low-latency and high-reliability communication mechanism, using PECI commands to read the processor's core error register can not only achieve efficient and accurate fault monitoring and management, but also enhance scalability and compatibility, making it particularly suitable for data center and server environments that require high reliability and high performance.
[0061] In one embodiment, the method of this disclosure may further include: after step S202, triggering a reset signal for the target pin. For example, sending a reset command for the target pin pulls the target pin low to trigger the reset signal for the target pin.
[0062] Based on this, the target pin can be reset in a timely manner, avoiding missing any core correctable errors that occur before the target pin is reset.
[0063] Figure 3 illustrates a flowchart of an error handling method according to an embodiment of the present disclosure. This method can be applied to the BIOS of a computing device, for example, executed by the BIOS shown in Figure 1. As shown in Figure 3, the method may include:
[0064] Step S301: In response to the presence of a target processor in at least one processor that has a core correctable error, a set signal of the target pin is triggered, wherein the set signal of the target pin is used to indicate the presence of a target processor, so that the board management controller records the information of the core correctable error of the target processor and sends a system control interrupt to the operating system, wherein the system control interrupt is used to notify the operating system to handle the core correctable error of the target processor.
[0065] According to the error handling method provided in Figure 3 of this embodiment, after the BIOS detects the target processor (i.e. the processor that has a core correctable error), it will trigger the set signal of the target pin to indicate that there is a processor that has a core correctable error. The subsequent error handling process will be transferred to the BMC so that the BMC can collect and record the core correctable error and realize fault monitoring.
[0066] Figure 4 shows a flowchart of an error handling method according to an embodiment of the present disclosure. This method can be applied to the BIOS of a computing device, for example, executed by the BIOS shown in Figure 1. As shown in Figure 4, the method may include:
[0067] Step S401: In response to the presence of a target processor in at least one processor where a core correctable error has occurred, execute a system management interrupt, which is used to trigger the set signal of the target pin and interrupt the operating system of the computing device;
[0068] Step S402: Exit the system management interrupt after the set signal of the target pin is triggered.
[0069] According to the error handling method provided in Figure 4 of this embodiment, after the BIOS detects a correctable core error, it enters SMI (System Error Management) and exits SMI after the target pin's set signal is triggered. This fast-in / fast-out approach reduces the impact on the operating system and applications by decreasing the SMI processing time. Furthermore, the SMI protection mode allows the BIOS to perform necessary error handling, such as initial recovery measures, to prevent malicious setting of the target pin, thus enabling error detection in a safer and more controllable environment. Additionally, SMI is typically used as a standard response mechanism in server or computing device specifications; therefore, the BIOS's fast-in / fast-out SMI also improves the compatibility of computing devices and servers.
[0070] It should be noted that the error handling methods shown in Figures 2, 3 and 4, their corresponding implementation methods and technical effects can be found in the corresponding descriptions of server 100 and computing device 110 in Figure 1, and will not be repeated here.
[0071] Figure 5 shows an application example of the error handling method according to an embodiment of the present disclosure. As shown in Figure 5, in this example, the error handling method includes: (1) a BIOS SMI simplification process, for example executed by the BIOS shown in Figure 1, or implemented by the BIOS executing steps S401 and S402; (2) a BMC process, for example executed by the BMC 113 shown in Figure 1, or implemented by the BMC executing steps S201 to S203; (3) an OS process, for example executed by the operating system of the computing device 110 shown in Figure 1.
[0072] In the simplified BIOS SMI process, the BIOS detects that the server is in normal operation. If a core correctable error occurs in a processor, such as CPU0 as shown in Figure 1, an SMI is triggered to the BIOS, causing the BIOS to enter the SMI interrupt procedure. Further, the BIOS pulls a preset GPIO high (set) and exits the SMI. The preset GPIO is the target pin; pulling the preset GPIO high pulls the target pin to a high level, thereby triggering the set signal of the target pin. After the set signal of the target pin is triggered, the BIOS exits the SMI. In other words, in the simplified BIOS SMI process, the BIOS triggers the set signal of the target pin and performs a fast-forward and fast-out SMI.
[0073] In the BMC process, the BMC detects a preset GPIO set, i.e., detects the set signal of the target pin. Responding to the target pin's set signal, the BMC polls the CPU CE register value via the PECI command, i.e., polls the core error register count value of each processor (e.g., CPU0 and CPU1 as shown in Figure 1). Further, the BMC checks if CPU CE Count != 0, i.e., determines whether the count values of these core error registers contain non-zero values. If so, it indicates that a processor has experienced a core correctable error. Thus, the BMC finds that CPU0's core error register has a non-zero value, thereby determining that CPU0 is the processor experiencing a core correctable error. Further, the BMC acquires information from the CPU0 CE register (CPU0's core error register) and records it to SEL, realizing fault acquisition on the BMC side. Further still, the BMC pulls a specific GPIO (reset) low to trigger the reset signal of the target pin; in parallel, the BMC triggers SCI to notify the OS.
[0074] In the OS process, the OS responds to SCI, collects information from the CPU0 CE register (CPU0's core error register), thereby forming a fault log on the OS side, and clears the error information recorded in the CPU0 CE register (CPU0's core error register).
[0075] It should be noted that the application scenarios or examples provided in this disclosure are for ease of understanding, and this disclosure does not specifically limit the application of the technical solutions. Furthermore, all information and data involved in this disclosure are authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0076] Figure 6 is a block diagram of an electronic device used to implement embodiments of the present disclosure. As shown in Figure 6, the electronic device includes a memory 601 and a processor 602. The memory 601 stores a computer program that can run on the processor 602. When the processor 602 executes the computer program, it implements the methods in the above embodiments. The number of memories 601 and processors 602 can be one or more. In a specific implementation, the electronic device may also include a communication interface 603 for communicating with external devices and performing data exchange and transmission.
[0077] In practical implementation, if the memory 601, processor 602, and communication interface 603 are implemented independently, they can be interconnected via a bus to communicate with each other. This bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be categorized as an address bus, data bus, control bus, etc. For ease of representation, only one thick line is used in Figure 6, but this does not indicate that there is only one bus or one type of bus.
[0078] Optionally, in a specific implementation, if the memory 601, processor 602 and communication interface 603 are integrated on a single chip, the memory 601, processor 602 and communication interface 603 can communicate with each other through an internal interface.
[0079] This disclosure provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the methods provided in this disclosure.
[0080] This disclosure provides a computer program product, including a computer program that, when executed by a processor, implements the methods provided in this disclosure.
[0081] This disclosure also provides a chip including a processor for calling and executing instructions stored in a memory, causing a communication device on which the chip is installed to perform the methods provided in this disclosure.
[0082] This disclosure also provides a chip, including: an input interface, an output interface, a processor, and a memory. The input interface, output interface, processor, and memory are connected through an internal connection path. The processor is used to execute code in the memory. When the code is executed, the processor is used to execute the method provided in the application embodiment.
[0083] It should be understood that the aforementioned processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. It is worth noting that the processor can be a processor supporting Advanced Reduced Instruction Set Machines (ARM) architecture.
[0084] Further, optionally, the aforementioned memory may include read-only memory and random access memory. The memory may be volatile memory or non-volatile memory, or may include both. Non-volatile memory may include read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may include random access memory (RAM), which serves as an external cache. By way of example, but not limitation, many forms of RAM are available. Examples include Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), and Direct Rambus RAM (DR RAM).
[0085] In the above embodiments, implementation can be achieved, in whole or in part, by software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions according to this disclosure are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another.
[0086] In the description of this disclosure, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this disclosure. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this disclosure, as well as the features of those different embodiments or examples.
[0087] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this disclosure, "a plurality of" means two or more, unless otherwise explicitly specified.
[0088] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process. Furthermore, the scope of the preferred embodiments of this disclosure includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functionality involved.
[0089] The logic and / or steps described in the flowchart or otherwise herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus or device (such as a computer-based system, a processor-included system or other system that can fetch and execute instructions from, an instruction execution system, apparatus or device).
[0090] It should be understood that various parts of this disclosure can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. All or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware, the program being stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiments.
[0091] Furthermore, the functional units in the various embodiments of this disclosure can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. This storage medium can be a read-only memory, a disk, or an optical disk, etc.
[0092] The above description is merely an exemplary embodiment of this disclosure, but the scope of protection of this disclosure is not limited thereto. Any person skilled in the art can easily conceive of various variations or substitutions within the technical scope described in this disclosure, and these should all be included within the scope of protection of this disclosure. Therefore, the scope of protection of this disclosure should be determined by the scope of the claims.
Claims
1. A computing device, comprising: At least one processor for running the operating system of the computing device; A bridge chip, connected to the processor, the bridge chip integrating a basic input / output system, the basic input / output system being configured to: in response to the presence of a target processor in the at least one processor where a core correctable error has occurred, trigger a set signal of a target pin, the set signal of the target pin being used to indicate the presence of the target processor; A baseboard management controller is connected to the bridge chip via the target pin. The baseboard management controller is configured to: determine the target processor from the at least one processor in response to a set signal of the target pin, record information about a core correctable error occurring in the target processor, and send a system control interrupt to the operating system. The operating system is configured to handle core correctable errors of the target processor in response to system control interrupts.
2. The computing device according to claim 1, wherein, The processor is provided with a core error register for recording a core correctable error when the corresponding processor experiences a core correctable error; the baseboard management controller is further configured to: for any processor, identify whether the processor has experienced a core correctable error by reading the processor's core error register, and determine the processor as the target processor if the processor experiences a core correctable error.
3. The computing device according to claim 2, wherein, The baseboard management controller is also configured to: obtain information about a core correctable error occurring in the target processor by reading the core error register of the target processor.
4. The computing device according to claim 2, wherein, The baseboard management controller has at least one platform environment control interface, and the at least one platform environment control interface is connected to the at least one processor in a one-to-one correspondence. The baseboard management controller is further configured to: send platform environment control interface commands to each of the processors through each of the platform environment control interfaces to read the core error register of each processor, and determine the processor corresponding to the core error register that records the core correctable error as the target processor.
5. The computing device according to claim 2, wherein, The operating system is also configured to: in response to the system control interrupt, read the kernel error register of the target processor to record information about a kernel correctable error occurring in the target processor, and clear the kernel correctable error recorded in the kernel error register of the target processor.
6. The computing device according to claim 5, wherein, The operating system is also configured to: in response to a system control interrupt, initiate a corresponding error handler, the error handler reading the kernel error register of the target processor via register instructions and recording information about a kernel correctable error in the operating system's system log.
7. The computing device according to claim 1, wherein, The baseboard management controller is also configured to record information about core correctable errors occurring in the target processor based on the system event log, and to provide a user interface for viewing the system event log.
8. The computing device according to claim 1, wherein, The baseboard management controller is also configured to: record information about core correctable errors occurring in the target processor based on the system event log, and send the system event log over the network.
9. The computing device according to claim 1, wherein, The baseboard management controller is also configured to record information about a core correctable error occurring in the target processor in the memory of the baseboard management controller.
10. The computing device according to any one of claims 1 to 9, wherein, The basic input / output system is specifically configured to: in response to the presence of the target processor in the at least one processor, execute a system management interrupt, and exit the system management interrupt after the set signal of the target pin is triggered, wherein the system management interrupt is used to trigger the set signal of the target pin and interrupt the operating system.
11. The computing device according to any one of claims 1 to 9, wherein, The baseboard management controller is also configured to trigger a reset signal on the target pin after recording information that a core correctable error has occurred in the target processor.
12. An error handling method applied to a baseboard management controller of a computing device, the computing device further comprising at least one processor, the at least one processor being configured to run an operating system of the computing device, the method comprising: In response to a set signal on a target pin, a target processor from the at least one processor is determined to have a core correctable error, wherein the set signal on the target pin is used to indicate the presence of the target processor; Record information about core correctable errors occurring in the target processor; A system control interrupt is sent to the operating system, which is used to notify the operating system to handle the core correctable error of the target processor.
13. The method according to claim 12, wherein, The step of determining the target processor from the at least one processor includes: A platform environment control interface command is sent to each of the processors. The platform environment control interface command is used to read the core error register of the corresponding processor. The core error register is used to record a core correctable error when a core correctable error occurs in the corresponding processor. The processor corresponding to the kernel error register that records core correctable errors is identified as the target processor.
14. The method according to claim 12 or 13, wherein, After recording the information about the core correctable error that occurred in the target processor, the method further includes: The reset signal of the target pin is triggered.
15. An error handling method applied to the basic input / output system of a computing device, the computing device further comprising a baseboard management controller and at least one processor, the at least one processor being used to run an operating system of the computing device, the method comprising: In response to the presence of a target processor in the at least one processor that has a core correctable error, a set signal of a target pin is triggered, wherein the set signal of the target pin is used to indicate the presence of the target processor, so that the board management controller records the information of the core correctable error occurring in the target processor and sends a system control interrupt to the operating system, wherein the system control interrupt is used to notify the operating system to handle the core correctable error of the target processor.
16. The method according to claim 15, wherein, The triggering of the target pin's set signal includes: executing a system management interrupt, wherein the system management interrupt is used to trigger the target pin's set signal and interrupt the operating system of the computing device; The method further includes: exiting the system management interrupt after the set signal of the target pin is triggered.
17. A server comprising a computing device according to any one of claims 1 to 11 and a memory.
18. An electronic device comprising a memory, a processor, and a computer program stored in the memory, wherein the processor, when executing the computer program, implements the method of any one of claims 12 to 16.
19. A computer-readable storage medium storing a computer program that, when executed by a processor, implements the method of any one of claims 12 to 16.
20. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 12 to 16.