Bus device uncorrectable error handling method and server
By reporting the first UCE error triggered by a PCIe device and setting a time window and suppression flag, the problem of frequent CPU entry into SMM mode caused by frequent UCE errors is solved, thus improving the reliability of the server system and the stability of business operations.
Patent Information
- Application Number
- CN202511240052.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-01
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2045-09-01
AI Technical Summary
PCIe devices may trigger UCE errors due to hardware failures or unstable links, causing the CPU to frequently enter SMM mode, creating an error reporting storm, consuming system resources, masking critical business needs, and affecting server system availability.
When a UCE error is first triggered by a bus device, it is reported, and a time window for uncorrectable error statistics is opened. A suppression flag is assigned, and the reporting is suppressed when the error count is less than the threshold within the time window. The reporting is resumed when the error count is greater than or equal to the threshold. By setting the suppression flag and the time window, the error reporting frequency is reduced to ensure that no initial error is missed.
This reduces the frequency of UCE error reporting, minimizes redundant reporting, avoids error reporting storms, improves the reliability of the server system and the impact on normal business operations, and ensures that no critical alarms are missed.
Smart Images

Figure CN120723527B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of servers, and particularly relates to a bus device uncorrectable error processing method and a server. BACKGROUND
[0002] In a server system, a PCIe (Peripheral Component Interconnect Express, high-speed serial computer expansion bus standard) device may trigger an UCE (Uncorrectable Error, uncorrectable error) due to hardware failure or unstable link.
[0003] In the related art, PCIe error processing mainly relies on a hardware register to record error states, and triggers an SMI (System Management Interrupt, system management interrupt) through a BIOS firmware priority reporting mechanism, and the BIOS (Basic Input / Output System, basic input / output system) processes and reports to the BMC (Baseboard Management Controller, baseboard management controller) and notifies the OS (Operating System, operating system). However, when the device frequently triggers the UCE error, the system will continuously trigger the SMI interrupt, causing the CPU (Central Processing Unit, central processing unit) to repeatedly enter the SMM (System Management Mode, system management mode). This not only consumes a large amount of system resources, but also may mask other critical business requirements, form an error reporting storm, and seriously affect the availability of the system. SUMMARY
[0004] The present application provides a bus device uncorrectable error processing method and a server to at least solve the problem that the uncorrectable error is frequently triggered in the related art, masks other business requirements, and affects the availability of the server system.
[0005] The present application provides a bus device uncorrectable error processing method, comprising the following steps: detecting whether a bus device triggers an uncorrectable error for the first time; if it is detected that the bus device triggers the uncorrectable error for the first time, triggering an uncorrectable error reporting event and starting a time window for uncorrectable error statistics; assigning a suppression flag to the bus device within the time window, obtaining an error count of the uncorrectable error within the time window, and the suppression flag is used to suppress the reporting of the uncorrectable error when the error count is less than a reporting threshold within the time window; if the error count is greater than or equal to the reporting threshold, triggering the uncorrectable error reporting event, if the time window is timed out and the error count is less than the reporting threshold, clearing the suppression flag of the bus device and resetting the error count of the uncorrectable error.
[0006] The application further provides a bus device uncorrectable error processing apparatus, comprising: a detection module, configured to detect whether a bus device triggers an uncorrectable error for the first time; a triggering module, configured to trigger an uncorrectable error reporting event and start a time window for uncorrectable error statistics if it is detected that the bus device triggers the uncorrectable error for the first time; an assignment module, configured to assign a suppression flag to the bus device within the time window and acquire error count of uncorrectable errors within the time window, the suppression flag being used to suppress the uncorrectable error reporting when the error count is less than a reporting threshold within the time window; and a processing module, configured to trigger the uncorrectable error reporting event if the error count is greater than or equal to the reporting threshold, and clear the suppression flag of the bus device and reset the error count of uncorrectable errors if the time window is overdue and the error count is less than the reporting threshold.
[0007] The application further provides a server, comprising: a memory, configured to store a computer program; and a processor, configured to execute the computer program to implement the steps of the bus device uncorrectable error processing method according to the above embodiments.
[0008] According to the application, since the bus device may trigger uncorrectable errors due to faults or unstable links, when the bus device frequently triggers uncorrectable errors, system management interrupts are continuously triggered, the central processing unit repeatedly enters the system management mode, an error reporting storm is formed, a large amount of server system resources are consumed, other key business requirements may be covered, and the availability of the server system is seriously affected,
[0009] Therefore, this embodiment reports an error when the bus device first triggers an uncorrectable error and opens a time window for uncorrectable error statistics. Within this time window, a suppression flag is assigned to the bus device. If the uncorrectable error count reaches the reporting threshold within the time window, no error reporting occurs. If the error count is greater than or equal to the reporting threshold, reporting occurs again. By setting the uncorrectable error suppression flag for the bus device and the time window for uncorrectable error statistics, the frequency of uncorrectable error reporting is reduced. Furthermore, reporting the uncorrectable error upon its first trigger captures it, ensuring no initial error is missed. Reporting again when the error count within the time window exceeds the reporting threshold avoids missing errors that exacerbate the problem, reducing redundant error reporting and suppressing error reporting storms. This reduces the impact of frequent error triggering on the availability of the server system and the normal operation of other services, improving server reliability. Thus, this solves the technical problem in related technologies where frequent uncorrectable errors mask other business needs and affect the availability of the server system, achieving the technical effect of reducing the impact of frequent error triggering on the availability of the server system and the normal operation of other services, and improving server reliability. Attached Figure Description
[0010] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 This is a schematic diagram of a bus device uncorrectable error handling method provided according to an embodiment of this application;
[0012] Figure 2 This is an architecture diagram of a bus device uncorrectable error handling system provided according to an embodiment of this application;
[0013] Figure 3 This is a schematic diagram of a bus device uncorrectable error handling apparatus provided according to an embodiment of this application;
[0014] Figure 4 This is a schematic diagram of the structure of a server provided according to an embodiment of this application. Detailed Implementation
[0015] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.
[0016] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0017] Before describing the solution of this application, let me first introduce some related content to aid in understanding the solution of this application, as follows.
[0018] There are many mechanisms in related technologies to deal with UCE error reporting storms, such as UCE threshold mechanisms and UCE error storm suppression. However, they all overlook the fact that, even without causing system crashes or restarts, frequent UCE error reporting storms can cause the system to frequently enter and exit SMM mode, thereby affecting the operation of critical system services. This is especially true for systems that have deployed fault monitoring, early warning, and processing services. They could have pushed a fault alarm for a certain PCIe device to the user, but frequent UCE storms may cause the alarm to fail to be sent or be delayed, leading to more serious system failures.
[0019] Furthermore, the typical approach to handling a Fatal UCE is to crash or restart the system and log the relevant information. However, if a system restarts or crashes and is then powered off and restarted, the faulty PCIe device may still exist, causing the system to crash or restart again. Critical logs recorded in the OS system may not be exported, and services cannot be restored in a timely manner. Therefore, there is a lack of measures to automatically suppress (mask) UCE errors after a crash or restart caused by a UCE.
[0020] In summary, the solutions provided by these technologies have the following shortcomings:
[0021] 1. Error reporting storm issue: When the device continuously triggers UCE errors, the high frequency of SMI interrupts will significantly reduce system performance and may cause other business anomalies.
[0022] 2. Static configuration limitations: Mechanisms in related technologies typically configure error handling strategies through fixed thresholds or hardware registers, which cannot adapt to dynamically changing error scenarios.
[0023] 3. Lack of suppression mechanism: No UCE error reporting suppression strategy is provided for specific devices, resulting in redundant error information overwhelming critical alarms.
[0024] Therefore, embodiments of this application provide a server uncorrectable error handling system to solve at least one of the above problems.
[0025] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0026] The embodiments of this application provide a method for handling uncorrectable errors in bus devices. The method is described in detail below in conjunction with the execution flow of the method for handling uncorrectable errors in bus devices.
[0027] Figure 1 This is a flowchart of a method for handling uncorrectable errors in a bus device according to an embodiment of this application.
[0028] like Figure 1 As shown, the uncorrectable error handling method for this bus device includes the following steps:
[0029] In step S101, it is detected whether the bus device has triggered an uncorrectable error for the first time.
[0030] Among them, the bus device can be a PCIe device.
[0031] In some embodiments of this application, detecting whether a bus device has triggered an uncorrectable error for the first time includes: detecting whether the bus device has triggered an uncorrectable error; when the bus device triggers an uncorrectable error, identifying whether the bus device carries a suppression flag; if the bus device carries a suppression flag, determining that the bus device has not triggered an uncorrectable error for the first time, and if it does not carry a suppression flag, determining that the bus device has triggered an uncorrectable error for the first time.
[0032] The suppression flag can be represented by UCE_SUPPRESSION_FLAG. The suppression flag can be used to determine whether the bus device is triggering an uncorrectable error for the first time, or it can be used to suppress the reporting of uncorrectable errors.
[0033] It is understood that the embodiments of this application can determine whether it is a first-time trigger of an uncorrectable error by detecting the suppression flag of the bus device. When the bus device carries a suppression flag, it is determined that the bus device is not a first-time trigger of an uncorrectable error. When the bus device does not carry a suppression flag, it is determined that the bus device is a first-time trigger of an uncorrectable error. By accurately determining whether it is a first-time trigger of an error, more accurate processing of uncorrectable errors can be performed subsequently.
[0034] In step S102, if an uncorrectable error is detected for the first time by a bus device, an uncorrectable error reporting event is triggered, and a time window for uncorrectable error statistics is opened.
[0035] The time window can be represented by UCE_ERROR_WINDOW_TIMER. The length of the time window can be set according to the specific situation and there is no specific limitation. For example, it can be set to 5 minutes, 6 minutes, etc. The reported event is an uncorrectable error reported to at least one of the baseboard management controller and the operating system of the server.
[0036] It is understood that, in this embodiment of the application, when an uncorrectable error is detected for the first time by a bus device, an uncorrectable error reporting event is triggered to promptly report the uncorrectable error, capture the uncorrectable error, ensure that no initial error is missed, and open a time window for uncorrectable error statistics so as to facilitate further processing of the uncorrectable error triggered by the bus device in the future.
[0037] In some embodiments of this application, if an uncorrectable error is detected for the first time by a bus device, an uncorrectable error reporting event is triggered, including: if an uncorrectable error is detected for the first time by a bus device, calling the interrupt service function of the bus device and using the interrupt service function to trigger a device interrupt of the bus device; and reporting the uncorrectable error after triggering the device interrupt of the bus device.
[0038] The interrupt service function of the bus device can trigger a bus device interrupt, which can be understood as an SMI interrupt. After the SMI interrupt is triggered, the processor will enter SMM mode.
[0039] It is understood that, in the embodiments of this application, when an uncorrectable error is detected for the first time by a bus device, the interrupt service function of the bus device is called, the interrupt service function is used to trigger a device interrupt of the bus device, and after the interrupt, the uncorrectable error is reported, thereby realizing timely and effective error handling, recording and diagnosis of the bus device, so as to investigate and repair the bus device that triggered the error in the future.
[0040] It should be noted that the SMI interrupt is a special type of hardware interrupt in the server, controlled by the firmware, used to handle system-level emergencies. Its core characteristic is that it has the highest priority. The priority of SMI is higher than that of the CPU's maskable interrupts, non-maskable interrupts, and even hardware reset signals. Once triggered, the CPU will immediately suspend all current tasks (including the operating system kernel and applications) and enter a special processing mode. SMI is non-maskable (similar to NMI, but SMI is more special), has the highest priority in the x86 architecture, and enters SMM after being triggered.
[0041] The SMI processing flow is as follows: After a UCE error is triggered, an interrupt signal is sent to the CPU via the SMI pin. The CPU immediately suspends the current instruction stream; saves key information such as register states and program counters to SMRAM (System Management RAM); switches to System Management Mode (SMM mode), executes the preset firmware processing program in SMRAM (usually provided by BIOS); after processing is complete, the CPU restores the previously saved state from SMRAM, executes the RSM (Return from SMM) instruction, exits SMM mode, and continues to execute the interrupted task.
[0042] SMM mode is a high-privilege, isolated operating mode designed specifically for firmware to handle system-level core tasks. It is the execution environment triggered by the SMI interrupt. Its core characteristic is complete independence from the operating system and other operating modes. SMM has its own dedicated physical memory area, SMRAM, which is hardware-protected and cannot be directly accessed by the operating system, applications, or even other CPU modes (such as real mode and protected mode), ensuring that firmware code and data are not tampered with. SMM has higher privileges than all other CPU modes (including kernel mode), and can directly manipulate hardware registers, modify system configurations (such as memory mapping and interrupt vector tables), and even suspend the entire system. SMM is usually implemented by BIOS / UEFI firmware.
[0043] In some embodiments of this application, after triggering the reporting event of an uncorrectable error, the method further includes: obtaining the number of times the bus device reports the event during its current lifecycle, wherein the current lifecycle is the period from power-on startup to power-off shutdown of the bus device; if the number of triggers is greater than a disable threshold, then disabling the reporting function of the bus device for uncorrectable errors, wherein the reporting function is used to trigger the reporting event of an uncorrectable error.
[0044] The lifecycle refers to the period from power-on startup to power-off shutdown of the bus device, which can also be understood as the same running lifecycle (i.e., Runtime) of the server; the disable threshold can be set according to specific circumstances, without specific limitations, such as 20 times or 15 times.
[0045] It is understood that, in this application embodiment, when the number of times the bus device reports events during its current lifecycle exceeds the disable threshold, the reporting function of the bus device for uncorrectable errors will be turned off to prevent it from frequently reporting errors, which would lead to frequent triggering of the processor interrupt mechanism, thus avoiding affecting the normal operation of other services and preventing the blocking of operating system performance.
[0046] Furthermore, since this application has already determined that the bus device has encountered an error, and the operating system is also aware of it, disabling the reporting will not affect the operating system. On the contrary, if it is not disabled, frequent reporting will cause performance blocking of the operating system.
[0047] Specifically, in this application embodiment, if a single bus device triggers N threshold events (N≥1, configurable) within the same Runtime lifecycle (from power-on to the next restart), the reporting function of uncorrectable errors of the corresponding error type of the bus device is disabled.
[0048] For example, if the threshold N is set to 20 times, and bus device A reports 21 uncorrectable errors in the same runtime lifecycle, and the error type is type a, then the error reporting function for type a corresponding to bus device A will be disabled. If bus device B reports 19 uncorrectable errors in the same runtime lifecycle, then there is no need to disable the uncorrectable error reporting function for bus device B.
[0049] In some embodiments of this application, the bus device includes a register.
[0050] Specifically, the register can be the AER (Advanced Error Reporting) register.
[0051] In some embodiments of this application, disabling the reporting function of uncorrectable errors of a bus device includes: generating a disabling configuration for the reporting function if the number of triggers exceeds a disabling threshold; writing the disabling configuration into the register of the bus device; the bus device reading the disabling configuration from the register; and disabling the reporting function of uncorrectable errors of the bus device according to the disabling configuration.
[0052] It is understood that, in the embodiments of this application, when the disable threshold of the bus device in the current life cycle is greater than the disable threshold, a disable configuration for the reporting function is generated, the disable configuration is written into the register of the bus device, the bus device reads the disable configuration in the register, and disables the reporting function of the bus device for uncorrectable errors according to the disable configuration.
[0053] Specifically, in this embodiment of the application, the UCE Mask in the AER register can be set by configuring the PCIe configuration space of the RootPort or slot where the bus device is located.
[0054] In some embodiments of this application, after disabling the reporting function of uncorrectable errors of the bus device, the method further includes: restoring the reporting function of uncorrectable errors of the bus device when the server restarts; and after restoring the reporting function of uncorrectable errors of the bus device, clearing the suppression flag of the bus device and resetting the error count of uncorrectable errors.
[0055] It is understood that the embodiments of this application can restore the reporting function of uncorrectable errors of the bus device when the server restarts. After restoring the reporting function of uncorrectable errors of the bus device, the suppression flag of the bus device is cleared and the error count of uncorrectable errors is reset, so as to re-execute the uncorrectable error processing in a new running cycle after the server restarts, eliminate the interference of historical errors, and avoid misjudging new errors.
[0056] In some embodiments of this application, before restoring the reporting function of the bus device for uncorrectable errors when the server restarts, the method further includes: reading the server's operation log; determining the reason for the server's restart based on the operation log; if the reason for the restart is a startup caused by the bus device triggering a reporting event, then generating a device isolation event in the baseboard management controller and using the device isolation event to isolate the bus device.
[0057] The operation log includes the reason for the server restart, as well as the bus device that triggered the uncorrectable error and the error type; uncorrectable errors include errors that will cause the server to restart or crash, and errors that will not cause the server to restart or crash.
[0058] It is understood that, before restoring the reporting function of the bus device for uncorrectable errors when the server restarts, the embodiments of this application determine the reason for the server restart based on the server's operation log. If the reason for the restart is the start caused by the bus device triggering a reporting event, a device isolation event is generated in the baseboard management controller. The bus device is isolated by the device isolation event to limit the impact of the bus device with the uncorrectable error on other devices and the overall operation of the server. This avoids the problem that the server log cannot be exported and the service cannot be restored in time due to the server still crashing or restarting after restarting.
[0059] It should be noted that the core objective of device isolation events is to limit the impact of a bus device with an uncorrectable error on other devices and the overall operation of the server. By cutting off the fault propagation link, the impact is controlled to a minimum. For example, when a bus device is isolated due to an uncorrectable error, its interaction with other components can be restricted through hardware or software mechanisms (such as prohibiting the bus device from accessing shared resources, stopping the assignment of tasks to the bus device, or cutting off communication links with other devices). Through these measures, the error of the bus device with the uncorrectable error cannot be transmitted to other devices, thereby preventing a single point of failure from spreading into a system-wide failure.
[0060] In step S103, a suppression flag is assigned to the bus device within the time window, and the error count of uncorrectable errors within the time window is obtained. The suppression flag is used to suppress the reporting of uncorrectable errors when the error count within the time window is less than the reporting threshold.
[0061] The reporting threshold can be determined based on specific circumstances, without any specific limitation, such as setting it to 15 times or 10 times. The reporting threshold can be preset with a default value in the code (saved to flash), or it can be set as an option in BIOS stepup, and can be modified later through stepup.
[0062] It is understood that the embodiments of this application can assign a suppression flag to the bus device within a time window and obtain the error count of uncorrectable errors within the time window, so as to determine the processing of the bus device that triggered the uncorrectable error based on the error count. The suppression flag is used to suppress the reporting of uncorrectable errors when the error count is less than the reporting threshold within the time window, thereby reducing redundant error reporting, suppressing error reporting storms, and thus reducing the impact of frequent error triggering on the availability of the server system and the normal operation of other services, and improving the reliability of the server.
[0063] In step S104, if the error count is greater than or equal to the reporting threshold, an uncorrectable error reporting event is triggered. If the time window expires and the error count is less than the reporting threshold, the suppression flag of the bus device is cleared and the error count of the uncorrectable error is reset.
[0064] It is understood that in this embodiment of the application, when the error count is greater than or equal to the reporting threshold within the time window, the reporting event of uncorrectable errors is triggered again to avoid omission of errors that exacerbate the error. When the time window expires and the error count is less than the reporting threshold, it indicates that the bus device will not cause an error reporting storm. Therefore, the suppression flag of the bus device is cleared and the error count of uncorrectable errors is reset. When uncorrectable errors occur in the future, they can be reported normally, and it is convenient to re-count them in the next time window.
[0065] For example, in this embodiment of the application, the error statistics period is 5 minutes and the reporting threshold is 10 times. When bus device A triggers an uncorrectable error for the first time, it triggers an SMI interrupt. After processing by the BIOS, the error is reported to the board management controller and the operating system. An uncorrectable error suppression flag (such as UCE_SUPPRESSION_FLAG) is set for bus device A, and an error statistics time window (UCE_ERROR_WINDOW_TIMER) is opened for 5 minutes. If the number of uncorrectable error reports by bus device A is 10 within the 5-minute error statistics time window, reaching the reporting threshold of 10 times, then it is reported to the board management controller and the operating system again. If the number of uncorrectable error reports by bus device A is 9 times, not reaching the reporting threshold of 10 times, and the error statistics time window times out, then the uncorrectable error suppression flag (such as UCE_SUPPRESSION_FLAG) is cleared, and the uncorrectable error count (uce_error_count=0) is reset.
[0066] Specifically, since bus devices may trigger uncorrectable errors due to faults or unstable links, frequent uncorrectable errors will continuously trigger SMI interrupts, causing the CPU to repeatedly enter SMM mode, forming an error reporting storm, consuming a large amount of server system resources, and potentially masking other critical business needs, severely affecting the availability of the server system. Therefore, this application embodiment can reduce the frequency of uncorrectable error reporting by setting an uncorrectable error suppression flag for bus devices and a time window for uncorrectable error statistics. When an uncorrectable error is first triggered, it is reported to capture the uncorrectable error and ensure that the initial error is not missed. Furthermore, when the error count in the time window exceeds the reporting threshold, it will be reported again to avoid missing errors that exacerbate the problem, reduce redundant error reporting, suppress the error reporting storm, and thus reduce the impact of frequent error triggering on the availability of the server system and the normal operation of other businesses, thereby improving the reliability of the server.
[0067] In some embodiments of this application, if the error count is greater than or equal to the reporting threshold, after triggering the reporting event of an uncorrectable error, the method further includes: reopening the time window and adjusting the reporting threshold after triggering the reporting event of an uncorrectable error, wherein the adjusted reporting threshold is greater than the original reporting threshold; maintaining the suppressed state of the bus device and continuing the error count of uncorrectable errors.
[0068] It is understood that, after triggering the reporting event of an uncorrectable error, the embodiments of this application reopen the time window and adjust the reporting threshold to dynamically adapt to the changing trend of uncorrectable errors. The adjusted reporting threshold is greater than the original reporting threshold to better suppress the error reporting frequency, maintain the suppression state of the bus device, and continue the error counting of uncorrectable errors.
[0069] In addition, it should be noted that the solution in this application provides dual implementation paths for BIOS and CPU firmware to be compatible with different hardware architecture requirements.
[0070] In summary, this application reduces redundant error reporting through dynamic threshold control and error counting mechanisms; introduces a UCE error statistics time window and suppression flag to avoid the CPU frequently entering SMM mode and reduce the frequency of SMI interrupts triggered by UCE; and suppresses redundant error information of non-critical devices while ensuring the reporting of critical alarms.
[0071] Based on the above-described method for handling uncorrectable errors in bus devices, this application also provides a system for handling uncorrectable errors in bus devices, the system architecture of which is as follows: Figure 2 As shown, by embedding an intelligent suppression module in the UEFI Runtime phase, including an error counter, a time window timer, and a device state machine, closed-loop control is achieved through SMI interrupt service routines. Figure 2 In PCIe AER, AER is a mechanism for detecting and reporting errors that occur in PCIe devices, allowing PCIe devices to detect and report various types of errors.
[0072] The following is combined Figure 2 This describes the entire implementation process of the uncorrectable error handling method for bus devices, including:
[0073] 1. UCE error reporting and suppression process.
[0074] (1) Mechanism for immediate reporting of first-time errors.
[0075] When a PCIe device triggers a UCE error for the first time, an SMI interrupt is immediately triggered. The BIOS processes the interrupt and reports it to the BMC and notifies the OS. The device path and UCE error that triggered the error are also recorded.
[0076] The BIOS processing includes error logging, error parsing (including parsing the specific device address where the error occurred), then reporting it to the BMC and notifying the OS for further processing; the Device path is the unique path that represents a PCIe device during the UEFI BIOS boot phase.
[0077] The BIOS sets the UCE suppression flag (such as UCE_SUPPRESSION_FLAG) to enable the error statistics time window (UCE_ERROR_WINDOW_TIMER).
[0078] The design process for the error time statistics window can be as follows: Set the timer to 5 minutes. During the timer period, count the UCE errors. When the error count exceeds the set threshold, report it once, then clear the error count and continue counting. When the timer expires, if the UCE error count is still less than the set threshold, clear the UCE error count and start the timer again.
[0079] (2) Dynamic threshold suppression mechanism for UCE errors.
[0080] Within the time window, in the SMI interrupt service function, count the UCE errors of devices that have reported a UCE error once (uce_error_count, regardless of the UCE error type).
[0081] If the count reaches the preset threshold, report the error to the BMC and notify the OS again; maintain the suppressed state and do not reset the UCE error counter (uce_error_count).
[0082] In addition, it should be noted that the UCE error counting and threshold judgment in this application are performed in the SMI handler. However, these logical processes take up SMM time in the microsecond range and exit immediately after processing, so the impact on OS performance is minor.
[0083] (3) The suppression state is automatically released.
[0084] If the following conditions are met simultaneously, the UCE error count of the same device within the window period is less than the preset threshold, and the error statistics time window (UCE_ERROR_WINDOW_TIMER) times out; clear the UCE suppression flag (such as UCE_SUPPRESSION_FLAG) and reset the UCE error count (uce_error_count=0).
[0085] (4) Device-level permanent disable mechanism.
[0086] If a single device triggers N threshold events (N≥1, configurable) within the same Runtime lifecycle (from power-on to the next reboot), the reporting function for the corresponding error type of that device will be disabled.
[0087] Specifically, this is achieved by configuring the PCIe configuration space of the RootPort or slot where the device is located, and setting the UCE Mask in the AER register.
[0088] (5) System-level state reset.
[0089] When the server restarts, clear the UCE suppression flag on all devices, reset all UCE error counters, and restore the UCE error reporting function on all devices by default.
[0090] 2. Backtracking suppression mechanism after system crash or restart.
[0091] (1) Enabled by retrospective inhibition mechanism.
[0092] When the system crashes or restarts due to a UCE error, the BIOS parses the error log (stored in NVRAM) from the last run during the POST phase; identifies the fatal UCE device that triggered the crash (via the Fatal ErrorReceived flag in the PCIe AER log); dynamically suppresses the reporting of all UCE error types for that device; and generates a device isolation event in the BMC (EventLog: "PCIe Device [BDF] disabled due to fatal UCE").
[0093] (2) The retrospective inhibition mechanism is cancelled.
[0094] When maintenance personnel discover a server device crash or restart alarm, they can troubleshoot the faulty device based on the recorded logs and repair or replace it. Then, they can set a PCIe device fault recovery flag in the BMC. When the server device restarts again, the BIOS sends an IPMI command to the BMC to obtain the PCIe device fault recovery flag. If the flag is set, the device's backtracking suppression mechanism is canceled.
[0095] Clear the UCE suppression flag on all devices, reset all UCE error counters, and restore the UCE error reporting function on all devices by default.
[0096] In summary, this application can reduce the frequency of redundant PCIe UCE error reporting and SMI interrupt triggering by using an immediate first error reporting mechanism, dynamic threshold control, error counting and time window management, and a backtracking suppression mechanism after a crash or restart. This enables closed-loop processing of crash scenarios. By using hardware register expansion and a three-level suppression strategy, it solves the vicious cycle problem of repeated system crashes caused by faulty devices, effectively improving the reliability of the server system and the performance of system services.
[0097] In summary, the main technical means of this application include:
[0098] 1. Pioneering error-crash cause-effect chain analysis: Accurately pinpoint faulty devices by using the pre-death log stored in NVRAM (error logs saved after the last runtime crash or restart).
[0099] 2. Three-tiered suppression strategy: temporary suppression (window period) → permanent suppression (within the life cycle) → forced suppression (lethal device).
[0100] 3. Lifecycle awareness: Continuously track device health throughout the Runtime cycle.
[0101] 4. Non-destructive reset: Full monitoring will be automatically restored after restarting.
[0102] 5. Hardware-level isolation: Extend the PCIe AER register to enable hard shutdown of the device-level reporting channel.
[0103] The execution flow of the server uncorrectable error handling system of this application is described below through a specific embodiment.
[0104] PCIe network card A of a certain server frequently triggers UCE due to hardware link noise.
[0105] 1. First UCE error triggered (1st error).
[0106] Error occurred: Network card A triggered the first UCE error at 09:00:00.
[0107] System response: The hardware automatically triggers an SMI interrupt, the CPU enters SMM mode, and the BIOS takes over the processing.
[0108] The BIOS reads the network card's PCIe configuration space, records error information (device BDF, error type, timestamp), and writes it to the log.
[0109] The BIOS reports the error to the BMC (logging: "PCIe Device A UCE error (1st occurrence)"), and at the same time notifies the OS (the OS generates an alarm: "Network card A detected an uncorrectable error").
[0110] Threshold and cycle initialization: The initial error reporting threshold is set to "3 times / 5 minutes". After the first error, the counter = 1, and the first 5-minute statistical cycle (09:00:00-09:05:00) begins.
[0111] 2. The error is triggered again during the window period (the second error).
[0112] Error occurred at 09:01:30, the network card triggered the same UCE error again (threshold not reached).
[0113] System response: The error count is incremented to 2 (still < threshold 3), and intermediate errors are not reported repeatedly (to avoid SMI storm).
[0114] The BIOS only updates internal counters and does not send duplicate notifications to the BMC / OS, ensuring that CPU resources are used for normal business operations.
[0115] 3. When the error count reaches the threshold, a second report is triggered and the threshold is adjusted.
[0116] Error occurred at 09:03:00, the network card triggered the UCE error for the third time (count = 3, threshold reached).
[0117] System response: The BIOS triggers the processing flow again: record the error information, and report to the BMC (log: "PCIe DeviceA UCE error (3rd occurrence, threshold reached)") and the OS (alarm: "Network card error frequency exceeded, please pay attention").
[0118] Adjust threshold and enter new cycle: To enhance sensitivity to subsequent errors, the reporting threshold for the next cycle is adjusted from "3 times" to "2 times". Reset counter = 0 and enter the second statistical cycle (09:03:00-09:08:00).
[0119] 4. Errors persist during the new cycle, triggering isolation.
[0120] Error occurred at 09:06:00 (within the new cycle), the network card triggered the UCE error for the fourth time (count = 1); at 09:07:00, it triggered the error for the fifth time (count = 2, reaching the new threshold of 2).
[0121] System response: The system re-reported to the BMC and OS, indicating "error continues to escalate." Due to the repeated reaching of the threshold within a short period, the system determined it to be a "persistent fault" and triggered device isolation.
[0122] The BIOS disables the network card via the PCIe bus (cutting off its data transmission channel). The BMC logs the isolation event: "PCIe Device A isolated due to repeated UCE errors".
[0123] According to the bus device uncorrectable error handling method proposed in the embodiments of this application, when the bus device first triggers an uncorrectable error, an error report is made, and a time window for uncorrectable error statistics is opened. Within the time window, a suppression flag is assigned to the bus device. If the error count within the time window is below the reporting threshold, no error report is made. If the error count is greater than or equal to the reporting threshold, a report is made again. By setting the suppression flag for uncorrectable errors of the bus device and the time window for uncorrectable error statistics, the frequency of uncorrectable error reporting is reduced. Furthermore, when an uncorrectable error is first triggered, it is reported to capture the uncorrectable error and ensure that the initial error is not missed. When the error count within the time window is greater than the reporting threshold, a report is also made again to avoid missing the processing of errors that aggravate the error, reduce redundant error reporting, suppress error reporting storms, and thus reduce the impact of frequent error triggering on the availability of the server system and the normal operation of other services, thereby improving the reliability of the server.
[0124] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.
[0125] Embodiments of this application also provide a bus device uncorrectable error handling apparatus.
[0126] Figure 3 This is a schematic diagram of a bus device uncorrectable error handling apparatus provided according to an embodiment of this application.
[0127] like Figure 3 As shown, the bus device uncorrectable error handling device 10 includes: a detection module 100, a trigger module 200, an assignment module 300, and a processing module 400.
[0128] The detection module 100 is used to detect whether the bus device has triggered an uncorrectable error for the first time.
[0129] The trigger module 200 is used to trigger an uncorrectable error reporting event and open a time window for uncorrectable error statistics if an uncorrectable error is detected for the first time by a bus device.
[0130] The assignment module 300 is used to assign a suppression flag to the bus device within a time window, obtain the error count of uncorrectable errors within the time window, and the suppression flag is used to suppress the reporting of uncorrectable errors when the error count within the time window is less than the reporting threshold.
[0131] The processing module 400 is used to trigger an uncorrectable error reporting event if the error count is greater than or equal to the reporting threshold, and to clear the suppression flag of the bus device and reset the error count of the uncorrectable error if the time window expires and the error count is less than the reporting threshold.
[0132] In some embodiments of this application, the detection module 100 is further configured to: detect whether the bus device triggers an uncorrectable error; when the bus device triggers an uncorrectable error, identify whether the bus device carries a suppression flag; if the bus device carries a suppression flag, determine that the bus device has not triggered an uncorrectable error for the first time; if it does not carry a suppression flag, determine that the bus device has triggered an uncorrectable error for the first time.
[0133] In some embodiments of this application, the bus device uncorrectable error handling device 10 of this application embodiment further includes: an adjustment module.
[0134] The adjustment module is used to, if the error count is greater than or equal to the reporting threshold, trigger an uncorrectable error reporting event, then reopen the time window and adjust the reporting threshold, wherein the adjusted reporting threshold is greater than the original reporting threshold; maintain the suppressed state of the bus device and continue the uncorrectable error counting.
[0135] In some embodiments of this application, the triggering module 200 is further configured to: if an uncorrectable error is detected for the first time by the bus device, call the interrupt service function of the bus device and use the interrupt service function to trigger a device interrupt of the bus device; after triggering the device interrupt of the bus device, report the uncorrectable error.
[0136] In some embodiments of this application, the reported event is at least one of the baseboard management controller and the operating system of the server that reports an uncorrectable error.
[0137] In some embodiments of this application, the bus device uncorrectable error handling device 10 of this application embodiment further includes: a shutdown module.
[0138] The shutdown module is used to obtain the number of times the bus device has reported the event during its current lifecycle after triggering the reporting event of an uncorrectable error. The current lifecycle is the period from power-on to power-off of the bus device. If the number of triggers is greater than the disable threshold, the reporting function of the bus device for uncorrectable errors is disabled. The reporting function is used to trigger the reporting event of an uncorrectable error.
[0139] In some embodiments of this application, the bus device includes a register.
[0140] In some embodiments of this application, the shutdown module is further configured to: generate a disabled configuration for the reporting function if the number of triggers exceeds the disable threshold; write the disabled configuration into the register of the bus device; the bus device reads the disabled configuration from the register; and disables the reporting function of the bus device for uncorrectable errors according to the disabled configuration.
[0141] In some embodiments of this application, the bus device uncorrectable error handling device 10 of this application embodiment further includes: a recovery module.
[0142] The recovery module is used to restore the reporting function of uncorrectable errors of the bus device when the server restarts after the uncorrectable error reporting function of the bus device is disabled; after restoring the reporting function of uncorrectable errors of the bus device, it clears the suppression flag of the bus device and resets the error count of uncorrectable errors.
[0143] In some embodiments of this application, the bus device uncorrectable error handling device 10 of this application embodiment further includes: a determination module.
[0144] The determination module is used to read the server's operation log before restoring the reporting function of the bus device for uncorrectable errors when the server restarts; determine the reason for the server restart based on the operation log; if the reason for the restart is the start caused by the bus device triggering a reporting event, then the baseboard management controller generates a device isolation event and uses the device isolation event to isolate the bus device.
[0145] It should be noted that the description of the features in the embodiment corresponding to the bus device uncorrectable error handling device can be found in the relevant description of the embodiment corresponding to the bus device uncorrectable error handling method, and will not be repeated here.
[0146] Embodiments of this application also provide a server, such as Figure 4 As shown, it includes a memory 401 and a processor 402. The memory 401 stores a computer program, and the processor 402 is configured to run the computer program to perform the steps in any of the above embodiments of the bus device uncorrectable error handling method.
[0147] Embodiments of this application also provide a non-volatile computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above embodiments of the bus device uncorrectable error handling method when running.
[0148] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.
[0149] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above embodiments of the bus device uncorrectable error handling method.
[0150] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0151] The foregoing has provided a detailed description of a method for handling uncorrectable errors in a bus device and a server provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and its core ideas. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. A method for handling uncorrectable errors in a bus device, characterized in that, Includes the following steps: Detect whether the bus device has triggered an uncorrectable error for the first time; The method of detecting whether a bus device has triggered an uncorrectable error for the first time includes: detecting whether the bus device has triggered the uncorrectable error; when the bus device triggers the uncorrectable error, identifying whether the bus device carries a suppression flag; if the bus device carries the suppression flag, determining that the bus device has not triggered the uncorrectable error for the first time, and if it does not carry the suppression flag, determining that the bus device has triggered the uncorrectable error for the first time, wherein the suppression flag is used to determine whether the bus device has triggered the uncorrectable error for the first time; If the bus device is detected to have triggered the uncorrectable error for the first time, an uncorrectable error reporting event is triggered, and a time window for uncorrectable error statistics is opened. Within the time window, a suppression flag is assigned to the bus device, and the error count of the uncorrectable error within the time window is obtained. The suppression flag is used to suppress the reporting of the uncorrectable error when the error count within the time window is less than the reporting threshold. If the error count is greater than or equal to the reporting threshold, an uncorrectable error reporting event is triggered. If the time window times out and the error count is less than the reporting threshold, the suppression flag of the bus device is cleared, and the error count of the uncorrectable error is reset. If the error count is greater than or equal to the reporting threshold, after triggering the uncorrectable error reporting event, the method further includes: reopening the time window and adjusting the reporting threshold after triggering the uncorrectable error reporting event, wherein the adjusted reporting threshold is greater than the original reporting threshold; maintaining the suppression state of the bus device and continuing the uncorrectable error counting; after triggering the uncorrectable error reporting event... It also includes: monitoring the status of the server; if the server crashes or restarts, and the reason for the crash or restart is that the bus device triggered the reported event, then the fault flag of the bus device is set to a first flag; after the bus device is repaired or replaced, the fault flag of the bus device is set to a second flag; when the server restarts next time, the suppression flags of all bus devices are cleared, the error count of uncorrectable errors of all bus devices is reset, and the reported events of all bus devices when the uncorrectable error is triggered are restored, wherein the first flag indicates that the bus device has an uncorrectable error, and the second flag indicates that the uncorrectable error of the bus device has been recovered.
2. The method for handling uncorrectable errors in bus devices according to claim 1, characterized in that, If the bus device is detected to have triggered the uncorrectable error for the first time, an uncorrectable error reporting event is triggered, including: If the bus device is detected to have triggered the uncorrectable error for the first time, the interrupt service function of the bus device is invoked, and the device interrupt of the bus device is triggered by the interrupt service function. After triggering a device interrupt on the bus device, the uncorrectable error is reported.
3. The method for handling uncorrectable errors in bus devices according to claim 1 or 2, characterized in that, The reported event is at least one of the baseboard management controller and operating system that reports the uncorrectable error to the server.
4. The method for handling uncorrectable errors in bus devices according to claim 1, characterized in that, Following the reporting event that triggers the uncorrectable error, the following is also included: The number of times the reported event is triggered by the bus device during its current lifecycle is obtained, wherein the current lifecycle is the period from power-on startup to power-off shutdown of the bus device; If the number of triggers exceeds the disable threshold, then the reporting function of the uncorrectable error of the bus device is disabled, wherein the reporting function is used to trigger the reporting event of the uncorrectable error.
5. The method for handling uncorrectable errors in bus devices according to claim 4, characterized in that, The bus device includes registers, and disabling the reporting function of the uncorrectable error of the bus device includes: If the number of triggers exceeds the disable threshold, a disable configuration for the reporting function is generated; The disabled configuration is written into the register of the bus device, the bus device reads the disabled configuration from the register, and disables the reporting function of the bus device for uncorrectable errors according to the disabled configuration.
6. The method for handling uncorrectable errors in bus devices according to claim 4, characterized in that, After disabling the reporting function of the uncorrectable error of the bus device, the following is also included: Restore the reporting function of the uncorrectable error of the bus device when the server restarts; After restoring the reporting function of the uncorrectable error of the bus device, the suppression flag of the bus device is cleared and the error count of the uncorrectable error is reset.
7. The method for handling uncorrectable errors in bus devices according to claim 6, characterized in that, Before restoring the reporting function of the uncorrectable error of the bus device upon server restart, the following is also included: Read the server's operation logs; The reason for the server restart was determined based on the operation logs. If the restart is caused by the bus device triggering the reported event, then a device isolation event is generated in the baseboard management controller, and the bus device is isolated using the device isolation event.
8. A server, characterized in that, Its features include: Memory, used to store computer programs; A processor, configured to implement the steps of the bus device uncorrectable error handling method as described in any one of claims 1 to 7 when executing the computer program.
Citation Information
Patent Citations
System and method for counting storage device-related errors utilizing a sliding window
US8156382B1