Troubleshooting methods, systems, and media
By building an adaptive storm determination mechanism, combining historical fault information and current system load, and using sliding time window to count fault information, the problem of misjudgment and misjudgment in hardware fault processing is solved, achieving higher fault identification accuracy and rapid response.
Patent Information
- Application Number
- CN202510838987.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2045-06-20
AI Technical Summary
In the prior art, the hardware fault handling mechanism relies on static thresholds to cause misjudgment or misjudgment, which cannot adapt to the environmental state and fault trends at different operating stages, and the fault handling accuracy is low.
By combining historical fault information and current system load status, an adaptive storm determination mechanism is built, a sliding time window is used to count the fault information, and the storm trigger flag is set when the preset conditions are met to achieve dynamic adjustment and hierarchical response.
It improves the accuracy and real-time nature of fault identification, reduces misjudgment and misjudgment, ensures rapid response and implements targeted measures in case of fault-intensive situations, and improves the accuracy and adaptability of fault handling.
Smart Images

Figure CN120354246B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of information processing technology, and in particular to a fault handling method, a fault handling system, and a medium. Background Art
[0002] In current data centers and computing platforms, hardware error monitoring mechanisms are often implemented to ensure stable system operation. For example, through the Reliability, Availability, and Serviceability (RAS) architecture, various hardware fault errors generated during system operation are detected, reported, and responded to.
[0003] In related technologies, hardware error handling mechanisms often rely on statically set fixed thresholds. Once the threshold trigger condition is met, the error is reported to the operating system or a preset policy is executed. Because the threshold is statically configured, it cannot be dynamically adjusted to the operating environment status or fault trends at different operating stages. This can lead to over-responses to some fault events, increasing the probability of misjudgment. Other abnormal behaviors may be ignored because they do not reach the fixed threshold, resulting in missed judgments. Therefore, a fault handling method is urgently needed to address the current technical problem of low fault handling accuracy. Summary of the Invention
[0004] The present application provides a fault handling method, a fault handling system, and a medium to at least solve the problem of low accuracy of fault handling in related technologies.
[0005] In a first aspect, the present application provides a fault handling method, comprising:
[0006] receiving fault information of at least one hardware component sent by a basic input / output system, and acquiring historical fault information of at least one hardware component based on the stored fault information;
[0007] The fault information is obtained by the basic input and output system when a hardware component fails and triggers an interrupt;
[0008] Obtain the current system load information, and based on the historical fault information and the current system load information, obtain the current storm judgment conditions corresponding to each hardware component;
[0009] Obtaining cumulative window fault information of at least one hardware component according to the current sliding time window, and setting a storm trigger flag of the hardware component when the cumulative window fault information of the hardware component meets its corresponding current storm determination condition, and recording the hardware component as a storm trigger component;
[0010] Fault information of at least one storm triggering component is received, and when the fault information of at least one storm triggering component meets any preset suppression trigger condition, a suppression flag of at least one storm triggering component is set.
[0011] In a second aspect, the present application provides a fault handling method, comprising:
[0012] triggering an interrupt when at least one hardware component fails, and obtaining failure information of the at least one hardware component;
[0013] The fault information of at least one hardware component is sent to the baseboard control manager so that the baseboard control manager receives the fault information of at least one hardware component and obtains the historical fault information of at least one hardware component based on the stored fault information; obtains the current system load information, and obtains the current storm judgment condition corresponding to each hardware component based on the historical fault information and the current system load information; obtains the window cumulative fault information of at least one hardware component according to the current sliding time window, and sets the storm trigger flag of the hardware component when the window cumulative fault information of the hardware component meets its corresponding current storm judgment condition, and records the hardware component as a storm trigger component; receives the fault information of at least one storm trigger component, and sets the suppression flag of at least one storm trigger component when the fault information of at least one storm trigger component meets any preset suppression trigger condition.
[0014] In a third aspect, the present application provides a fault handling method, comprising:
[0015] The basic input and output system triggers an interrupt when at least one hardware component fails, and obtains failure information of the at least one hardware component;
[0016] The basic input and output system sends a fault message of at least one hardware component;
[0017] The baseboard management controller receives fault information of at least one hardware component and obtains historical fault information of at least one hardware component based on the stored fault information;
[0018] The baseboard management controller obtains the current system load information and obtains the current storm judgment condition corresponding to each hardware component based on the historical fault information and the current system load information;
[0019] The baseboard management controller obtains the window cumulative fault information of at least one hardware component according to the current sliding time window, and sets the storm trigger flag of the hardware component when the window cumulative fault information of the hardware component meets its corresponding current storm judgment condition, and records the hardware component as a storm trigger component;
[0020] The basic input and output system monitors the storm flag in real time, and after reading the storm flag of the storm trigger component, it blocks the fault information of the storm trigger component from being passed to the upper layer;
[0021] The baseboard management controller receives fault information of at least one storm triggering component, and sets a suppression flag of at least one storm triggering component to a position when the fault information of at least one storm triggering component meets any preset suppression trigger condition;
[0022] The basic input and output system monitors the suppression flag in real time, and executes the suppression strategy corresponding to the suppression flag after reading that the suppression flag is set.
[0023] In a fourth aspect, the present application further provides a baseboard management controller, comprising:
[0024] a data receiving module, configured to receive fault information of at least one hardware component sent by a basic input / output system, and obtain historical fault information of at least one hardware component based on stored fault information; wherein the fault information is obtained by the basic input / output system when an interrupt is triggered by a hardware component fault;
[0025] The condition generation module is used to obtain the current system load information and obtain the current storm judgment conditions corresponding to each hardware component based on the historical fault information and the current system load information;
[0026] A storm judgment module is used to obtain the window cumulative fault information of at least one hardware component according to the current sliding time window, and when the window cumulative fault information of the hardware component meets its corresponding current storm judgment condition, set the storm trigger flag of the hardware component and record the hardware component as a storm trigger component;
[0027] The suppression monitoring module is used to receive fault information of at least one storm triggering component and set the suppression flag of at least one storm triggering component to a position when the fault information of at least one storm triggering component meets any preset suppression triggering condition.
[0028] In a fifth aspect, the present application further provides a basic input and output system, including:
[0029] an information acquisition module, configured to trigger an interrupt when at least one hardware component fails, and obtain failure information of the at least one hardware component;
[0030] An information sending module is used to send the fault information of at least one hardware component to the baseboard control manager, so that the baseboard control manager receives the fault information of at least one hardware component and obtains the historical fault information of at least one hardware component based on the stored fault information; obtains the current system load information, and obtains the current storm judgment condition corresponding to each hardware component based on the historical fault information and the current system load information; obtains the window cumulative fault information of at least one hardware component according to the current sliding time window, and sets the storm trigger flag of the hardware component when the window cumulative fault information of the hardware component meets its corresponding current storm judgment condition, and records the hardware component as a storm trigger component; receives the fault information of at least one storm trigger component, and sets the suppression flag of at least one storm trigger component when the fault information of at least one storm trigger component meets any preset suppression trigger condition.
[0031] In a sixth aspect, the present application further provides a fault handling system, comprising: a baseboard management controller, a basic input and output system, and a processor;
[0032] The baseboard management controller includes a microprocessor and a management control memory, the management control memory is used to store a management control program, and the microprocessor is used to execute the management control program stored in the management control memory to perform the steps of the fault handling method according to the first aspect;
[0033] The basic input and output system is used to store basic input and output programs;
[0034] The processor is used to execute a basic input and output program stored in the basic input and output system to perform the steps of the fault handling method according to the second aspect.
[0035] In a seventh aspect, the present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned fault handling methods when executing the computer program.
[0036] In an eighth aspect, the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program implements the steps of any of the above-mentioned fault handling methods when executed by a processor.
[0037] In a ninth aspect, the present application also provides a computer program product, comprising a computer program, which implements the steps of any of the above-mentioned fault handling methods when executed by a processor.
[0038] This application obtains current system load information and combines it with historical fault information to obtain the storm judgment conditions of hardware components, thereby constructing a dynamic perception and judgment mechanism for error storms with adaptive characteristics. It can dynamically adjust the judgment threshold according to load fluctuations and historical fault trends, reduce misjudgments or missed judgments caused by unreasonable threshold settings, and ensure that the storm judgment results are more targeted and accurate. Then, by statistically analyzing the window-accumulated fault information of the hardware components within the current sliding time window and judging whether the corresponding storm judgment conditions are met, it achieves rapid identification of error storms and improves the real-time and accuracy of fault detection. After identifying the storm triggering component, it continues to receive its fault information. When the fault information of the storm triggering component meets any preset condition, its suppression flag is set to the position for subsequent execution of targeted response strategies, ensuring that when faults continue to occur intensively, it can respond quickly and implement corresponding measures. Therefore, this application can solve the current technical problem of low fault processing accuracy and achieve the technical effect of improving fault identification accuracy and rapid response. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0040] Figure 1 Schematic diagram of the process of the troubleshooting method provided in the embodiment of the application Figure 1 ;
[0041] Figure 2 A schematic flow chart of a method for determining the current sliding time window provided in an embodiment of the present application;
[0042] Figure 3 Schematic diagram of the troubleshooting method provided in this application embodiment Figure 2 ;
[0043] Figure 4 Schematic diagram of the process of the troubleshooting method provided in the embodiment of the application Figure 3 ;
[0044] Figure 5 A timing diagram of the fault handling method provided in an embodiment of the present application;
[0045] Figure 6 A schematic diagram of the structure of a baseboard management controller provided in an embodiment of the present application;
[0046] Figure 7 A schematic diagram of the structure of the basic input and output system provided in an embodiment of the present application;
[0047] Figure 8 This is a schematic diagram of the structure of the electronic device provided in this application. DETAILED DESCRIPTION
[0048] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0049] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.
[0050] In related technologies, hardware error handling logic generally uses fixed threshold strategies, which are inflexible when faced with dynamically changing load conditions, different hardware lifecycle states, and error outbreak patterns. For example, in high-load or complex application environments, a large number of transient errors may occur. These errors trigger response strategies without sufficient analysis, which can easily lead to misjudgments and excessive processing.
[0051] Based on the above technical problems and needs, the inventive concept of this application is to propose a fault handling method, which combines historical fault information with the current system load status to construct an adaptive storm judgment mechanism, and realizes hierarchical response control through hierarchical suppression.
[0052] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0053] Figure 1 Schematic diagram of the process of the troubleshooting method provided in the embodiment of the application Figure 1 ,like Figure 1 Shown, including:
[0054] S11 , receiving fault information of at least one hardware component sent by a basic input / output system, and acquiring historical fault information of at least one hardware component according to stored fault information.
[0055] In this embodiment, fault information of at least one hardware component sent by the Basic Input Output System (BIOS) is first received to sense the abnormal state of the current hardware level. The Basic Input Output System can monitor various hardware components and generate interrupt events before the operating system is loaded. Once an abnormality is detected in a hardware component during operation, the BIOS will obtain and send fault information according to the preset interrupt response mechanism. After receiving the fault information, the historical fault information corresponding to the hardware component will be retrieved from the local or other storage structures. By obtaining historical faults and retrospectively modeling the historical behavior patterns of hardware components, a data foundation is provided for subsequent steps such as error trend analysis, storm identification, and dynamic threshold calculation.
[0056] S12, obtaining current system load information, and obtaining current storm determination conditions corresponding to each hardware component based on historical fault information and current system load information.
[0057] In this embodiment, by obtaining the current system load information and combining it with historical fault information, the storm judgment conditions of each hardware component in the current operating state are determined, aiming to provide a dynamic threshold basis for subsequent fault storm identification. Among them, load information refers to comprehensive status data that reflects the operating pressure of the entire system at a certain moment, which usually comes from the operating status monitoring of multiple subsystems such as processors, memory, storage, and networks. The method of obtaining system load information can be based on the resource monitoring interface, underlying hardware sensors, or embedded management modules provided by the operating system, and can be determined according to the actual deployment environment. Historical fault information reflects the fault characteristics exposed by each hardware component in past operation, such as error frequency, type, timing pattern, etc., and is important basic data for analyzing the stability of hardware components.
[0058] By combining historical fault information with the current load status, a strategy for dynamic judgment is constructed to adjust the tolerance for fault behavior according to the current load level of the system, making it more sensitive to similar errors in a high-load environment, thereby improving the timeliness and accuracy of fault response. On the contrary, when the load is light, the tolerance threshold can be appropriately increased to reduce overreaction to minor occasional errors. Finally, based on the above analysis results, a current storm judgment condition is generated for each hardware component, which serves as the judgment benchmark for identifying abnormal fault behavior in subsequent steps. This mechanism realizes real-time updating and personalized setting of the judgment thresholds of different hardware components, solving the problem of insufficient accuracy caused by the inability of traditional static thresholds to adapt to changes in operating status, while improving the adaptability and robustness of the entire fault handling.
[0059] S13, according to the current sliding time window, obtain the window cumulative fault information of at least one hardware component, and when the window cumulative fault information of the hardware component meets its corresponding current storm judgment condition, set the storm trigger flag of the hardware component and record the hardware component as a storm trigger component.
[0060] In this embodiment, the current sliding time window is first determined to limit the observation interval of fault statistics. The sliding time window is a continuously updated time period used to track and analyze the distribution of fault events during recent operation. Its setting ensures the timeliness and continuity of fault statistics. After the time window is determined, the cumulative fault information of at least one hardware component within the time window is collected. The cumulative fault information is the aggregated statistical result of historical fault data within the sliding time window, and is usually recorded separately by fault type, such as the frequency of correctable errors and uncorrectable errors. This information can reflect the operational stability of the hardware component within the time period.
[0061] Subsequently, the accumulated fault information of the current window is compared according to the storm determination conditions obtained previously. The storm determination conditions are used to define the degree of fault density that should be considered as an abnormal storm state. If the accumulated fault information of a hardware component within the sliding time window meets its corresponding storm determination conditions, the storm trigger flag of the hardware component is set, indicating that the hardware component has entered a state that may trigger a fault storm. The hardware component that is set will be marked as a storm trigger component so that targeted fault suppression and isolation strategies can be executed later. This embodiment effectively identifies components with intensive fault outbreaks by introducing a sliding time window statistical mechanism, thereby achieving early detection and response to error storms in fault handling, and improving the automation and accuracy of fault handling.
[0062] S14, receiving fault information of at least one storm triggering component, and setting a suppression flag of at least one storm triggering component when the fault information of at least one storm triggering component meets any preset suppression triggering condition.
[0063] In this embodiment, by continuously monitoring the fault behavior of storm-triggering components and further determining their status based on preset suppression trigger conditions, dynamic control and proactive intervention of the impact of fault storms are achieved. After identifying a storm-triggering component, subsequent fault information generated by the storm-triggering component is continuously received to indicate whether the current state of the storm-triggering component is deteriorating or has a tendency to continuously disrupt its operation. After receiving fault information from at least one storm-triggering component, the current fault status is compared and determined based on pre-set suppression trigger conditions. These suppression trigger conditions are used to reflect different levels of risk scenarios and may include multiple dimensions such as failure frequency thresholds, error density, short-term outbreak trends, and the presence of critical component linkage. The diversity and hierarchical configuration of pre-set suppression trigger conditions enable more detailed and accurate response strategies for complex fault scenarios. When the fault information of a storm-triggering component meets any suppression trigger condition, the suppression flag bit of the storm-triggering component is set. Setting the suppression flag bit not only indicates the current fault status of the storm-triggering component but also indicates that the storm-triggering component has triggered a specific prevention and control mechanism. This embodiment achieves dynamic scenario stratification and accurate intervention after storm identification by continuously tracking and analyzing the subsequent fault behavior of storm-triggered components, ensuring that the fault storm is effectively contained before it spreads widely, reducing the occurrence of performance degradation or unrecoverable failures caused by persistent failures.
[0064] Figure 2 A flow chart of a method for determining the current sliding time window provided in an embodiment of the present application. Based on the above embodiment, the method includes:
[0065] S201, obtaining a time slot width corresponding to a previous sliding time window, and obtaining window accumulated fault information corresponding to the previous sliding time window;
[0066] S202, calculating and obtaining the error density corresponding to the previous sliding time window based on the accumulated fault information in the window;
[0067] S203, obtaining a time slot width corresponding to the current sliding time window according to the error density, the time slot width, and a preset time slot width sequence;
[0068] S204: Determine the start time and end time of the current sliding time window based on the time slot width corresponding to the current sliding time window, and complete the determination of the current sliding time window.
[0069] In this embodiment, in order to ensure the dynamic adaptability of the storm triggering conditions, the time slot width of the sliding time window is adjusted to match the occurrence density of fault events in the current system. First, the time slot width corresponding to the previous sliding time window is obtained. The time slot width is the unit time length that divides the entire time window into several continuous time periods, which is used for the subsequent time attribution and classification statistics of fault events. At the same time, the window cumulative fault information corresponding to the previous sliding time window is also obtained, that is, the various types of fault event data collected and summarized within the time window. Based on the fault information, the error density is further calculated. The error density refers to the number of fault events occurring per unit time, reflecting the activity level of the current fault. The process of obtaining the error density usually involves statistics of all fault events in the time window, and normalization is performed in combination with the total length of the time window to obtain a density value with time correlation.
[0070] Subsequently, based on the acquired error density, time slot width and preset time slot width sequence, determine whether the time slot width corresponding to the current sliding time window needs to be adjusted. The preset time slot width sequence is a set of time lengths of multiple different granularities, which is used to achieve dynamic granularity control. When the error density is high, a smaller time slot granularity can be used to improve event response accuracy; when the error density is low, a larger time slot granularity can be used to reduce system load. Finally, based on the time slot width corresponding to the current sliding time window, the start and end times of the current sliding time window are determined. This process needs to refer to the end time of the last time window, combine the new slot width to calculate the new time window boundary, and complete the update of the sliding time window. Through this embodiment, dynamic adjustment of the fault event monitoring range in the time dimension is achieved, providing a more accurate time window basis for subsequent storm judgment and suppression mechanisms.
[0071] In a specific embodiment, a specific implementation of step S23 is provided. Based on the above embodiment, the time slot width corresponding to the current sliding time window is obtained according to the error density, the time slot width, and a preset time slot width sequence, including:
[0072] S2301, obtaining a preset time slot width sequence; wherein the time slot width sequence includes a plurality of slot width granularities arranged in descending order;
[0073] S2302: If the error density is greater than or equal to a preset first density threshold, and the time slot width is not the minimum slot width granularity in the time slot width sequence, select the next slot width granularity corresponding to the time slot width from the time slot width sequence, and record the next slot width granularity as the time slot width corresponding to the current sliding time window;
[0074] S2303: If the error density is greater than the preset second density threshold and less than the preset first density threshold, set the time slot width to the time slot width corresponding to the current sliding time window;
[0075] S2304: If the error density is less than or equal to the preset second density threshold, and the time slot width is not the maximum slot width granularity in the time slot width sequence, then select the previous slot width granularity corresponding to the time slot width from the time slot width sequence, and record the previous slot width granularity as the time slot width corresponding to the current sliding time window.
[0076] In this embodiment, in order to achieve dynamic adjustment of the time slot granularity under different fault density levels to improve the accuracy of fault detection and the adaptability of the response, the time slot width used by the current sliding time window is flexibly configured through a preset time slot width sequence. First, a preset time slot width sequence is obtained, which is a set of slot width granularities with different time lengths and is arranged in order from large to small. This sequence is intended to select appropriate time granularity for window division when the density of fault events changes, thereby ensuring detection sensitivity while taking into account the overall computational load. On this basis, the error density corresponding to the current time window is compared with two preset density thresholds.
[0077] If the current error density is greater than or equal to the first density threshold, and the currently used time slot width is not the minimum slot width granularity listed in the time slot width sequence, the current fault activity is considered high, and the monitoring granularity needs to be further refined to improve the timeliness of error event location and response. Therefore, the current time slot width is moved back one position in the sequence, and the corresponding next slot width granularity is selected as the new time slot width for the current sliding time window.
[0078] If the current error density is between the first and second density thresholds, the density level is considered intermediate, neither abnormal nor negligibly low. Therefore, the current time slot width is maintained to maintain overall operational stability and continuity. In this case, the time slot granularity is not reduced or increased; it remains at the granularity determined by the previous sliding time window.
[0079] If the current error density is less than or equal to the second density threshold, and the current time slot width is not the maximum slot width granularity listed in the time slot width sequence, the fault density is determined to be low, and the time granularity needs to be increased to reduce the system's statistical burden and the possibility of over-response to events. In this case, the previous slot width granularity (i.e., the larger slot width) in the time slot width sequence is selected as the new slot width for the current sliding time window.
[0080] This embodiment implements an adaptive update mechanism for time window granularity by determining and adjusting different fault density levels. This provides a more reasonable time division basis for subsequent window fault statistics and storm detection. This mechanism enhances the ability to capture fault outbreak trends while reducing unnecessary resource usage during inactive periods.
[0081] Furthermore, in another specific embodiment, determining the current sliding time window further includes:
[0082] S211, obtaining the time slot width corresponding to the previous sliding time window, and receiving the fault information of the hardware component in real time, and calculating the real-time error density of the hardware component after the end time of the previous sliding time window;
[0083] S212: If the real-time error density is greater than or equal to a preset third density threshold, and the time slot width is not the minimum slot width granularity in the time slot width sequence, then select the next slot width granularity corresponding to the time slot width from the time slot width sequence, and record the next slot width granularity as the time slot width corresponding to the current sliding time window;
[0084] S213, if the real-time error density is greater than a preset fourth density threshold and less than a preset third density threshold, setting the time slot width to the time slot width corresponding to the current sliding time window;
[0085] S214: If the real-time error density is less than or equal to a preset fourth density threshold, and the time slot width is not the maximum slot width granularity in the time slot width sequence, then select the previous slot width granularity corresponding to the time slot width from the time slot width sequence, and record the previous slot width granularity as the time slot width corresponding to the current sliding time window;
[0086] S215 , based on the time slot width corresponding to the current sliding time window, determine the start time and the end time of the current sliding time window, and complete the determination of the current sliding time window.
[0087] In this embodiment, as in the previous embodiment, the time slot width corresponding to the previous sliding time window is obtained. Subsequently, fault information reported by each hardware component is continuously received. Based on these real-time reported fault events, the real-time error density from the end of the previous sliding time window to the current moment is calculated. This real-time error density is calculated by dividing the number of errors per unit time by the length of the corresponding time interval, and is used to reflect the current level of fault activity.
[0088] If the statistically calculated real-time error density is greater than or equal to the preset third density threshold, and the currently used time slot width is not the smallest slot width granularity in the time slot width sequence (i.e., further refinement is still possible), the next slot width granularity corresponding to the current slot width is selected from the order of the time slot width sequence and set as the time slot width corresponding to the current sliding time window. By reducing the time slot granularity, sudden changes in errors can be captured within a shorter time scale, thereby enhancing the ability to identify storm errors early.
[0089] If the real-time error density is between the preset fourth density threshold and the third density threshold, the current time slot width is maintained unchanged. The existing time granularity can well adapt to the current fault frequency, and there is no need to adjust the granularity setting of the sliding time window.
[0090] If the real-time error density is less than or equal to the fourth density threshold, and the currently used time slot width is not the largest slot width granularity in the time slot width sequence (i.e., it can still be relaxed), the previous slot width granularity corresponding to the current slot width is selected from the time slot width sequence and set as the time slot width for the current sliding time window. By relaxing the time slot granularity, unnecessary fine-grained processing can be reduced in low-error activity scenarios, thereby reducing resource usage and improving monitoring efficiency.
[0091] Finally, based on the currently selected time slot width and the latest error density changes, the start and end times of the current sliding time window are determined, thereby completing the definition of the current sliding time window's boundaries. This embodiment enables adaptive adjustment of the sliding time window to the real-time error density of hardware components, effectively improving the sensitivity and stability of fault handling.
[0092] In one embodiment, a specific implementation method for obtaining historical fault information of at least one hardware component in step S11 is provided herein, including:
[0093] S111, receiving fault metadata of at least one hardware component sent by a basic input / output system;
[0094] S112, receiving and parsing the fault metadata, and storing the parsed fault metadata in a system event log; wherein the system event log stores historical fault information corresponding to at least one fault metadata;
[0095] S113, based on the historical fault information stored in the system event log, mapping the historical fault information to a preset fault monitoring buffer according to a preset time length of the sliding time window;
[0096] S114, based on the start and end times of the current sliding time window, filter and extract the historical fault data of each hardware component from the preset fault monitoring buffer; wherein the fault monitoring buffer uses timestamps as indexes to record the fault events of each hardware component at different time points and the fault types corresponding to the fault events.
[0097] In this embodiment, fault metadata of at least one hardware component transmitted by the basic input and output system is received. The fault metadata is usually automatically triggered and reported by the underlying firmware after an abnormality occurs in the hardware component, and includes fields such as the timestamp of the fault occurrence, the fault type identifier, and the fault source component identifier. After receiving the fault metadata, a parsing operation is performed to extract fields and normalize the format of the original fault information, and the normalized fault metadata structure is stored in the system event log. As a unified error event storage medium, the system event log can continuously record the historical fault information of multiple hardware components, ensuring that subsequent analysis can perform judgments based on the complete historical trajectory.
[0098] Subsequently, based on the preset duration of the sliding time window, historical fault information recorded in the system event log is mapped by timestamp and written to a preset fault monitoring buffer. This buffer is divided into multiple storage units with fixed time spans, each corresponding to a time slot. Fault events for different hardware components at various time points are categorized and recorded using the fault timestamp as an index. This mapping process ensures that fault data is accurately assigned to the corresponding time slot by comparing the event timestamps with the time boundaries of each slot in the sliding time window, facilitating subsequent time window statistical processing. Based on this, the data content within the corresponding time period in the fault monitoring buffer is retrieved based on the start and end times of the current sliding time window. Fault information for each hardware component occurring within that time period is filtered and extracted to form the historical fault data set for the current window. By constructing a time-indexed buffer structure and sliding window mapping mechanism, the fault history of hardware components can be organized and rapidly accessed, providing a controllable and traceable error data foundation for fault handling methods and improving the accuracy and efficiency of fault analysis.
[0099] For example, access the system event log. During a log query, a pre-set program execution command is executed to obtain the 26th event record. This record includes detailed fields such as the event number, record type, timestamp, event source device, sensor type, event type, and event direction. For example, the event's sensor type is A and the event description is B, indicating that the event corresponds to a correctable bus-level hardware error.
[0100] In a specific embodiment, an implementation of the above step S12 is provided. Based on the above embodiment, according to the historical fault information and the current system load information, the current storm determination condition corresponding to each hardware component is obtained, including:
[0101] S121, collecting and obtaining current system load information, and normalizing the current system load information to obtain a normalized value Load corresponding to the current load information;
[0102] S122, obtaining a historical correctable error density of the hardware component based on the current sliding event window; wherein the historical correctable error density is the correctable error density of the previous sliding event window; and / or obtaining a historical uncorrectable error density of the hardware component; wherein the historical uncorrectable error density is the uncorrectable error density of the previous sliding event window;
[0103] S123, using the formula:
[0104]
[0105] Calculate and obtain a first trigger density Threshold1, and obtain a correctable error storm trigger value of the hardware component based on the first trigger density and a preset time window length; wherein Base1 is a preset correctable error base threshold, α1 is a preset first load adjustment coefficient, β1 is a preset first error density adjustment coefficient, D1 is a historical correctable error density, and k1 is a preset first historical impact sensitivity; and / or
[0106] Calculate and obtain the second trigger density Threshold2; and obtain the uncorrectable error storm trigger value of the hardware component based on the second trigger density and the preset time window length; wherein Base2 is the preset uncorrectable error basic threshold, α2 is the preset second load adjustment coefficient, β2 is the preset second error density adjustment coefficient, D2 is the historical correctable error density, and k2 is the preset second historical impact sensitivity.
[0107] In this embodiment, current system load information is collected and acquired. System load information typically includes, but is not limited to, parameters such as current processor utilization, memory usage, I / O bandwidth usage, and task queue length. To facilitate use in subsequent algorithms, the raw system load information must be processed to generate a normalized value with a unified dimension. For example, any load information (such as processor utilization) can be selected and its corresponding normalized value calculated as the normalized value corresponding to the current load information. Alternatively, multiple load information (such as processor utilization and memory usage) can be selected, and normalized values calculated for each load information. Then, a weighted sum is performed based on the weights corresponding to each load information to obtain the normalized value corresponding to the current load information. It should be noted that normalization can be performed using existing technical means and is not specifically limited here. Based on the obtained normalized current load value, the historical correctable error density and / or historical uncorrectable error density of each hardware component is further obtained based on the current sliding time window. Historical error density represents the cumulative frequency of errors per unit time and is a key metric for measuring the frequency of error occurrence. Correctable error density reflects the disturbance intensity of non-fatal faults in the system, while uncorrectable error density focuses more on the level of system reliability.
[0108] Based on system load information and historical error density, a predefined nonlinear combination function is used to calculate the correctable and uncorrectable error storm trigger values for each hardware component. The formula includes three key variables: a base threshold term, representing the triggering criteria under ideal load and historically clean conditions (typically set between 3 and 10 errors per second for correctable errors); a load adjustment term, which dynamically reflects the effect of the current load on the trigger threshold; and a historical impact term, which adjusts the response strength by introducing error density and using a historical impact sensitivity parameter. This combination function takes into account both the current load level and historical error behavior, enabling adaptive adjustment of the storm detection threshold as system conditions change, reducing false positives or missed detections caused by static thresholds. This embodiment introduces two dynamic factors—normalized system load and historical error density—combined with a predefined adjustment coefficient and a nonlinear adjustment formula to construct a storm trigger detection criterion that automatically adjusts to the system's operating status. This achieves a sensitive response to abnormal fluctuations in hardware components and effectively enhances the adaptability and accuracy of the error storm detection mechanism.
[0109] In a specific embodiment, a specific implementation method is provided for obtaining the window cumulative fault information of the hardware component in the above step S13. Based on the above embodiment, it includes:
[0110] S131, obtaining the start time and end time of the current sliding time window, and the time slot width corresponding to the current sliding time window; wherein the time slot width is used to divide the current sliding time window into multiple continuous non-overlapping time slots;
[0111] S132, extracting historical fault information of one or more hardware components from a preset fault monitoring buffer according to the start time and the end time of the current sliding time window;
[0112] S133, mapping the historical fault information to each time slot in the current sliding time window according to the occurrence timestamp corresponding to the historical fault information, and mapping the collected real-time fault information to the corresponding time slot;
[0113] S134: Based on the fault type of the historical fault information, classify and count the fault information of each hardware component to obtain the cumulative number of correctable errors and / or the cumulative number of uncorrectable errors of at least one hardware component in the current sliding time window.
[0114] In this embodiment, the start and end times of the current sliding time window are obtained, and the corresponding time slot width is determined. The sliding time window represents the time range currently used to analyze hardware fault behavior. Its start and end times together define the window's time span. Setting the time slot width allows for granular temporal location and aggregated statistics of fault events. Subsequently, based on the start and end times of the current sliding time window, historical fault information for the target hardware component is extracted from a preset fault monitoring buffer. The fault monitoring buffer is used to cache and store fault records reported by multiple hardware components during operation in timestamp order. Each piece of historical fault information contains fields such as the timestamp of the fault occurrence, the corresponding hardware component identifier, and the fault type identifier. After extracting the historical fault information, the timestamp corresponding to each piece of information is used to determine whether it falls within the current sliding time window. The historical fault information that meets the criteria is then mapped to the corresponding time slot within the current time window. Furthermore, to ensure real-time statistics, newly collected real-time fault information within the current cycle is also mapped to the corresponding time slot, thereby ensuring the integrity and accuracy of the fault data within the time window.
[0115] After completing the fault information mapping at the time slot level, it is necessary to further classify and count the fault information of each hardware component in each time slot based on the fault type identification. Common fault types may include correctable errors (such as memory error correction code check errors) and uncorrectable errors (such as fatal failures of the central processing unit). The results of the classification counting will be summarized as the cumulative number of correctable errors and the cumulative number of uncorrectable errors of each hardware component in the current sliding time window, which will be used for subsequent storm judgment. This embodiment realizes the structured management and statistical aggregation of hardware fault events by introducing the error mapping and classification counting mechanism at the time slot granularity, providing accurate data support for error density analysis and storm trigger judgment within the sliding window, thereby improving the response efficiency and judgment reliability when facing high-frequency error events.
[0116] Furthermore, in a specific embodiment, in the above step S13, the accumulated fault information of the hardware component window meets the corresponding current storm judgment condition, and a specific judgment form is provided here. Based on the above embodiment, it includes:
[0117] S135, obtaining the cumulative number of correctable errors corresponding to the hardware component and the correctable error storm trigger value corresponding to the hardware component based on the window cumulative fault information of the hardware component;
[0118] S136, if the cumulative number of correctable errors is greater than or equal to the correctable error storm trigger value, determining that the window cumulative fault information of the hardware component meets its corresponding current storm determination condition;
[0119] S137, obtaining the cumulative number of uncorrectable errors corresponding to the hardware component, and obtaining the uncorrectable error storm trigger value corresponding to the hardware component;
[0120] S138: If the cumulative number of uncorrectable errors is greater than or equal to the uncorrectable error storm trigger value, determine whether the window cumulative fault information of the hardware component meets its corresponding current storm determination condition.
[0121] In this embodiment, based on the window cumulative fault information obtained through previous statistics, the cumulative number of correctable errors and the cumulative number of uncorrectable errors of each hardware component within the current sliding time window are obtained respectively. The window cumulative fault information is the result of aggregating and statistically analyzing the historical fault data collected and classified within the current sliding time window. Correctable errors generally refer to non-fatal errors that occur at the hardware level but can be recovered through an automatic repair mechanism. They do not immediately affect the stability of the system, but their frequent occurrence may indicate potential hardware degradation. Uncorrectable errors generally indicate more serious faults that occur during the operation of the equipment and cannot be repaired by self-recovery means.
[0122] Next, obtain the correctable error storm trigger value and the uncorrectable error storm trigger value of each current hardware component. Compare the cumulative number of correctable errors recorded in the window cumulative fault information with its corresponding correctable error storm trigger value. If the former is greater than or equal to the latter, it indicates that the frequency of correctable errors of the hardware component in the current sliding time window has reached the storm judgment threshold, and it is determined to meet its current storm judgment condition. Similarly, if the cumulative number of uncorrectable errors of the component in the current sliding time window is greater than or equal to its corresponding uncorrectable error storm trigger value, it can also be determined to meet the storm judgment condition. This embodiment realizes the judgment process of whether the error behavior of the hardware component meets the error storm judgment standard by setting a storm trigger threshold based on dynamic calculation and comparing the fault aggregation statistics within the sliding time window. Not only does it take into account the severity of the error type, but it also integrates the dynamic adjustment of the system operation status and historical characteristics, improves the accuracy and scenario adaptability of the fault storm judgment, and thus effectively supports the triggering and execution of the suppression strategy in subsequent fault processing.
[0123] In one embodiment, the fault handling method further includes:
[0124] S15: If the cumulative number of correctable errors of the storm triggering component is less than or equal to the preset number threshold within the preset sliding waiting time window, the storm triggering flag of the storm triggering component is canceled.
[0125] In this embodiment, a preset sliding wait time window is set, within which the error behavior of each hardware component with the storm trigger flag set is continuously monitored. During the sliding wait time window, the cumulative number of correctable errors of the storm trigger component is updated at set time intervals, and these values are compared with a preset number threshold. The number threshold reflects the tolerance range for slight error fluctuations under stable operating conditions. If the cumulative number of correctable errors of the storm trigger component is less than or equal to the number threshold during the entire sliding wait time window, it indicates that the error behavior of the hardware component has not continued to deteriorate or has stabilized.
[0126] When the above judgment conditions are met, it is determined that the hardware component no longer has the risk of continuously triggering a storm, and therefore the previously set storm trigger state will be revoked. Specifically, the setting of the storm trigger flag corresponding to the hardware component will be canceled, so that the reporting status of the hardware component returns to normal, and it will no longer be the target component in the subsequent suppression strategy. The revocation of this flag reduces unnecessary suppression measures for the component. This embodiment introduces a sliding waiting time window to implement a dynamic exit mechanism when the erroneous behavior is alleviated, thereby improving the stability of the entire error storm judgment and suppression. In addition, the mechanism also improves the adaptability to the natural fluctuations of correctable errors, ensuring sufficient fault tolerance and time elasticity in the storm identification process.
[0127] In a specific embodiment, a specific triggering embodiment for the flag bit setting condition is provided for the above step S14, that is, multiple suppression triggering conditions are preset, and different levels of error flag bits are triggered according to different conditions. Based on the above embodiment, it includes:
[0128] S141, within a first preset time range, if the cumulative number of correctable errors of the storm triggering component is greater than or equal to a first threshold, and the corresponding correctable error density is less than the first density threshold, then setting the observation level flag of the storm triggering component to a bit.
[0129] In this embodiment, the cumulative number of correctable errors, such as single-bit memory errors and transmission error retries, for all storm-triggering components is continuously monitored. A first preset time range is set to limit the observation period for error behavior analysis. For example, this time range can be set from several minutes to ten minutes. Within this time range, correctable error data for the target storm-triggering component is acquired, the cumulative number of correctable errors within the time range is calculated, and the error density within this time period is calculated. If the cumulative number of correctable errors for the storm-triggering component is greater than or equal to a preset first threshold, it indicates that the component has reached a certain level of failures during the observation period. At the same time, the correctable error density for the component is still below the first density threshold. This indicates that although the cumulative number is high, the error behavior is relatively evenly distributed and has not yet exhibited the typical characteristics of a high-density error storm, indicating that overall operation has not been severely disrupted. For example, if the cumulative number of correctable errors for the same storm-triggering component within 10 minutes is greater than or equal to 100, and the corresponding correctable error density is less than 0.5 errors / second, the observation level flag for the storm-triggering component is set if all of the above conditions are met. The observation-level flag, the lowest level in the hierarchical suppression process, indicates that the component is currently in a suspicious state and has not yet triggered actual suppression measures. By establishing this observation-level flag, potential issues can be identified and categorized for early management, helping to provide trend warnings before a potential problem breaks out and improving the accuracy of fault prediction.
[0130] S142: If the cumulative number of correctable errors of the storm triggering component is greater than or equal to a second threshold within a second preset time range, setting a warning level flag of the storm triggering component to a position.
[0131] In this embodiment, the correctable error behavior of each storm-triggering component is continuously tracked, and a second preset time range is set. Unlike the observation level, the warning level focuses on the absolute increase in the number of errors, rather than whether the error density changes dramatically. Within the second preset time range, correctable error data for the target component is obtained, and the total number of correctable errors occurring in the component during that time period is counted. If the cumulative number is greater than or equal to a second threshold preset by the system, the component is determined to be experiencing frequent errors, indicating that its hardware performance or electrical characteristics may be continuously degrading, and there is a risk of functional failure in the short to medium term. After confirming that the cumulative number of errors has reached the second threshold, the warning level flag for the storm-triggering component is set. For example, if the cumulative number of correctable errors for the same storm-triggering component within one minute is greater than or equal to 50, the corresponding suppression strategy for the warning level flag is generally more stringent than that for the observation level. For example, it may trigger local performance throttling, load shifting recommendations, or initiate detailed hardware fault diagnosis. This embodiment enhances the speed of response to fault trends and the accuracy of handling them, while ensuring system availability and reducing the cascading risks of hardware failures.
[0132] S143: When receiving the uncorrectable error information of the storm triggering component, setting the warning level flag of the storm triggering component.
[0133] In this embodiment, a higher-priority response strategy is implemented for uncorrectable errors to ensure that the source of the fault can be identified and controlled immediately when a serious hardware fault signal occurs. Uncorrectable errors typically indicate that some hardware component functionality has been compromised or that the error cannot be recovered through software. If not promptly responded to, it may cause system crashes, data loss, or hardware damage. To this end, upon receiving uncorrectable error information from a storm-triggering component, the warning level flag for the storm-triggering component is immediately set, without waiting for cumulative analysis. By setting the warning level flag, the hardware component can be included in the risk management mechanism as soon as possible. For example, once this flag is triggered, it can simultaneously notify the operating system to execute emergency measures such as device isolation, driver uninstallation, or service migration, while also generating a high-priority alarm entry in the system event log for subsequent review by operations and maintenance personnel. By strengthening the response capability for uncorrectable errors, this embodiment enables the rapid location and intervention of serious errors in the fault storm handling system, effectively reducing the spread of hardware damage or service interruptions caused by delayed responses, thereby ensuring continuous availability and fault tolerance in complex load environments.
[0134] S144: If the cumulative number of uncorrectable errors of the storm triggering component is greater than or equal to a third threshold within a third preset time range, setting the isolation level flag of the storm triggering component.
[0135] In this embodiment, within a preset third preset time range, each hardware component that has been marked as a storm trigger component is continuously monitored and counted for uncorrectable error events. If the cumulative number of uncorrectable errors is greater than or equal to the set third threshold, the isolation level flag of the hardware component is immediately set. For example, the cumulative number of uncorrectable errors of the same storm trigger component within 10 seconds is greater than or equal to 3 times. Uncorrectable errors usually include bus communication interruption, device self-test failure and other faults that cannot be automatically repaired at the software level. Their repeated occurrence often indicates that the component is in a state of severe degradation or is about to fail. By setting the third preset time range, it is ensured that the monitoring window has sufficient granularity to capture high-frequency error outbreaks, while not misjudging it as a serious fault due to short-term abnormal fluctuations. Once the cumulative number exceeds the set threshold, the operating status of the hardware component will be upgraded to the isolation level through the flag mechanism.
[0136] S145. Within a fourth preset time range, record the number of storm-triggering components that receive a number of correctable errors greater than or equal to a fourth threshold. If the number is greater than or equal to the first number threshold, set the isolation level flag corresponding to one or more storm-triggering components to a position.
[0137] In this embodiment, horizontal correlation analysis is performed based on group failure behavior to identify potential local failure outbreaks or common cause failure spread. Within a fourth preset time range, all hardware components in a storm-triggered state are continuously monitored and recorded, with a focus on screening storm-triggered components whose cumulative number of correctable errors is greater than or equal to a fourth threshold, and their number is counted. If this number reaches or exceeds the first threshold, it is determined that there may be multiple hardware components in the current system that are abnormally active at the same time in a similar time period, which may cause problems such as resource contention, interrupt storms, or decreased system stability.
[0138] Exemplarily, the statistical process employs a sliding time window mechanism and, through a pre-set event aggregation strategy, uniformly analyzes the abnormal behavior of multiple components over time. This ensures that not only individual abnormally active components are identified, but also trends in overall abnormal activity can be detected. The fourth threshold limits the minimum error count required for each component to ensure that the components being counted are indeed in a persistently abnormal state. The first threshold defines the lower limit at which the system deems the number of abnormal components to have reached a threshold requiring systemic intervention. For example, the cumulative number of correctable errors (i.e., the number of correctable errors received) for three storm-triggering components within one minute is greater than or equal to 100. Once the aforementioned aggregated abnormality triggering condition is met, a unified suppression action is performed on these storm-triggering components that have met the error count threshold. This action is accomplished by setting the isolation level flag for the corresponding component to activate the corresponding isolation mechanism, such as shutting down the corresponding channel, suspending the component's task scheduling, restricting access priority, or removing it from the running resource pool. This embodiment, through a dual assessment based on component number and error accumulation, not only enhances the ability to identify collective error storms but also improves the coordination of fault response and overall fault tolerance.
[0139] S146 , within a fifth preset time range, recording the number of storm triggering components receiving uncorrectable errors, and if the number is greater than or equal to a second number threshold, setting severity level indicators of multiple storm triggering components to bits.
[0140] In this embodiment, within the fifth preset time range, the fault reporting information from each storm triggering component is continuously monitored, and at the same time, the cumulative number of storm triggering components that have experienced uncorrectable errors within the current time range is counted. If the number reaches or exceeds the second quantity threshold, it means that serious failure events have occurred simultaneously in multiple key components within a short period of time. For example, greater than or equal to 3 storm triggering components trigger uncorrectable errors within 30 seconds. This failure mode often has systemic risks and may cause partial business interruption, service quality degradation, or affect the execution of high availability strategies. To this end, based on the suppression strategy framework, the severity level flags of all these storm triggering components are set to trigger the corresponding level of suppression action. By using uncorrectable errors as the judgment basis and the number of faulty components as the measurement standard, a unified identification and marking mechanism for multi-point serious hardware failures is constructed, which achieves rapid response and control of the system's serious risk status and enhances the accuracy and comprehensiveness of fault handling.
[0141] S147, if the joint density of correctable errors and uncorrectable errors of multiple storm triggering components per unit time is greater than a preset joint density threshold, and the physical topological structures of the multiple storm triggering components are associated, then the severity level flags corresponding to the multiple storm triggering components are set to positions.
[0142] In this embodiment, by setting the joint density of correctable errors and uncorrectable errors as a composite indicator, and combining the physical topology between multiple storm triggering components, a determination mechanism for the trend of fault aggregation in local hardware areas is established. Specifically, first, within a unit time window, the frequency of correctable errors and uncorrectable errors of each storm triggering component is counted in real time, and the total occurrence of the two types of errors is normalized to a unified error density value. In order to determine whether there is a local fault area with a risk of collaborative failure, the error density values of multiple storm triggering components are accumulated and their joint density is calculated. The joint density is the normalized value of the total frequency of correctable errors and uncorrectable errors reported by all storm triggering components in the area within a unit time, and is a parameter reflecting the intensity of local error outbreaks. If the joint density is greater than the preset joint density threshold, it indicates that the area is currently in a high-intensity error outbreak state.
[0143] On this basis, the deployment relationship of these storm-triggering components in the physical space is further analyzed to determine whether there is a direct or indirect association in the physical topology. For example, these components may be concentrated in the same cabinet, the same motherboard or the same backplane connection link, and have a common power supply or signal path. The above-mentioned topological relationship is determined by reading the component topology information, hardware layout map or association table. If the combined error density of multiple storm-triggering components not only exceeds the threshold, but also there is a structural association in the physical topology, the severity level flags of all related storm-triggering components are set. For example, the combined density of correctable errors and uncorrectable errors is greater than 5 seconds / time, and meets the topological correlation. This embodiment realizes the effective perception and timely response to the outbreak of high-risk regional faults through the dual criteria of combined error density and topological correlation, and strengthens the active defense capability of fault handling in complex operating environments.
[0144] Figure 3 Schematic diagram of the troubleshooting method provided in this application embodiment Figure 2 .like Figure 3 As shown, including:
[0145] S31, triggering an interrupt when at least one hardware component fails, and obtaining failure information of the at least one hardware component.
[0146] In this embodiment, an interrupt is triggered when at least one hardware component fails. An interrupt mechanism is an asynchronous signal proactively triggered by hardware and used to promptly respond to critical events, particularly in reliability, availability, and serviceability architectures. By triggering an interrupt, exception handling can be initiated immediately upon the occurrence of a hardware failure, thereby improving the real-time and accuracy of fault response. When an interrupt is triggered, parameter information corresponding to the fault event is collected. This fault information typically includes the timestamp of the fault, the fault type (e.g., correctable or uncorrectable), the hardware identifier of the fault source, the fault severity, the sensor number, and the raw data of the error event. This acquired fault information serves as the basis for subsequent fault analysis and judgment, providing data for further determining whether an error storm has occurred. This embodiment, through the use of an interrupt-based fault triggering mechanism, enables proactive reporting of hardware component fault information, reducing the latency and resource waste associated with traditional polling mechanisms. Interrupt responses can detect fault conditions with extremely low latency and provide the data foundation for subsequent multi-level judgment mechanisms based on historical information and system load status, ensuring efficient and stable fault response capabilities.
[0147] S32: Send fault information of at least one hardware component to the baseboard control manager.
[0148] So that the baseboard control manager receives the fault information of at least one hardware component, and obtains the historical fault information of at least one hardware component based on the stored fault information; obtains the current system load information, and obtains the current storm judgment condition corresponding to each hardware component based on the historical fault information and the current system load information; obtains the window cumulative fault information of at least one hardware component according to the current sliding time window, and sets the storm trigger flag of the hardware component when the window cumulative fault information of the hardware component meets its corresponding current storm judgment condition, and records the hardware component as a storm trigger component; receives the fault information of at least one storm trigger component, and sets the suppression flag of at least one storm trigger component when the fault information of at least one storm trigger component meets any preset suppression trigger condition.
[0149] In this embodiment, fault information about at least one hardware component is sent to the baseboard management controller (BMC). As a controller within the server platform, the BMC monitors, manages, and controls the underlying hardware status. It performs logical processing on the hardware status and maintains normal operation even when the operating system fails or the main processor experiences an anomaly. By reporting fault information to the BMC, the BMC ensures continuous tracking and analysis of hardware component faults even in the event of system-level anomalies.
[0150] After receiving the fault information of at least one hardware component, the baseboard control manager determines whether the hardware component meets the storm triggering condition and the suppression triggering condition. The corresponding description has been made in the above embodiment and will not be repeated here.
[0151] Figure 4 Schematic diagram of the process of the troubleshooting method provided in the embodiment of the application Figure 3 On the basis of the above embodiments, Figure 4 As shown, it also includes:
[0152] S41, monitoring the storm flag in real time, and after reading the storm flag of the storm triggering component, shielding the fault information of the storm triggering component from being transmitted to the upper layer.
[0153] In this embodiment, by real-time monitoring of the storm flag status, it is continuously detected whether each hardware component is in a storm-triggered state. The storm flag is a logical identification bit set for each hardware component and is typically dynamically updated by a management control module (e.g., a baseboard control manager) based on fault statistics and judgment logic. When a hardware component's fault behavior is detected to meet its storm judgment criteria, the flag is set to an active state, indicating that the component is in a fault storm state. To prevent the large amount of error information continuously generated during a storm state from disrupting the system, upon detecting the storm flag being set, the fault information of the storm-triggered component is immediately shielded from upstream transmission. Specifically, the fault information's transmission path to the upper layer is logically interrupted or marked as ignored. This shielding can be achieved by modifying relevant RAS registers, for example, to ensure that subsequent errors generated by the component are not transmitted to the upper layer for a certain period of time. This shielding operation effectively reduces the chain reactions caused by error storms, such as load surges, redundant fault event recording, and false error response triggering, thereby improving overall operational stability and the reliability of exception handling. Furthermore, this shielding mechanism does not affect the BMC's continuous monitoring of the component's status; fault information for the component remains available in the management channel during the shielding period. This embodiment achieves rapid suppression of fault storm behavior and interference isolation by reading the storm flag in real time and dynamically controlling the flow of error information.
[0154] S42, monitoring the suppression flag in real time, and after reading the suppression flag of the storm triggering component, executing the suppression strategy corresponding to the suppression flag.
[0155] In this embodiment, the suppression flag of each hardware component is continuously monitored in real time. The suppression flag is a status identifier set for the corresponding hardware component after the fault behavior of the component is judged against the preset suppression trigger conditions, and is used to indicate whether the component is currently in a state where suppression measures need to be taken. If the suppression flag of any hardware component is detected to be in position during the monitoring process, it indicates that the current fault behavior of the storm-triggered component has met the risk conditions of the corresponding level, and the preset suppression strategy corresponding to the level of the suppression flag will be executed. Suppression flags of different levels will trigger suppression actions of different intensities and ranges to ensure that the suppression strategy is targeted. This embodiment achieves rapid intervention in faults through real-time reading and response to the suppression flag, thereby improving the overall stability and automation level of error handling.
[0156] In a specific embodiment, a specific suppression strategy is provided for the above step S42. Based on the above embodiment, it includes:
[0157] S421, when the observation level flag of the storm triggering component is read, the fault information of the storm triggering component is recorded in the system event log and a health report is output;
[0158] S422, when the warning level flag of the storm trigger component is read, the execution frequency of the subsystem corresponding to the storm trigger component is reduced, and an alarm notification carrying information about the storm trigger component is output;
[0159] S423, when the isolation level flag of the storm triggering component is read and it is set, the storm triggering component is isolated; or,
[0160] S424 , when the severity flag of the storm triggering component is read and is set, the system-level operation mode is downgraded, and a health check instruction is output to the baseboard management controller.
[0161] In this embodiment, by reading the different levels of flags of each storm-triggering component in real time, fine-grained perception and response control of component fault states is achieved. When the observation level, warning level, isolation level, or severe level flag of a storm-triggering component is set, the corresponding fault tolerance or suppression action is taken according to the preset policy corresponding to that level.
[0162] Specifically, when a storm-triggered component's observation-level flag is set, it indicates that the component has experienced a certain frequency of correctable errors. At this point, relevant fault information (including error type, timestamp, component ID, and error density) is recorded in the system event log for fault analysis and trend tracking. A health report is also generated for reference by system administrators or fault management services.
[0163] When the warning level flag for a storm-triggering component is set, it indicates that the component's erroneous behavior is clearly abnormal and has the potential to deteriorate into an uncorrectable error or trigger a chain reaction. This triggers targeted action mechanisms, such as proactively reducing the execution frequency of the subsystem to which the storm-triggering component belongs (e.g., memory subsystem access rate, link bandwidth, etc.) to prevent further spread of the fault. Furthermore, an alert notification is issued containing the component's identification, fault level, and recommended action, prompting manual intervention or remote diagnosis.
[0164] When the isolation level flag of a storm-triggering component is detected, it indicates that the component may no longer be able to maintain normal functionality or is continuously generating serious errors. To reduce the spread of system-level impact, physical or logical isolation operations will be immediately performed on the storm-triggering component. Such operations may include disconnecting the component from the backbone data path, shutting down the corresponding device port, stopping service allocation, or dynamically redirecting tasks to redundant components to ensure fault isolation with minimal loss.
[0165] When the severity flag of a storm-triggering component is detected, it indicates that multiple components in the system are simultaneously experiencing high-density faults and have correlated or synchronous error characteristics. At this point, a system-level operating mode downgrade will be executed, such as switching to a safe operating mode, limiting the number of concurrent tasks, or suspending some services to quickly stabilize the system operating environment. At the same time, a health check instruction will be output to the baseboard management controller, instructing the baseboard management controller to initiate a system-wide component status inspection, temperature and power consumption verification, or remote diagnosis tasks, providing basic information for subsequent fault tracing and repair decisions.
[0166] This embodiment constructs multi-level, fine-grained error handling through graded responses to flag bits of different levels, realizes full-process linkage control from primary perception to high-level intervention, and effectively improves controllability and stability in the face of complex fault scenarios.
[0167] Figure 5 This is a timing diagram of the fault handling method provided in the embodiment of the present application. Figure 5 As shown, including:
[0168] S51, the basic input and output system triggers an interrupt when at least one hardware component generates a fault, and obtains fault information of the at least one hardware component;
[0169] S52, the basic input and output system sends fault information of at least one hardware component;
[0170] S53, the baseboard management controller receives fault information of at least one hardware component, and obtains historical fault information of at least one hardware component based on the stored fault information;
[0171] S54, the baseboard management controller obtains current system load information, and obtains current storm determination conditions corresponding to each hardware component based on historical fault information and current system load information;
[0172] S55: The baseboard management controller obtains window accumulated fault information of at least one hardware component according to the current sliding time window, and sets a storm trigger flag of the hardware component when the window accumulated fault information of the hardware component meets its corresponding current storm determination condition, and records the hardware component as a storm trigger component;
[0173] S56, the basic input and output system monitors the storm flag in real time, and after reading the storm flag of the storm trigger component, blocks the fault information of the storm trigger component from being transmitted to the upper layer;
[0174] S57, the baseboard management controller receives fault information of at least one storm triggering component, and sets a suppression flag of at least one storm triggering component when the fault information of at least one storm triggering component meets any preset suppression trigger condition;
[0175] S58, the basic input and output system monitors the suppression flag in real time, and after reading that the suppression flag is set, executes the suppression strategy corresponding to the suppression flag.
[0176] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0177] Figure 6 This is a schematic diagram of the structure of the baseboard management controller provided in the embodiment of the present application. Figure 6 As shown, the baseboard management controller 6 includes:
[0178] The data receiving module 61 is configured to receive fault information of at least one hardware component sent by the basic input / output system, and obtain historical fault information of at least one hardware component based on the stored fault information; wherein the fault information is obtained by the basic input / output system when an interrupt is triggered by a hardware component fault;
[0179] Condition generation module 62, used to obtain current system load information and obtain current storm determination conditions corresponding to each hardware component based on historical fault information and current system load information;
[0180] Storm determination module 63 is configured to obtain cumulative window fault information of at least one hardware component according to the current sliding time window, and when the cumulative window fault information of the hardware component meets its corresponding current storm determination condition, set the storm trigger flag of the hardware component and record the hardware component as a storm trigger component;
[0181] The suppression monitoring module 64 is configured to receive fault information of at least one storm triggering component and set a suppression flag of at least one storm triggering component when the fault information of at least one storm triggering component meets any preset suppression triggering condition.
[0182] For the description of the features in the embodiment corresponding to the baseboard management controller, reference can be made to the relevant description of the embodiment corresponding to the above-mentioned fault handling method, which will not be repeated here.
[0183] Figure 7 This is a schematic diagram of the basic input and output system provided in the embodiment of the present application. Figure 6 As shown, the basic input and output system 7 includes:
[0184] An information acquisition module 71 is configured to trigger an interrupt when at least one hardware component fails, and obtain failure information of the at least one hardware component;
[0185] The information sending module 72 is used to send the fault information of at least one hardware component to the baseboard control manager, so that the baseboard control manager receives the fault information of at least one hardware component and obtains the historical fault information of at least one hardware component based on the stored fault information; obtains the current system load information, and obtains the current storm judgment condition corresponding to each hardware component based on the historical fault information and the current system load information; obtains the window cumulative fault information of at least one hardware component according to the current sliding time window, and sets the storm trigger flag of the hardware component when the window cumulative fault information of the hardware component meets its corresponding current storm judgment condition, and records the hardware component as a storm trigger component; receives the fault information of at least one storm trigger component, and sets the suppression flag of at least one storm trigger component when the fault information of at least one storm trigger component meets any preset suppression trigger condition.
[0186] For the description of the features in the embodiment corresponding to the basic input and output system, reference can be made to the relevant description of the embodiment corresponding to the above-mentioned fault handling method, which will not be repeated here.
[0187] The present application also provides a fault handling system, comprising: a baseboard management controller, a basic input and output system, and a processor;
[0188] The baseboard management controller includes a microprocessor and a management control memory, the management control memory is used to store a management control program, and the microprocessor is used to execute the management control program stored in the management control memory to perform the steps according to the corresponding fault handling method;
[0189] The basic input and output system is used to store basic input and output programs;
[0190] The processor is used to execute the basic input and output program stored in the basic input and output system to perform the steps according to the corresponding fault handling method.
[0191] Figure 8 This is a schematic diagram of the structure of the electronic device provided in this application. Figure 8 As shown, the electronic device 8 provided in this embodiment includes: at least one processor 81 and a memory 82. Optionally, the electronic device 8 also includes a communication component 83. The processor 81, the memory 82 and the communication component 83 are connected via a bus 84.
[0192] During the specific implementation process, at least one processor 81 executes the computer-executable instructions stored in the memory 82, so that the at least one processor 81 executes the above-mentioned fault handling method embodiment.
[0193] The specific implementation process of the processor 81 can be found in the above-mentioned method embodiment. Its implementation principle and technical effects are similar and will not be repeated here in this embodiment.
[0194] In the above embodiments, it should be understood that the processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the application may be directly executed by a hardware processor or by a combination of hardware and software modules within the processor.
[0195] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage.
[0196] A bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Buses can be categorized as address buses, data buses, and control buses. For ease of illustration, the buses in the drawings of this application are not limited to just one bus or just one type of bus.
[0197] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps of any of the above-mentioned fault handling method embodiments when running.
[0198] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0199] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above-mentioned fault handling method embodiments are implemented.
[0200] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above-mentioned fault handling method embodiments are implemented.
[0201] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0202] The above is a detailed introduction to a fault handling method provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core ideas of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.
Claims
1. A fault handling method, characterized in that: include: receiving fault information of at least one hardware component sent by a basic input / output system, and acquiring historical fault information of at least one hardware component based on the stored fault information; Wherein, the fault information is obtained by the basic input and output system when a hardware component fails and triggers an interrupt; Obtaining current system load information, and obtaining a current storm determination condition corresponding to each of the hardware components based on the historical fault information and the current system load information; Obtaining, according to the current sliding time window, window cumulative fault information of at least one of the hardware components, and when the window cumulative fault information of the hardware component meets its corresponding current storm determination condition, setting a storm trigger flag of the hardware component, and recording the hardware component as a storm trigger component; receiving fault information of at least one storm triggering component, and setting a suppression flag of the at least one storm triggering component when the fault information of the at least one storm triggering component meets any preset suppression trigger condition; The acquiring, according to the current sliding time window, window accumulated fault information of at least one of the hardware components includes: Obtaining the start time and end time of the current sliding time window, and the time slot width corresponding to the current sliding time window; wherein the time slot width is used to divide the current sliding time window into multiple continuous non-overlapping time slots; Extracting historical fault information of one or more hardware components from a preset fault monitoring buffer according to the start time and the end time of the current sliding time window; Mapping the historical fault information to each time slot within the current sliding time window according to the occurrence timestamp corresponding to the historical fault information, and mapping the collected real-time fault information to the corresponding time slot; Based on the fault type of the historical fault information, the fault information of each of the hardware components is classified and counted to obtain the cumulative number of correctable errors and / or the cumulative number of uncorrectable errors of at least one of the hardware components in the current sliding time window.
2. The fault handling method according to claim 1, characterized in that: The determination of the current sliding time window includes: Obtaining a time slot width corresponding to a previous sliding time window, and obtaining window accumulated fault information corresponding to the previous sliding time window; Calculate and obtain the error density corresponding to the previous sliding time window based on the accumulated fault information in the window; Obtaining a time slot width corresponding to the current sliding time window according to the error density, the time slot width, and a preset time slot width sequence; Based on the time slot width corresponding to the current sliding time window, the start time and the end time of the current sliding time window are determined, thereby completing the determination of the current sliding time window.
3. The fault handling method according to claim 2, characterized in that: The acquiring, according to the error density, the time slot width, and a preset time slot width sequence, a time slot width corresponding to the current sliding time window includes: Obtain a preset time slot width sequence; wherein the time slot width sequence includes a plurality of slot width granularities arranged in descending order; If the error density is greater than or equal to a preset first density threshold, and the time slot width is not the minimum slot width granularity in the time slot width sequence, then selecting the next slot width granularity corresponding to the time slot width from the time slot width sequence, and recording the next slot width granularity as the time slot width corresponding to the current sliding time window; or If the error density is greater than a preset second density threshold and less than a preset first density threshold, setting the time slot width to the time slot width corresponding to the current sliding time window; or, If the error density is less than or equal to a preset second density threshold, and the time slot width is not the maximum slot width granularity in the time slot width sequence, then the previous slot width granularity corresponding to the time slot width is selected from the time slot width sequence, and the previous slot width granularity is recorded as the time slot width corresponding to the current sliding time window.
4. The fault handling method according to claim 1, characterized in that: The receiving of the fault information of at least one hardware component sent by the basic input and output system and obtaining historical fault information of at least one hardware component based on the stored fault information includes: receiving fault metadata of at least one of the hardware components sent by the basic input and output system; Receive and parse the fault metadata, and store the parsed fault metadata in a system event log; wherein the system event log stores historical fault information corresponding to at least one fault metadata; Based on the historical fault information stored in the system event log, the historical fault information is mapped to a preset fault monitoring buffer according to a preset time length of the sliding time window; Filtering and extracting historical fault data of each of the hardware components from a preset fault monitoring buffer according to the start time and the end time of the current sliding time window; The fault monitoring buffer uses timestamps as indexes to record fault events occurring at different time points of each hardware component and the fault types corresponding to the fault events.
5. The fault handling method according to claim 1, characterized in that: The obtaining, based on the historical fault information and the current system load information, a current storm determination condition corresponding to each of the hardware components includes: Collect and obtain current system load information, and normalize the current system load information to obtain a normalized value Load corresponding to the current load information; Obtaining, based on the current sliding event window, a historical correctable error density of the hardware component; wherein the historical correctable error density is the correctable error density of the previous sliding event window; and / or obtaining a historical uncorrectable error density of the hardware component; wherein the historical uncorrectable error density is the uncorrectable error density of the previous sliding event window; Using the formula: Calculate and obtain a first trigger density Threshold1; and obtain a correctable error storm trigger value of the hardware component based on the first trigger density and the preset time window length; wherein Base1 is a preset correctable error base threshold, α1 is a preset first load adjustment coefficient, β1 is a preset first error density adjustment coefficient, D1 is a historical correctable error density, and k1 is a preset first historical impact sensitivity; and / or Calculate and obtain the second trigger density Threshold2; and obtain the uncorrectable error storm trigger value of the hardware component based on the second trigger density and the preset time window length; wherein Base2 is the preset uncorrectable error basic threshold, α2 is the preset second load adjustment coefficient, β2 is the preset second error density adjustment coefficient, D2 is the historical correctable error density, and k2 is the preset second historical impact sensitivity.
6. The fault handling method according to claim 1, characterized in that: The window accumulated fault information of the hardware component satisfies the corresponding current storm determination condition, including: Obtaining, according to the window cumulative fault information of the hardware component, the cumulative number of correctable errors corresponding to the hardware component and the correctable error storm trigger value corresponding to the hardware component; If the cumulative number of correctable errors is greater than or equal to the correctable error storm trigger value, then determining that the window cumulative fault information of the hardware component meets its corresponding current storm determination condition; or Obtaining a cumulative number of uncorrectable errors corresponding to the hardware component, and obtaining an uncorrectable error storm trigger value corresponding to the hardware component; If the cumulative number of uncorrectable errors is greater than or equal to the uncorrectable error storm trigger value, it is determined that the window cumulative fault information of the hardware component meets its corresponding current storm determination condition.
7. The fault handling method according to claim 1, further comprising: If the cumulative number of correctable errors of the storm triggering component is less than or equal to a preset number threshold within the preset sliding waiting time window, the setting of the storm triggering flag of the storm triggering component is canceled.
8. The method according to claim 1, characterized in that When the fault information of at least one storm triggering component meets any preset suppression trigger condition, setting the suppression flag of the at least one storm triggering component to a position includes: If, within a first preset time range, the cumulative number of correctable errors of the storm triggering component is greater than or equal to a first threshold, and the corresponding correctable error density is less than the first density threshold, the observation level flag of the storm triggering component is set to 1; or If the cumulative number of correctable errors of the storm triggering component is greater than or equal to a second threshold within a second preset time range, the warning level flag of the storm triggering component is set; or Upon receiving the uncorrectable error information of the storm triggering component, setting the warning level flag of the storm triggering component; or, If the cumulative number of uncorrectable errors of the storm triggering component is greater than or equal to a third threshold within a third preset time range, the isolation level flag of the storm triggering component is set; or Within a fourth preset time range, recording the number of storm-triggering components that receive a number of correctable errors greater than or equal to a fourth threshold, and if the number is greater than or equal to the first number threshold, setting the isolation level flag corresponding to one or more of the storm-triggering components to a position; or Within a fifth preset time range, recording the number of storm triggering components receiving uncorrectable errors, and if the number is greater than or equal to a second number threshold, setting severity level indicators of the plurality of storm triggering components to bits; If the joint density of correctable errors and uncorrectable errors of multiple storm triggering components per unit time is greater than a preset joint density threshold, and the physical topological structures of the multiple storm triggering components are associated, the severity level flags corresponding to the multiple storm triggering components are set to positions.
9. A fault handling method, characterized in that: include: triggering an interrupt when at least one hardware component fails, and obtaining failure information of the at least one hardware component; Sending the fault information of the at least one hardware component to the baseboard control manager, so that the baseboard control manager receives the fault information of the at least one hardware component, and obtains historical fault information of the at least one hardware component based on the stored fault information; obtains current system load information, and obtains current storm determination conditions corresponding to each of the hardware components based on the historical fault information and the current system load information; obtains window cumulative fault information of at least one of the hardware components based on the current sliding time window, and when the window cumulative fault information of the hardware component meets its corresponding current storm determination condition, sets the storm trigger flag of the hardware component and records the hardware component as a storm trigger component; receiving fault information of at least one storm triggering component, and setting a suppression flag of the at least one storm triggering component when the fault information of the at least one storm triggering component meets any preset suppression trigger condition; The acquiring, according to the current sliding time window, window accumulated fault information of at least one of the hardware components includes: Obtaining the start time and end time of the current sliding time window, and the time slot width corresponding to the current sliding time window; wherein the time slot width is used to divide the current sliding time window into multiple continuous non-overlapping time slots; Extracting historical fault information of one or more hardware components from a preset fault monitoring buffer according to the start time and the end time of the current sliding time window; Mapping the historical fault information to each time slot within the current sliding time window according to the occurrence timestamp corresponding to the historical fault information, and mapping the collected real-time fault information to the corresponding time slot; Based on the fault type of the historical fault information, the fault information of each of the hardware components is classified and counted to obtain the cumulative number of correctable errors and / or the cumulative number of uncorrectable errors of at least one of the hardware components in the current sliding time window.
10. The fault handling method according to claim 9, characterized in that: Also includes: Monitor the storm flag in real time, and after reading the storm flag of the storm triggering component, block the fault information of the storm triggering component from being transmitted to the upper layer; as well as The suppression flag is monitored in real time, and after reading that the suppression flag is set, a suppression strategy corresponding to the suppression flag is executed.
11. The fault handling method according to claim 10, characterized in that: The real-time monitoring of the suppression flag bit and, after reading that the suppression flag bit is set, executing the suppression strategy corresponding to the suppression flag bit includes: When the observation level flag of the storm triggering component is read, the fault information of the storm triggering component is recorded in the system event log and a health report is output; or When the warning level flag of the storm trigger component is read, the execution frequency of the subsystem corresponding to the storm trigger component is reduced, and an alarm notification carrying the information of the storm trigger component is output; or When the isolation level flag of the storm triggering component is read and it is set, isolating the storm triggering component; or, When the severity level flag of the storm triggering component is read and is set, the system-level operation mode is downgraded and a health check instruction is output to the baseboard management controller.
12. A fault handling method, characterized in that: include: The basic input and output system triggers an interrupt when at least one hardware component fails, and obtains failure information of the at least one hardware component; The basic input and output system sends a fault message of at least one hardware component; The baseboard management controller receives fault information of at least one hardware component and obtains historical fault information of at least one hardware component based on the stored fault information; The baseboard management controller obtains current system load information, and obtains current storm determination conditions corresponding to each of the hardware components according to the historical fault information and the current system load information; The baseboard management controller obtains the window accumulated fault information of at least one of the hardware components according to the current sliding time window, and sets the storm trigger flag of the hardware component when the window accumulated fault information of the hardware component meets the corresponding current storm determination condition, and records the hardware component as a storm trigger component; The basic input and output system monitors the storm flag in real time, and after reading the storm flag of the storm trigger component, blocks the fault information of the storm trigger component from being transmitted to the upper layer; The baseboard management controller receives fault information of at least one storm triggering component, and sets a suppression flag of the at least one storm triggering component when the fault information of the at least one storm triggering component meets any preset suppression trigger condition; The basic input and output system monitors the suppression flag in real time, and executes the suppression strategy corresponding to the suppression flag after reading that the suppression flag is set; The baseboard management controller obtains window accumulated fault information of at least one of the hardware components according to the current sliding time window, including: Obtaining the start time and end time of the current sliding time window, and the time slot width corresponding to the current sliding time window; wherein the time slot width is used to divide the current sliding time window into multiple continuous non-overlapping time slots; Extracting historical fault information of one or more hardware components from a preset fault monitoring buffer according to the start time and the end time of the current sliding time window; Mapping the historical fault information to each time slot within the current sliding time window according to the occurrence timestamp corresponding to the historical fault information, and mapping the collected real-time fault information to the corresponding time slot; Based on the fault type of the historical fault information, the fault information of each of the hardware components is classified and counted to obtain the cumulative number of correctable errors and / or the cumulative number of uncorrectable errors of at least one of the hardware components in the current sliding time window.
13. A fault handling system, characterized in that: include: baseboard management controller, basic input and output system, and processor; The baseboard management controller includes a microprocessor and a management control memory, wherein the management control memory is used to store a management control program, and the microprocessor is used to execute the management control program stored in the management control memory to perform the steps of the fault handling method according to any one of claims 1 to 8; The basic input and output system is used to store basic input and output programs; The processor is configured to execute the basic input / output program stored in the basic input / output system to perform the steps of the fault handling method according to any one of claims 9 to 11.
14. A computer-readable storage medium having a computer program stored therein, wherein: When the computer program is executed by a processor, the steps of the fault handling method according to any one of claims 1 to 8 or claims 9 to 11 are implemented.
Citation Information
Patent Citations
Method and device for processing network equipment alarm message storm
CN106656590A
Data processing method and device
CN119512801A