Fault management method and device, electronic equipment and computer program product

By managing high-bandwidth memory failures in a hierarchical manner, using BIOS and BMC to read and parse fault information, the problem of difficult HBM failures is solved, and the stability and reliability of the system are improved.

CN120336057APending Publication Date: 2025-07-18XFUSION DIGITAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510400281.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively manage high-bandwidth memory (HBM) faults, especially in complex and high-density packaging methods, troubleshooting is difficult and a perfect response mechanism is lacking.

Method used

A fault management method is provided, by determining the operating status of the operating system, and based on the basic input and output system (BIOS) and the fault management device (BMC), high bandwidth memory failures are managed in a graded manner, including reading and parsing fault information, recording minor faults and generating early warning information, and out-of-band collection and fault location in case of serious faults.

Benefits of technology

It realizes timely early warning and accurate positioning of high-bandwidth memory failures, improves the stability and reliability of the system, and ensures that the operating system can effectively deal with serious or fatal failures in abnormal states.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336057A_ABST
    Figure CN120336057A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computers, and particularly provides a fault management method and device, electronic equipment and a computer program product. The fault management method applied to the computing equipment comprises the following steps: in response to a high-bandwidth memory fault, determining a running state of an operating system; when the operation state meets the first state in response, first fault management is carried out based on received first fault information sent by the basic input and output system; and when the operation state meets the second state, performing second fault management. According to the method, comprehensive fault management is carried out on the operating system in a normal state and an abnormal state, and by recording and / or processing a plurality of correction problems and / or slightly uncorrected problems, timely early warning of faults and / or accurate positioning when serious or fatal uncorrected problems occur are / is realized, so that the stability and the reliability of the system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to a fault management method, apparatus, electronic device, and computer program product. Background Art

[0002] Generative artificial intelligence (AI) large models have extremely high requirements for data transmission speed. Currently, by integrating high bandwidth memory (HBM) into the processor package, the data transmission rate between the memory and the processor is improved to meet the requirements of AI large models.

[0003] Once a failure occurs in this complex and high-density packaging method, it is very difficult to troubleshoot, and currently, there is no perfect mechanism to deal with the occurrence of HBM failures. Summary of the Invention

[0004] In view of the above problems, the present disclosure is proposed. The present disclosure provides a fault management method, apparatus, electronic device, and computer program product.

[0005] According to one aspect of the present disclosure, there is provided a fault management method applied to a computing device. The method includes: in response to a high bandwidth memory failure, determining the operating state of the operating system; in response to the operating state satisfying a first state, performing first fault management based on the first fault information sent by the basic input / output system; and in response to the operating state satisfying a second state, performing second fault management.

[0006] In addition, in the fault management method according to one aspect of the present disclosure, performing the first fault management includes: reading and parsing the first fault information; based on a first preset rule, determining a first fault type of the parsed first fault information; and in the case where the first fault type satisfies a first type, recording the parsed first fault information and / or generating a warning message.

[0007] In addition, in the fault management method according to one aspect of the present disclosure, performing the second fault management includes: obtaining second fault information of a register; and generating a fault result and / or a warning message based on the second fault information, the first preset rule, and / or the parsed first fault information and / or a second preset rule.

[0008] In addition, a fault management method according to an aspect of the present disclosure, wherein a fault result including and / or a warning message is generated based on second fault information, a first preset rule, and / or the parsed first fault information and / or a second preset rule: reading and parsing the second fault information; determining a second fault type of the parsed second fault information based on the first preset rule; recording the parsed second fault information and / or generating a warning message when the second fault type meets the first type; and generating a fault result based on the parsed second fault information, the parsed first fault information, and / or the second preset rule when the second fault type does not meet the first type.

[0009] In addition, a fault management method according to an aspect of the present disclosure, wherein the method further includes: outputting the fault result and / or the warning message.

[0010] In addition, a fault management method according to an aspect of the present disclosure, wherein the first state includes: the operating system is running normally; the second state includes: the operating system is running abnormally; the first fault information includes: corrected errors and / or minor uncorrected errors; the second fault information includes: serious or fatal uncorrected errors; the first type includes: uncorrected errors that do not require any action.

[0011] In addition, a fault management method according to an aspect of the present disclosure, wherein the first fault information and the second fault information parameters include one or more of the following: system time log; time related to the fault information; source of the fault information; corresponding central processing unit socket information; corresponding central processing unit core information; corresponding central processing unit related sub-module information.

[0012] According to another aspect of the present disclosure, there is provided a fault management device, including: a determination module, configured to determine the operating state of the operating system in response to a high bandwidth memory fault; a management module, configured to perform first fault management based on the first fault information received from the basic input / output system when the operating state meets the first state; and configured to perform second fault management when the operating state meets the second state.

[0013] According to still another aspect of the present disclosure, there is provided an electronic device, including: a memory, configured to store computer-readable instructions; and a processor, configured to run the computer-readable instructions such that the electronic device executes the fault management method as described above.

[0014] According to yet another aspect of the present disclosure, there is provided a computer program product, including a computer program, wherein the computer program, when executed by a processor, implements the fault management method as described above.

[0015] As will be described in detail below, the fault management method according to an embodiment of the present disclosure comprehensively manages faults in both the normal state and the abnormal state of the operating system. By recording and / or processing multiple corrective problems and / or minor uncorrected problems, timely early warning of faults and / or accurate positioning in the event of serious or fatal uncorrected problems are achieved, thereby improving the stability and reliability of the system.

[0016] It is to be understood that both the foregoing general description and the following detailed description are exemplary and are intended to provide further explanation of the claimed technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] By describing the embodiments of the present disclosure in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present disclosure will become more apparent. The accompanying drawings are used to provide a further understanding of the embodiments of the present disclosure, and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the present disclosure and do not constitute a limitation to the present disclosure. In the drawings, the same reference numerals generally represent the same components or steps.

[0018] Figure 1 is a schematic diagram illustrating an HBM-related architecture.

[0019] Figure 2 is a schematic diagram of an application scenario of the fault management method according to an embodiment of the present disclosure.

[0020] Figure 3 is a flowchart of the method of the fault management method according to an embodiment of the present disclosure.

[0021] Figure 4 is a general flowchart of the fault management method according to an embodiment of the present disclosure.

[0022] Figure 5 is a schematic diagram of the fault management device according to an embodiment of the present disclosure.

[0023] Figure 6 is a hardware block diagram of an electronic device according to an embodiment of the present disclosure.

[0024] Figure 7 is a schematic diagram of a computer program product according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0025] In order to make the objectives, technical solutions, and advantages of the present disclosure more apparent, exemplary embodiments according to the present disclosure will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments of the present disclosure. It should be understood that the present disclosure is not limited by the exemplary embodiments described herein.

[0026] The fault management method according to the embodiments of the present disclosure can be applied to a computing device. The following will refer to Figure 1 - Figure 2 to outline the application scenarios according to the embodiments of the present disclosure.

[0027] Figure 1 is a schematic diagram showing the HBM-related architecture. Different from the traditional Double Data Rate Dual In-line Memory Module (DDR DIMM), HBM usually does not appear in the form of an independent memory module, but is bundled with a processor such as a CPU or GPU as a customized solution, belonging to In Package Memory (IPM).

[0028] As Figure 1 shown in (A) therein, the side view of the HBM-related architecture may include:

[0029] Multiple High Bandwidth Memory Dynamic Random Access Memory Dies (HBM DRAM Dies, or High Bandwidth Memory Dynamic Random Access Memory grains), HBM DRAM Dies are the most basic storage units that make up HBM, responsible for the specific data storage function, and have the capabilities of writing, reading, and storing data. Multiple HBM DRAM Dies achieve vertical interconnection between chip layers through Through-Silicon Vias (TSVs) and Microbumps to form a High Bandwidth Memory Stack (HBM Stack, or High Bandwidth Memory Stacking). This connection method can provide a high-speed and low-latency signal transmission channel to ensure the rapid transfer of data between different chip layers.

[0030] A Logic Die, which can be responsible for processing and coordinating data and other logical functions.

[0031] A high-performance computing engine, which can be a Graphics Processing Unit (GPU) / Central Processing Unit (CPU) / (System-on-Chip) Soc Die (for example: other Application Specific Integrated Circuits (ASICs)), etc.

[0032] An Interposer, which can play the role of electrical connection and signal transmission, connect the HBM Stack with the underlying Package Substrate, and further connect with the GPU / CPU / Soc Die, etc., to achieve data interaction.

[0033] As Figure 1 shown in (B) therein, the top view of the HBM-related architecture may include:

[0034] The central processing core, input / output interfaces, DDR controllers, and HBM controllers (Cores, IO, DDR Controllers, HBM Controllers) are the core control and data processing units of the system.

[0035] Among them, Cores (also known as cores) are the key parts of the system for data processing and operations. For example, CPU cores, GPU cores, etc. are responsible for executing instructions and processing data, and need to read data from memory (HBM Stack and DDR) and write the processing results.

[0036] DDR controllers and HBM controllers can be collectively referred to as memory controllers, which are responsible for managing operations such as data reading and writing of DDR and HBM Stack, and coordinating data transmission between memory and other components.

[0037] The HBM Stack is a memory structure formed by assembling multiple HBM DRAM Die chips together in a vertical stacking manner. The number of layers can generally reach 8 layers or even more (for example: High Bandwidth Memory 2nd Generation Enhanced (HBM2e) can reach 12 layers, and High Bandwidth Memory 3rd Generation (HBM3) can reach 16 layers). It is used to store data and is connected to the core processing unit through a memory controller to provide a high-speed data transmission channel.

[0038] DDR, double data rate synchronous dynamic random access memory, is connected to the central processing core through a controller and provides memory support for the system together with HBM.

[0039] As described above, HBM works in cooperation with other components through a vertical stacking connection method to achieve high-speed data transmission and storage functions to meet the requirements of high-performance computing and other scenarios. However, it is very difficult to troubleshoot such a complex and high-density packaging method once a failure occurs, and there is currently no perfect mechanism to deal with the occurrence of HBM failures.

[0040] Figure 2 It is a schematic diagram of the application scenario of the fault management method according to an embodiment of the present disclosure. As Figure 2 shown, the above computing device can be a server 20, which can at least include the following modules.

[0041] The high-performance computing chip 201. As described above, a memory controller, a processor, and an HBM Stack are encapsulated in the high-performance computing chip 201. Among them, the processor is the operation and control core of the server 20, responsible for executing instructions and processing data; the HBM Stack provides high-bandwidth memory support for the processor to meet the high requirements for data transmission speed; the memory controller is responsible for managing memory access, coordinating data transmission between the DIMM, the processor, and the HBM Stack, and ensuring the efficiency and accuracy of data reading and writing. The DIMM is the memory module of the server 20, used to store data, and can provide data storage and reading services for the processor and other components.

[0042] In addition, registers are also encapsulated in the high-performance computing chip 201. Further, it can be an HBM MCA (Machine Check Architecture) register. When the memory controller detects a hardware error, it will record the error information in the HBM MCA register. Specifically, it can include: a global control / status register and an error reporting register group (BANK). The global register is used to configure and monitor the overall state of the MCA, while the BANK register group is used to record the error information of specific hardware units.

[0043] The Basic Input Output System (BIOS) 202 is the program that runs first when the computer starts up and is responsible for initializing hardware, self-checking, etc. The BIOS can read the content of the HBM MCA register.

[0044] The fault management device 203 is used to handle the situation when an HBM fails. Specifically, it will be further described in conjunction with Figure 3 - Figure 5 for further description.

[0045] Figure 3 is a flowchart of the fault management method according to an embodiment of the present disclosure. As Figure 3 shown, the fault management method can at least include the following steps.

[0046] In step S301, in response to a high-bandwidth memory failure, determine the operating state of the operating system.

[0047] As described above, the present disclosure aims to solve the problem of how to handle HBM failures. HBM failures can be divided into Corrected Error (CE) and Uncorrected Error (UCE), and the UCE can be further divided according to its degree.

[0048] Among them, when a High Bandwidth Memory (HBM) fails, errors can be automatically corrected through techniques such as Error Correction Code (ECC). Errors that can be successfully corrected are called Correctable Errors (CE); correspondingly, errors that exceed the error correction ability of the ECC (for example, the errors are too severe) or the error correction mechanism itself fails, and errors that cannot be detected or corrected will be marked as Uncorrectable Errors (UCE).

[0049] In an embodiment of the present disclosure, CE may include: errors such as bit flip correction and single-byte error correction; UCE may include: errors such as multi-bit errors exceeding the error correction ability, inability to correct errors due to ECC hardware failure, and physical damage to memory chips.

[0050] According to different types and / or degrees, the impact of HBM failure on the operating system (OS) is also different. Therefore, the operating state of the OS can be determined first, and subsequent corresponding fault management can be carried out according to different operating states.

[0051] In step S302, when the operating state satisfies the first state, based on the first fault information received from the Basic Input / Output System (BIOS), the first fault management is performed; when the operating state satisfies the second state, the second fault management is performed. Specifically as follows:

[0052] (1) When the operating state of the OS satisfies the first state, based on the first fault information received from the BIOS, the first fault management is performed.

[0053] In an embodiment of the present disclosure, the first state may be that the operating system is running normally. That is to say, when the OS can still run normally, the data channel (or in-band) of the OS is available, and the in-band collection unit in the BIOS 202 can collect various information through this data channel, including HBM fault information (i.e., the first fault information), and then report the first fault information to the fault management device 203.

[0054] The fault management device 203 performs corresponding first fault management based on the received first fault information. Since the OS can still run normally at this time, the first fault information is basically CE and / or minor UCE, and the main way for the fault management device 203 to perform the first fault management on it is to read, parse, and record the first fault information and / or generate warning information.

[0055] (2) When the operating state of the OS satisfies the second state, the second fault management is performed.

[0056] In an embodiment of the present disclosure, the second state may be that the operating system is running abnormally (such as crashing). That is to say, at this time, the OS can no longer run normally, and the data channel (or in-band) of the OS is unavailable, and an out-of-band data channel independent of the in-band is needed to collect fault information.

[0057] At this time, out-of-band collection can be performed by the fault management device 203. That is to say, the fault management device 203 can operate independently of the OS and has a separate processor, memory, network connection, etc., and can perform second fault management on the hardware when the OS is not running or fails.

[0058] In an embodiment of the present disclosure, the fault management device 203 may be a BaseBoard Management Controller (BMC).

[0059] Specifically, performing the second fault management may include: obtaining second fault information of the register; generating a fault result and / or warning information based on the second fault information, the first preset rule and / or the parsed first fault information and / or the second preset rule.

[0060] The fault management device 203 actively collects the second fault information recorded in the HBM MCA register. Since the OS cannot run normally at this time, the second fault information is basically a severe or fatal UCE. After the fault management device 203 collects the above severe or fatal UCE, it analyzes it, locates the fault type and / or location, and gives a corresponding processing solution. Specifically, it will be described in detail in Figure 4 below.

[0061] Figure 4 FIG. is an overall flowchart of a fault management method according to an embodiment of the present disclosure. For convenience of description, the fault management device 203 is referred to as BMC203, but it does not constitute a limitation. As Figure 4 shown, the overall process of the fault management method may at least include the following steps:

[0062] In step S1, when a fault occurs in the HBM, the relevant register records the fault information of the HBM.

[0063] In step S2, it is determined whether the OS has crashed (i.e., the running state). As described above, fault management is performed at least based on the running state of the operating system. If the OS has not crashed, step S3a is executed; if the OS has crashed, step S3b is executed.

[0064] In step S3a, the BIOS obtains the first fault information. As described above, at this time the OS is running normally and the in-band data channel is available. The BIOS 202 can obtain the CE and / or minor UCE in the HBM MCA register, and then send it to the diagnostic unit in the BMC203 for processing and / or diagnosis.

[0065] In one embodiment of the present disclosure, the BIOS 202 can collect HBM fault information in relevant registers through CSMI (Corrected Machine Check Interrupt), and report it to the diagnostic unit in the BMC 203 through Intelligent Platform Management Interface (IPMI) commands.

[0066] Among them, IPMI is an open standard hardware management interface specification that defines a set of standardized commands and communication protocols for remotely managing and monitoring computer systems.

[0067] In step S3b, the BMC obtains the second fault information. As described above, at this time, the OS cannot run normally, the in-band data channel is unavailable, and the in-band collection unit of the BIOS 202 cannot work. At this time, the out-of-band collection unit in the BMC 203 independent of the OS is triggered to actively collect the severe or fatal UCE recorded in the current register, and then send it to the diagnostic unit of itself (BMC 203) for processing and / or diagnosis.

[0068] In one embodiment of the present disclosure, when the OS crashes, a severe fault out-of-band signal (for example: Catastrophic Error (CATERR)) can be generated, and this signal can trigger the BMC 203 to collect the second fault information through the out-of-band data channel.

[0069] In one exemplary embodiment of the present disclosure, the first fault information and / or the second fault information collected by the out-of-band collection unit in the BIOS 202 and / or the BMC 203 may include one or more of the following:

[0070] System Event Log (SEL), which is used to record various important events and fault information that occur in the server system;

[0071] Fault information related time (time), such as the timestamp when the fault occurs, the system running time, the fault duration, etc.;

[0072] Fault information source, for example, it can be the BIOS 202 and / or the BMC 203;

[0073] Corresponding Central Processing Unit socket (CPU socket) information, such as socket type, socket quantity, the status of each socket, etc.;

[0074] Corresponding to the central processing unit core information, such as the number of cores, core frequency, core temperature, core utilization rate, core failure information, etc.;

[0075] Corresponding to the central processing unit related sub-module information, such as the names of the CPU's sub-module IPs (Intellectual Property, or IP cores). Specifically, it can be the IP names related to the memory controller of the Intel Xeon processor platform, such as High Bandwidth Memory-Mesh to Memory (HBM-M2M), High Bandwidth Memory-PseudoCh0 (HBM-PS0), High Bandwidth Memory-PseudoCh1 (HBM-PS1).

[0076] In step S4, the BMC processes and / or diagnoses the first failure information and / or the second failure information. As described above, the first failure information and the second failure information are reported in step S3a and step S3b respectively. After receiving this information, the diagnostic unit of the BMC203 will process and / or diagnose it as follows:

[0077] 1) When receiving the first failure information sent by the BIOS202, the diagnostic unit of the BMC203 can perform the following steps:

[0078] 1a) Read and parse the first failure information.

[0079] In an exemplary embodiment of the present disclosure, processing the first fault information may include: reading the Intel Architecture 32-bit Machine Check Interrupt status (IA32_MCi_status) registers of HBM-M2M and HBM-PS, as well as some auxiliary registers, such as the Intel Architecture 32-bit Machine Check Interrupt Address (IA32_MCi_ADDR), the Intel Architecture 32-bit Machine Check Interrupt Miscellaneous (IA32_MCi_MISC), etc., parsing out the HBM Stack, Memory Controller (MC), channel information, and reading the code registers, such as: Machine Check Architecture Cause of Error Code (MCACOD), Machine Check Status Code (MSCOD), and parsing out the error code and the cause of the error.

[0080] 1b) Based on the first preset rule, determine the first fault type of the parsed first fault information.

[0081] In an exemplary embodiment of the present disclosure, the first preset rule may include determining the error type based on the bit status in the MAC. For example: judging the type of the fault information according to these bits such as the Enable (EN), Uncorrected (UC), Previous Machine Check Error Cleared (PCC), Severity (S), and Action Required (AR).

[0082] The first type of fault may include Fatal Error, Catastrophic Error, Software Recoverable Action Optional (SRAO), Software Recoverable Action Required (SRAR), Uncorrected Non-Actionable Error (UCNA), etc.

[0083] 1c) When the first type of fault meets the first type, record the parsed first fault information and / or generate a warning message.

[0084] In an exemplary embodiment of the present disclosure, the first type may be UCNA. As described above, after the judgment in 1b) above, when the first type of fault is UCNA, it is recorded and saved in the UCNA set. That is to say, faults of the UCNA type do not require immediate action, and BMC203 can record them to form a fault set so that the location of the fault can be more quickly located when a serious or fatal UCE occurs subsequently.

[0085] Further, when the first fault information is a certain fault in UCNA and a warning may be required to prompt its risk, a warning message is generated, and then step S5 is executed.

[0086] 2) When receiving the second fault information sent by the out-of-band collection unit of BMC203, the diagnostic unit of BMC203 may perform the following steps:

[0087] 2a) Read and parse the second fault information. The specific method is the same as 1a) above, except that the fault information read and parsed is different, and will not be elaborated here.

[0088] 2b) Based on the first preset rule, determine the second type of fault of the parsed second fault information. The specific method is the same as 1b) above, except that the determined fault information is different, and will not be elaborated here.

[0089] 2c) When the second type of fault meets the first type, record the parsed second fault information and / or warning message. The specific method is the same as 1c) above, except that the recorded fault information is different, and will not be elaborated here.

[0090] Similarly further, when the second fault information is a certain fault in UCNA and a warning may be required to prompt its risk, a warning message is generated, and then step S5 is executed.

[0091] 2d) In the case where the second fault type does not meet the first type, generate a fault result based on the parsed second fault information, the first preset rule, the parsed first fault information, and / or the second preset rule.

[0092] In an exemplary embodiment of the present disclosure, the first type may be UCNA, and the second preset rule may be an empirical rule base. For example, a mapping set of past UCNA data sets and information such as fault locations and fault solutions. That is to say, the second fault type may be a fatal error, a catastrophic error, SRAO, SRAR, etc. At this time, the fault is very serious and requires accurate location and problem solving. The diagnostic unit of BMC203 can combine the UCNA fault set formed by the CE and / or minor UCE collected by the previous BIOS202, and / or the UCNA type faults collected by the out-of-band collection unit of BMC203, with faults of types such as fatal errors, catastrophic errors, SRAO, SRAR, etc., and perform comprehensive judgment based on the empirical rule base. For example, it is possible to compare and / or match the information such as CPUSocket, HBMStack, MC, channel, etc. obtained in step S3b and / or in 1a) and / or 2a) in step S4 with the UCNA fault set formed in 1c) and / or 2c) in step S4, and / or the empirical rule base, and finally determine the HBM fault result, where the HBM fault result may include information such as the fault location and the solution.

[0093] It should be noted that in the case of an OS crash, 2c) and 2d) above can exist simultaneously. That is to say, multiple HBM faults may occur, one or more of which belong to UCNA, and one or more of the others belong to one or more of fatal errors, catastrophic errors, SRAO, SRAR, etc.

[0094] In step S5, the BMC outputs the fault result and / or warning information.

[0095] As described above, when there is an HBM fault result that needs to be warned and / or determined in 1c) and / or 2c) and / or 2d) above, it can be sent by the diagnostic unit of BMC203 to the display unit in BMC203 for output display.

[0096] So far, the overall process of the fault management method in the embodiments of the present disclosure has been introduced.

[0097] Based on the above overall process, the present disclosure provides an application example, which is as follows:

[0098] The HBM fails, but the OS can still run normally, indicating a CE and / or minor UCE;

[0099] BIOS202 obtains the first fault information and reports it to BMC203;

[0100] BMC203 processes and diagnoses. Specifically, on the Intel Xeon processor EGS platform, the PS0 / 1 register in the MCA record sub-module HBM-M2M records register information such as the fault address, status, and control of the HBM. It is found that a minor UCE occurs at the physical address of HBM (socket0, stack1, channel0, ps0, BankGroup0, Bank0, row20000, Column400), and the system address is 0x003898768; then parse MCACOD, MSCOD, and the corresponding UCE severity level, and judge the error type according to the EN, UC, PCC, S, and AR bits, and judge that the type of this UCE is UCNA;

[0101] BMC203 puts the above fault into the UCNA fault set;

[0102] After that, multiple minor UCEs may occur, and their types are all UCNA, then continuously record these faults into the UCNA fault set;

[0103] Then one day the OS suddenly crashes, and BMC203 actively collects the second fault information in the register;

[0104] After the processing and diagnosis of BMC203, it is found that at this time a severe or fatal UCE occurs at the physical address of HBM (socket0, stack1, channel0, ps0, BankGroup0, Bank1, row10000, Column200), and the system address is 0x0040987843; based on this physical address, compare information such as socket ID, stack, channel, and ps, if it is found that it matches the previously recorded minor UCE, then the HBM fault diagnosis is successful, and solution information can also be matched according to the experience rule library;

[0105] Finally, the diagnosis result is output.

[0106] Figure 5 It is a schematic diagram of a fault management device according to an embodiment of the present disclosure. As Figure 5 shown, the fault management device 500 may at least include the following modules:

[0107] A determination module 501, configured to determine the operating state of the operating system in response to a high-bandwidth memory fault.

[0108] The management module 502 is configured to perform first fault management based on the first fault information received from the basic input / output system in response to the running state satisfying the first state; and to perform second fault management in response to the running state satisfying the second state.

[0109] Further, the management module 502 may include:

[0110] The out-of-band collection unit 5021 is configured to obtain the second fault information of the register when the running state satisfies the second state;

[0111] The diagnosis unit 5022 is configured to read and parse the first fault information when the running state satisfies the first state; determine the first fault type of the parsed first fault information based on the first preset rule; and record the parsed first fault information and / or generate a warning message when the first fault type satisfies the first type.

[0112] And, when the running state satisfies the second state, obtain the second fault information of the register, and generate a fault result and / or a warning message based on the second fault information, the first preset rule, and / or the parsed first fault information and / or the second preset rule. Specifically, it may be to read and parse the second fault information; determine the second fault type of the parsed second fault information based on the first preset rule; record the parsed second fault information and / or generate a warning message when the second fault type satisfies the first type; and generate a fault result based on the parsed second fault information, the parsed first fault information, and / or the second preset rule when the second fault type does not satisfy the first type.

[0113] The output unit 5023 is configured to output the fault result and / or the warning message.

[0114] Above, the first state includes: the operating system is running normally; the second state includes: the operating system is running abnormally; the first fault information includes: corrected errors and / or minor uncorrected errors; the second fault information includes: serious or fatal uncorrected errors; the first type includes: uncorrected errors that do not require any action.

[0115] The parameters of the first fault information and the second fault information include one or more of the following: system time log; time related to the fault information; source of the fault information; corresponding central processing unit slot information; corresponding central processing unit core information; corresponding central processing unit related sub-module information.

[0116] Figure 6It is a hardware block diagram of an electronic device according to an embodiment of the present disclosure. The electronic device according to an embodiment of the present disclosure at least includes a processor; and a memory for storing computer-readable instructions. When the computer-readable instructions are loaded and run by the processor, the processor executes the fault management method as described above.

[0117] Figure 6 The illustrated electronic device 600 specifically includes: a central processing unit (CPU) 601, a graphics processing unit (GPU) 602, and a memory 603. These units are interconnected via a bus 604. The central processing unit (CPU) 601 and / or the graphics processing unit (GPU) 602 can be used as the above-mentioned processor, and the memory 603 can be used as the memory for storing the computer-readable instructions. In addition, the electronic device 600 may further include a communication unit 605, a storage unit 606, an output unit 607, an input unit 608, and an external device 609, and these units are also connected to the bus 604.

[0118] Figure 7 It is a schematic diagram of a computer program product according to an embodiment of the present disclosure. As Figure 7 shown, on the computer program product 700 according to an embodiment of the present disclosure, a computer program 701 is stored. When the computer program 701 is executed by a processor, the fault management method described with reference to the above drawings is executed. The computer program product includes, but is not limited to, for example, volatile memory and / or non-volatile memory. The volatile memory may, for example, include random access memory (RAM) and / or cache memory, etc. The non-volatile memory may, for example, include read-only memory (ROM), hard disk, flash memory, optical disc, magnetic disk, etc.

[0119] As described above, with reference to the drawings, a fault management method, apparatus, electronic device, and computer program product according to an embodiment of the present disclosure are described. According to the fault management method of the embodiment of the present disclosure, the method performs comprehensive fault management on the operating system in normal and abnormal states, and through the recording and / or processing of multiple corrective problems and / or minor uncorrected problems, realizes timely early warning of faults and / or accurate positioning when serious or fatal uncorrected problems occur, thereby improving the stability and reliability of the system.

[0120] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present disclosure.

[0121] The basic principles of the present disclosure have been described in connection with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present disclosure are only examples and not limitations. It cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present disclosure. Additionally, the specific details disclosed above are only for illustrative and facilitating understanding purposes, rather than limitations. These details do not limit the present disclosure to necessarily adopt these specific details for implementation.

[0122] The block diagrams of the devices, apparatuses, equipment, and systems involved in the present disclosure are only illustrative examples and do not intend to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any way. Words such as "comprising," "including," "having," etc. are open-ended terms, meaning "including but not limited to," and can be used interchangeably with each other. The word "or" and "and" used herein refer to the word "and / or" and can be used interchangeably with it, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to" and can be used interchangeably with it.

[0123] Additionally, as used herein, the "or" used in the listing of items starting with "at least one" indicates a disjunctive listing. So, for example, the listing of "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Moreover, the term "exemplary" does not mean that the described examples are preferred or better than other examples.

[0124] It should also be noted that in the systems and methods of the present disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of the present disclosure.

[0125] Various changes, substitutions, and alterations to the technologies described herein can be made without departing from the teachings of the technology defined by the appended claims. Additionally, the scope of the claims of the present disclosure is not limited to the specific aspects of the processes, machines, manufactures, compositions of events, means, methods, and acts described above. Current or later-developed processes, machines, manufactures, compositions of events, means, methods, or acts that perform substantially the same function or achieve substantially the same result as the corresponding aspects described herein can be utilized. Thus, the appended claims include such processes, machines, manufactures, compositions of events, means, methods, or acts within their scope.

[0126] The foregoing description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present disclosure. Thus, the present disclosure is not intended to be limited to the aspects shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0127] The foregoing description has been presented for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of the present disclosure to the forms disclosed herein. Although several example aspects and embodiments have been discussed above, those skilled in the art will recognize some of their variations, modifications, alterations, additions, and subcombinations.

Claims

1. A fault management method, applied to a computing device, characterized in that The method includes: Determining the operating state of the operating system in response to a high - bandwidth memory failure; In response to the operating state satisfying the first state, performing first - fault management based on the first fault information received from the basic input / output system; in response to the operating state satisfying the second state, performing second - fault management.

2. The fault management method according to claim 1, wherein Performing the first - fault management includes: Reading and parsing the first fault information; Based on a first preset rule, determining the first fault type of the parsed first fault information; When the first fault type satisfies the first type, recording the parsed first fault information and / or generating a warning message.

3. The fault management method according to claim 2, characterized in that, Performing the second - fault management includes: Obtaining second fault information of a register; Based on the second fault information, the first preset rule, and / or the parsed first fault information and / or a second preset rule, generating a fault result and / or a warning message.

4. The fault management method according to claim 3, wherein Based on the second fault information, the first preset rule, and / or the parsed first fault information and / or the second preset rule, generating a fault result and / or a warning message includes: Reading and parsing the second fault information; Based on the first preset rule, determining the second fault type of the parsed second fault information; When the second fault type satisfies the first type, recording the parsed second fault information and / or generating the warning message; When the second fault type does not satisfy the first type, generating the fault result based on the parsed second fault information, the parsed first fault information, and / or the second preset rule.

5. The fault management method according to claim 1, characterized in that, The method further includes: Outputting the fault result and / or the warning message.

6. The fault management method according to any one of claims 1 - 5, wherein: The first state includes: the operating system is running normally; The second state includes: the operating system is running abnormally; The first fault information includes: correctable errors and / or minor uncorrectable errors; The second fault information includes: severe or fatal uncorrectable errors; The first type includes: uncorrectable errors that do not require action.

7. The fault management method according to any one of claims 1-6, characterized in that The parameters of the first fault information and the second fault information include one or more of the following: System time log; Time related to the fault information; Source of the fault information; Corresponding central processing unit (CPU) socket information; Corresponding CPU core information; Corresponding CPU - related sub - module information.

8. A fault management device, characterized in that, The device includes: A determination module, configured to determine the operating state of the operating system in response to a high - bandwidth memory failure; A management module, configured to perform first - fault management based on the first fault information received from the basic input / output system when the operating state satisfies the first state; and configured to perform second - fault management when the operating state satisfies the second state.

9. An electronic device, characterized in that, It includes: A memory, configured to store computer - readable instructions; And A processor, configured to run the computer - readable instructions, so that the electronic device executes the fault management method according to any one of claims 1 to 7.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the fault management method according to any one of claims 1 to 7.