Fault information processing method, electronic device, storage medium and program product
By employing a collaborative approach between in-band and out-of-band fault processors, the problem of excessive restarts caused by hardware and software separation in graphics card failures is resolved. This enables precise fault location and rapid processing, improving system reliability and availability while reducing maintenance costs.
Patent Information
- Application Number
- CN202511220825.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-29
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2045-08-29
AI Technical Summary
In existing technologies, hardware monitoring and software optimization for graphics card failures suffer from hardware-software separation, leading to excessive restarts, which affect data processing efficiency, data integrity, and user experience. Furthermore, the maintenance costs are high, and it is difficult to achieve accurate fault location and early warning.
A method of in-band and out-of-band fault processors working together is adopted. The in-band fault processor performs software-level fault detection and processing, while the out-of-band fault processor performs hardware-level fault processing, forming a dual-channel, zero-blind-zone fault detection mechanism. Based on multi-dimensional target conditions, the fault information is comprehensively judged and processed.
It enables accurate location and rapid processing of fault information, avoids the blind spots of traditional single-channel detection, improves system reliability and availability, reduces operation and maintenance costs, and ensures data integrity and user experience.
Smart Images

Figure CN120743637B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, in particular to a fault information processing method, an electronic device, a storage medium and a program product. BACKGROUND
[0002] In the technical fields of large model training and application, game entertainment, professional creation, scientific calculation, automatic driving and medical health, the graphics processing unit (GPU) in the display card can provide efficient parallel processing capability. In actual application, GPU card drop will cause serious harm, affecting data processing efficiency, data integrity and user experience. In related technologies, for display card failure, mainly through hardware monitoring and software optimization to troubleshoot. However, the binary thinking of traditional hardware monitoring and software optimization has the problem of separation of software and hardware, and excessive restart can easily lead to abnormality of the whole task and system, data security is difficult to guarantee and operation and maintenance cost is high. SUMMARY
[0003] In view of the above problems, the present application provides a fault information processing method, device, equipment, medium and program product.
[0004] According to a first aspect of the present application, a fault information processing method is provided, applied to an in-band fault processor, comprising: in response to detecting that the state information of a processing module in a to-be-detected device is abnormal information, performing a first update operation on a target process in the processing module to obtain a first operation result; in the case that the first operation result, the fault type of the to-be-detected device, and the detection result based on the detection instruction meet a target condition, sending fault information to an out-of-band fault processor, so that the out-of-band fault processor performs a second update operation on a target module in a plurality of modules based on the fault information to obtain a second operation result, the in-band fault processor is a processor associated with the to-be-detected device, used for performing in-band fault processing on the to-be-detected device, the out-of-band fault processor is a processor independent of the to-be-detected device, used for performing out-of-band fault processing on the to-be-detected device in the case that the first operation result, the fault type of the to-be-detected device, and the detection result based on the detection instruction meet the target condition; wherein the first operation result, the fault type of the to-be-detected device, and the detection result based on the detection instruction meet the target condition include at least one of the following: the first operation result indicates that the first update operation fails; the fault type indicates that the fault type of the to-be-detected device is a hardware fault type; the instruction execution result based on the detection instruction indicates that the processing module or the plurality of modules is abnormal; the detection instruction executes abnormally, and the detection instruction is used to detect the module state of a plurality of modules associated with the processing module.
[0005] The second aspect of the present application provides a fault information processing method applied to an out-of-band fault processor, and the method comprises the following steps: in response to receiving fault information sent by an in-band fault processor, performing a second update operation on a target module based on the fault information, and obtaining a second operation result, wherein the fault information is obtained by the above method.
[0006] The third aspect of the present application provides a fault information processing device applied to an in-band fault processor, and the device comprises the following modules: a first update module, which is configured to perform a first update operation on a target process in a processing module in a to-be-detected device in response to detecting that the state information of the processing module is abnormal information, and obtain a first operation result; and a sending module, which is configured to send fault information to an out-of-band fault processor in the case that the first operation result, the fault type of the to-be-detected device, and the detection result based on a detection instruction meet a target condition, so that the out-of-band fault processor performs a second update operation on a target module in a plurality of modules based on the fault information, and obtains a second operation result, wherein the in-band fault processor is a processor associated with the to-be-detected device, and is configured to perform in-band fault processing on the to-be-detected device, the out-of-band fault processor is a processor independent of the to-be-detected device, and is configured to perform out-of-band fault processing on the to-be-detected device in the case that the first operation result, the fault type of the to-be-detected device, and the detection result based on the detection instruction meet the target condition; and wherein the first operation result, the fault type of the to-be-detected device, and the detection result based on the detection instruction meeting the target condition comprises at least one of the following: the first operation result indicating that the first update operation fails; the fault type indicating that the fault type of the to-be-detected device is a hardware fault type; the detection result based on the detection instruction indicating that the processing module or the plurality of modules are abnormal; and the detection instruction being abnormal, wherein the detection instruction is used to detect the module state of the plurality of modules associated with the processing module.
[0007] The fourth aspect of the present application further provides a fault information processing device applied to an out-of-band fault processor, and the device comprises the following modules: a second update module, which is configured to perform a second update operation on a target module based on fault information in response to receiving the fault information sent by an in-band fault processor, and obtain a second operation result, wherein the fault information is obtained according to the above method.
[0008] The fifth aspect of the present application provides an electronic device, which comprises: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.
[0009] The sixth aspect of the present application further provides a computer-readable storage medium, which stores a computer program or instructions, wherein the computer program or instructions are executed by a processor to implement the steps of the above method.
[0010] The seventh aspect of the present application also provides a computer program product comprising computer programs or instructions, which, when executed by a processor, implement the steps of the above method. BRIEF DESCRIPTION OF DRAWINGS
[0011] The above and other objects, features and advantages of the present application will become more apparent from the following description when taken in conjunction with the accompanying drawings, in which:
[0012] Figure 1 An application scenario diagram of a fault information processing method, device, equipment, medium and program product according to an embodiment of the present application is shown;
[0013] Figure 2 A flowchart of a fault information processing method according to an embodiment of the present application is shown;
[0014] Figure 3 A flowchart of another fault information processing method according to an embodiment of the present application is shown;
[0015] Figure 4 A block diagram of a fault information processing system according to an embodiment of the present application is shown;
[0016] Figure 5 A structural block diagram of a fault information processing device according to an embodiment of the present application is shown;
[0017] Figure 6 A structural block diagram of another fault information processing device according to an embodiment of the present application is shown;
[0018] Figure 7 A block diagram of an electronic device suitable for implementing a fault information processing method according to an embodiment of the present application is shown;
[0019] Figure 8 A block diagram of another electronic device suitable for implementing a fault information processing method according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0020] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. It is to be understood, however, that the description is merely exemplary and is not intended to limit the scope of the present application. In the following detailed description of the embodiments of the present application, numerous specific details are set forth in order to provide a thorough understanding of the embodiments of the present application. However, it would be apparent to those skilled in the art that the embodiments of the present application can be practiced without these specific details. In other instances, well-known structures and techniques have not been described in detail in order to avoid obscuring aspects of the present application.
[0021] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. As used herein, the terms "comprises", "comprising", "includes", "including" and the like are specifically intended to be open-ended and to mean that other features, steps, operations, and / or components can be added.
[0022] All terms used herein including technical and scientific terms have the meanings commonly understood by one of ordinary skill in the art unless otherwise defined herein. It should be noted that the terms "comprise", "comprising", "comprises" and the like can have the meaning ascribed to it in U.S. Patent law; it can mean "includes", "including", and the like; it can mean "consists of and the like.
[0023] In the case of using expressions similar to "at least one of A, B, and C, etc.", it will be generally understood that the same is intended to mean any one of A, B, and C or any combination of A, B, and C, etc.
[0024] In the large model training scene, GPU card drop not only causes program crash and data loss, but also seriously affects work efficiency and increases operation and maintenance cost. Historical statistics show that GPU problems account for a high proportion (up to 58.7%) in the interruption of intensive training. On the other hand, large-scale training tasks have strict requirements for the performance and stability of GPUs. Long-time high-load operation, GPU is prone to failure due to overheating, power supply shortage and other problems, and the compatibility of software and hardware in the cluster environment is more complex. Conflicts between the driver program, operating system and application program can all be potential factors for GPU card drop.
[0025] In some examples, troubleshooting for graphics cards mainly includes hardware troubleshooting, software optimization and regular maintenance. Hardware troubleshooting mainly includes checking bus slots, power supply, heat dissipation and gold finger contact; fault diagnosis is achieved by upgrading power supply, using single independent power supply line, and analyzing received error codes.
[0026] However, hardware troubleshooting relies on human experience and is difficult to find intermittent poor contact or micro solder point fracture, resulting in failure to locate logical layer failure. Error codes only indicate results, not reasons, and most of them are difficult to prevent after diagnosis, and direct abnormal restart can damage hardware devices, resulting in high resource cost and long time consumption.
[0027] In some examples, software optimization mainly includes updating graphics card drivers and motherboard basic input / output system to the latest stable version and reporting driver errors through log analysis. However, driver and firmware updates cannot solve the card drop problem caused by motherboard compatibility, power transient and poor bus signal integrity, and log analysis requires professional interpretation and is difficult to achieve early warning.
[0028] When GPU card drop occurs, the traditional method is to check the hardware connection state, ensure that the graphics card power supply line or signal line is correctly connected, and the graphics card is completely installed into the bus slot. After determining that the hardware connection is normal, the system is restarted to verify whether the graphics card can be recovered. However, in the scenario of distributed data parallel processing tasks, for example, large model data processing, large model training or inference tasks usually require continuous training for several days or even weeks, and abnormal restart may cause the loss of unsaved intermediate states (such as gradients, weights, and optimizer states), which need to be recovered from the last checkpoint or even start training from scratch. If the checkpoint saving frequency is low (for example, once an hour), the restart may lose several hours of computing results, significantly prolonging the overall training time. Therefore, in distributed data parallel or model parallel training, node restart may cause parameter synchronization failure, requiring all nodes to be re-coordinated, or even causing the entire task to terminate.
[0029] Abnormal restart also causes GPU resource limitations, and frequent abnormal restarts may accelerate the wear and tear of GPUs, power supplies, storage devices, and other hardware, especially abnormal shutdowns caused by power fluctuations or heat dissipation problems. After restarting, the task may need to be re-queued for resource allocation, resulting in a decrease in cluster resource utilization. If the direct abnormal restart involves critical nodes (such as parameter servers, schedulers, and storage gateways), it may cause the entire cluster to malfunction, and there may be problems such as memory data loss, file system damage, and monitoring interruption.
[0030] In some examples, the temperature and voltage of the graphics card are monitored through monitoring tools outside the system (such as baseboard management controllers), and fault detection is performed in the case of data anomalies; or the system internal tools are run to determine whether the graphics card is abnormal according to the state data reported by the driver. However, relying solely on monitoring tools outside the system or state data inside the system for detection is prone to data islands, resulting in state discontinuity and delayed response.
[0031] In view of this, the application provides a fault information processing method applied to an in-band fault processor, including: in response to detecting that state information of a processing module in a to-be-detected device is abnormal information, performing a first update operation on a target process in the processing module to obtain a first operation result; in a case where the first operation result, a fault type of the to-be-detected device, and a detection result based on a detection instruction meet a target condition, sending fault information to an out-of-band fault processor, so that the out-of-band fault processor performs a second update operation on a target module in a plurality of modules based on the fault information to obtain a second operation result, the in-band fault processor being a processor associated with the to-be-detected device and being used for performing in-band fault processing on the to-be-detected device, the out-of-band fault processor being a processor independent of the to-be-detected device and being used for performing out-of-band fault processing on the to-be-detected device in a case where the first operation result, the fault type of the to-be-detected device, and the detection result based on the detection instruction meet the target condition; wherein the first operation result, the fault type of the to-be-detected device, and the detection result based on the detection instruction meeting the target condition include at least one of the following: the first operation result indicating that the first update operation fails; the fault type indicating that the fault type of the to-be-detected device is a hardware fault type; the instruction execution result based on the detection instruction indicating that the processing module or the plurality of modules is abnormal; and the detection instruction executing abnormally, the detection instruction being used for detecting module states of a plurality of modules associated with the processing module.
[0032] According to the embodiments of the application, by in-band processing of the to-be-detected device by the in-band fault processor and out-of-band processing by the out-of-band fault processor, the fault information processing architecture is cooperatively composed, the double-channel and zero-blind-zone fault detection is formed, and thus in a case where the in-band processor cannot successfully process the fault information based on the multi-dimensional target condition comprehensive judgment, the fault information is sent to the out-of-band fault processor for processing. Since the fault information is judged based on the multi-dimensional target condition, the target module to be updated can be accurately and quickly located, and thus the out-of-band fault processor can accurately perform the second update operation on the target module with a hardware abnormality, the fault information processing efficiency is improved, the joint analysis of the fault information and the hardware level is realized, the in-band and out-of-band parallel fault information processing mechanism is formed, the blind area of the traditional single-channel detection is avoided, the comprehensive detection and processing of the fault information are realized, the fault fine-grained accuracy positioning is met, and the reliability and usability of the system are further improved.
[0033] Figure 1 An application scenario diagram of the fault information processing method, apparatus, device, medium and program product according to the embodiments of the application is shown.
[0034] As Figure 1As shown, the application scenario 100 according to the embodiment can include an in-band fault handler 101, an out-of-band fault handler 102, a device to be detected 103, and a network 104. The network 104 is a medium for providing a communication link between the in-band fault handler 101, the out-of-band fault handler 102, and the device to be detected 103. The network 104 can include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0035] The in-band fault handler 101 can be an in-band handler associated with the device to be detected, and can be used to identify a software-level fault and perform a task-level recovery. For example, the in-band fault handler 101 performs an update operation on the device to be detected and a processing module in the device to be detected based on the detected information of the device to be detected. For example, when a fault cannot be repaired through a software-level operation, the in-band fault handler 101 can send fault information to the out-of-band fault handler 102 for subsequent fault processing.
[0036] The out-of-band fault handler 102 can be an out-of-band handler independent of the device to be detected, such as a handler or a baseboard management controller that can independently perform hardware fault processing. The out-of-band fault handler 102 can be responsible for fault identification and processing at the hardware level, and perform a hardware-level recovery operation.
[0037] The device to be detected 103 can be a monitored object, and can also be a fault trigger source, and can provide bidirectional data through in-band registers and out-of-band sensors. The in-band fault handler 101 is associated with the device to be detected 103.
[0038] It should be noted that the fault information processing method provided by the embodiment of the present application can generally be executed by the in-band fault handler 101 or the out-of-band fault handler 102. Accordingly, the fault information processing apparatus provided by the embodiment of the present application can generally be arranged in the in-band fault handler 101 or the out-of-band fault handler 102. The fault information processing method provided by the embodiment of the present application can also be executed by a server or a server cluster different from the in-band fault handler 101, the out-of-band fault handler 102, and capable of communicating with the in-band fault handler 101 and the out-of-band fault handler 102. Accordingly, the fault information processing apparatus provided by the embodiment of the present application can also be arranged in a server or a server cluster different from the in-band fault handler 101 and the out-of-band fault handler 102, and capable of communicating with the in-band fault handler 101 and the out-of-band fault handler 102.
[0039] It should be understood that Figure 1 The number of in-band fault handlers, out-of-band fault handlers, networks, and devices to be detected in the application scenario 100 is only illustrative. According to the implementation needs, there can be any number of in-band fault handlers, out-of-band fault handlers, networks, and devices to be detected.
[0040] Figure 2 A flowchart of a fault information processing method according to an embodiment of the present application is shown.
[0041] As Figure 2 shown, the fault information processing method of this embodiment is applied to an in-band fault processor, and the fault information processing method includes operation S210 to operation S220.
[0042] In operation S210, in response to detecting that the state information of the processing module in the to-be-detected device is abnormal information, a first update operation is performed on a target process in the processing module, and a first operation result is obtained.
[0043] In an embodiment of the present application, the to-be-detected device can be a processing device for large model processing, game entertainment, professional creation, scientific computing, autonomous driving, and medical health, etc. scenes, and can include multiple functional modules, including but not limited to processing modules, storage modules, connection modules, power supply modules, sensing modules, and channel management. The state information can be determined by an automatic detection system in the operating system. The target process can be a process in the processing module that is in an abnormal state. The first update operation can include deleting, resetting, etc. operations on the target process, and the in-band processing module can be a module that includes the identification and processing functions of the processing module.
[0044] For example, the detection system can be used to detect each functional module of the to-be-detected device, and in the case where the detection result of the processing module indicates that the processing module is in an abnormal state, the state information of the processing module is generated, and the in-band fault processor performs a reset operation on the target process in the abnormal state based on the state information, and obtains a reset result.
[0045] In operation S220, if the first operation result, the fault type of the device under test, and the detection result based on the detection instruction meet the target conditions, fault information is sent to the out-of-band fault processor. This causes the out-of-band fault processor to perform a second update operation on the target module among multiple modules based on the fault information, obtaining a second operation result. The in-band fault processor is a processor associated with the device under test, used to perform in-band fault processing on the device under test. The out-of-band fault processor is a processor independent of the device under test, used to perform out-of-band fault processing on the device under test if the first operation result, the fault type of the device under test, and the detection result based on the detection instruction meet the target conditions. The first operation result, the fault type of the device under test, and the detection result based on the detection instruction meeting the target conditions include at least one of the following: the first operation result indicates that the first update operation failed; the fault type indicates that the fault type of the device under test is a hardware fault; the instruction execution result based on the detection instruction indicates that the processing module or multiple modules are abnormal; or the detection instruction execution is abnormal, and the detection instruction is used to detect the module status of multiple modules associated with the processing module.
[0046] In embodiments of the present invention, the modules associated with the processing module may include a storage module, a power supply module, a temperature control module, a connection module, a read / write module, and an interface module. Fault information can be used to indicate relevant information about potential faults in the device under test. An in-band fault processor may be a processor embedded in the device under test, capable of independently performing fault detection and processing tasks. An out-of-band fault processor may be a processor independent of the operating system and the device under test. The fault types may include software fault types and hardware fault types. The in-band processor can handle software faults in the device under test, while the out-of-band processing module can handle hardware faults in the device under test.
[0047] For example, after the first update operation is performed on the target process, the target process does not complete the reset operation. The status of multiple hardware modules corresponding to the connection module can be detected based on the preset executable command to obtain status detection information. If the status detection information indicates that multiple hardware modules exist, but no registers or memory space can be read, it can be preliminarily determined that the graphics processor has an electrical fault, firmware crash, or link layer fault, so that the software-level reset command cannot be transmitted to the graphics processor at all.
[0048] In one feasible embodiment, the in-band processor may include a first module and a second module. The first module may include an in-band identification module and an in-band processing module, and the second module may include an information management module and an information processing module. The information processing module can be used to perform fault processing based on the hardware fault information when the in-band identification module identifies the fault type as hardware fault information, to initially process the hardware fault information. If the fault processing result is still a processing failure, the hardware fault information is sent to the out-of-band fault processor for further processing. Preset state conditions may be conditions for determining whether the in-band fault processor can successfully execute fault processing, and conditions for determining whether out-of-band processing is triggered. Detection instructions may be used to detect and manage the status of local or remote processing modules, and may be command-line tools for fault diagnosis of the device under test.
[0049] The execution result of a detection command indicates an anomaly in the processing module or multiple modules. This can manifest as follows: the detection command executes normally, but it cannot acquire or control information about the device under test. For example, if the execution result shows the processing module as "error state," it indicates that the processing module is malfunctioning. Similarly, the anomaly detection result returned by the detection command can indicate problems such as a kernel module failing to load, a driver malfunction, or an abnormal number of devices under test.
[0050] An error in the execution of a detection command indicates that the command itself cannot be executed correctly. For example, if an error message is returned when executing a detection command, it can be determined that the driver is not installed or has not been loaded correctly; or the detection command may not respond after execution, may hang, or may fail to exit normally.
[0051] According to embodiments of the present invention, by accurately judging different types of triggering conditions, "software or driver recoverable faults" and "hardware-level unrecoverable faults" can be quickly distinguished, thereby achieving automatic isolation of fault domains and avoiding invalid driver reinstallation or system restart operations.
[0052] According to embodiments of the present invention, by having an in-band fault processor perform in-band processing of the device under test and an out-of-band fault processor perform out-of-band processing, a fault information processing architecture can be formed in collaboration, enabling dual-channel, zero-blind-spot fault detection. Based on a comprehensive judgment of multi-dimensional target conditions, if the in-band processor cannot successfully process the fault information, the fault information is sent to the out-of-band fault processor for processing. Since the fault information judgment is based on multi-dimensional target conditions, the target module to be updated can be accurately and quickly located. This allows the out-of-band fault processor to accurately execute a second update operation on the target module with hardware anomalies based on the fault information, improving fault information processing efficiency while achieving joint analysis of fault information and hardware. This forms a parallel in-band and out-of-band fault information processing mechanism, avoiding the blind spots of traditional single-channel detection, achieving comprehensive fault information detection and processing, satisfying fine-grained fault location accuracy, and further improving the system's reliability and availability.
[0053] In one feasible embodiment, to address situations where both automatic recovery at the software level (e.g., a reset mechanism within the driver) and routine recovery operations (e.g., restarting the application, uninstalling / reloading the kernel module) fail, a fault information processing system can perform real-time detection on the device under test. When an anomaly or crash occurs, the reset function of the device under test can be triggered, restoring the device from an abnormal state to a normal working state without requiring changes or repairs to the computing tasks. This achieves rapid fault isolation and functional recovery, maintains normal server operation, and enhances server stability.
[0054] According to an embodiment of the present invention, the processing module includes multiple sub-processing modules, and the exception information includes a first identifier of the exception sub-processing module and a second identifier of the target process; performing a first update operation on the target process in the processing module to obtain a first operation result includes: determining the target process from multiple processes in the processing module based on the first identifier and the second identifier, so as to stop the target process.
[0055] In embodiments of the present invention, the first identifier may be an identifier or code of a sub-processing module, and the second identifier may be an identifier or code of a process in an abnormal state. The abnormal information may be obtained by an in-band fault processor by comparing the detected state information with preset condition information.
[0056] For example, a process monitoring system can be used to dynamically monitor multiple processes in a processing module. Monitoring tools or custom scripts can be used to automatically identify GPU process identifiers (PIDs), generate an identifier list, dynamically update monitoring targets, and load the PID list through a file-based service discovery mechanism; alerts can be issued by monitoring process GPU utilization metrics, such as DCGM_FI_PROC_{PID}_GPU_UTIL=0 for 5 consecutive minutes.
[0057] For example, in the event of a process failure in the handling module, the in-band fault processor can combine the first and second identifiers to obtain combined information, and then trigger a reset signal to the in-band fault handling module based on the combined information. The in-band fault handling module then uses the reset signal (xx-smi-gpu-reset-i)<GPU_ID> -p <pid>terminating the target process.
[0058] According to the embodiment of the application, the mapping relationship from the "first identifier" to the "second identifier" can be driven internally, when an abnormal alarm or error of the identifier of the sub-processing module is monitored, the corresponding process identifier can be sent to the kernel for process termination through the mapping relationship, and the running of other processes is not affected, through the mapping relationship, the abnormal process can be located at the nanosecond level, and there is no need for user state polling, and zero external dependence is achieved.
[0059] According to the embodiment of the application, the fault information processing method further includes: in the case that the first operation result is abnormal and the state of the target process meets a preset state condition, generating an update instruction; and updating the plurality of processes based on the update instruction to obtain an update result.
[0060] In the embodiment of the application, the preset state condition can include that the target process is not interruptable and is in a deadlock state. The update instruction can be used to control the in-band fault handler to perform updating on the processing module.
[0061] In the case that the target process cannot be terminated to meet the fault processing, the state of the driver-level deadlock and the hardware-level downtime is updated through the update instruction, and the availability and recovery success rate in the fault condition are met.
[0062] For example, the processing module identifier of the stuck process and the target process identifier of the stuck process on the GPU can be found using a system management interface command (xx-smi); and the xx-smi --gpu-reset-i<GPU_ID>-p <pid>Precisely terminate a single process; if the process cannot be terminated normally, and the signal has been sent but the process has not exited, the entire GPU can be reset using xx-smi--gpu-reset-i<GPU_ID>#.
[0063] According to an embodiment of the present application, the fault information includes fault type information of each of the plurality of modules; and the fault information processing method further includes: determining a target module from the plurality of modules based on the fault type information, a module state of each of the plurality of modules, and association information of a module associated with the plurality of modules.
[0064] In an embodiment of the present application, the fault type information can include GPU failure, graphics memory failure, fan failure, interface failure, power supply module failure, and bus failure in hardware types. The module state can be determined according to the key log features of the plurality of modules obtained in the out-of-band fault processor. The association information can include a plurality of association information for the plurality of modules. The target module can be a module in an abnormal state in the plurality of modules.
[0065] For example, the fault code can be obtained from a sensor, a log, or a diagnostic protocol, so as to extract the fault type information, and the physical indicators and logical states of each hardware module are collected to obtain the module state. Considering the necessity of multi-source information fusion diagnosis, the module associated with the current hardware module can be determined through the physical connection, data dependency, and functional coupling relationship among the hardware modules, and the state information of the associated module is obtained. The way of determining the target module can include mapping the fault signal to a specific corresponding module by using a fault propagation strategy and a priority strategy.
[0066] For example, a mapping relationship between the fault type information and the module can be constructed. The fault mode information library can predefine the mapping rule between the fault type information and the module based on fault mode and effect analysis; in the case of multiple modules being abnormal at the same time, the priority can be dynamically adjusted based on the weight; when fault A occurs, the affected associated modules and association information can be searched along the association graph, in the case of multiple modules being abnormal at the same time, the module with higher priority can be selected based on the weight comparison method, and the fault module can be analyzed based on the time sequence analysis strategy; and then the target module is determined through cross verification of multi-source data (such as comparison between log analysis, abnormal module state, and associated module state).
[0067] According to an embodiment of the present application, by fusing the fault type, the module state (such as temperature, voltage, and register exception) itself, and the association information between the modules, the cascade false alarm can be excluded, and the real fault root cause module can be directly locked. Without manual traversal of all modules, the system can automatically narrow down the scope of investigation, shorten the traditional hours-level fault positioning to seconds, and improve the efficiency of fault positioning.
[0068] According to an embodiment of the present application, the plurality of modules comprises a transmission module, a management module and a connection module; determining the target module from the plurality of modules based on the fault type information, the module state of each of the plurality of modules and the association information of the modules associated with the plurality of modules comprises: determining the target transmission module from the plurality of modules based on the first encoding type for the transmission module in the fault type information, the transmission state of the transmission module and the first association information for the transmission module; determining the target management module from the plurality of modules based on the second encoding type for the management module in the fault type information, the management state of the management module and the second association information for the management module, wherein the second association information comprises at least one of sensor information, interface state and performance state; determining the target connection module from the plurality of modules based on the protocol fault type information for the connection module in the fault type information, the connection state of the connection module and the third association information for the connection module.
[0069] In an embodiment of the present application, the transmission module can comprise a power supply module, a power supply circuit and a power device. The management module can comprise a firmware storage medium, a controller and a communication interface. The connection module can comprise a data link layer controller, a configuration space and an interrupt controller. The target module can comprise a target transmission module, a target management module and a target connection module. The first encoding type can be an alarm code and a level identifier corresponding to the transmission module; the transmission state can be determined by monitoring key electrical parameters; and the first association information can be state information of associated modules having physical connection, data dependency and functional coupling relationship with the transmission module.
[0070] The second encoding type can be an error code and a key field corresponding to the management module, the management state can be determined by a state key field of the management firmware, and the second association information can be state information of associated modules having physical connection, data dependency and functional coupling relationship with the management module. The protocol fault type information can comprise a data link protocol type, a flow control protocol type and a transaction layer data packet error type. The connection state can be obtained by checking a Peripheral Component Interconnect Express (PCI Express) training state machine and an error counter. The third association information can be state information of associated modules having physical connection, data dependency and functional coupling relationship with the connection module.
[0071] For example, after the first encoding type of the transmission module is collected, an association graph analysis can be used to construct the electrical association relationship between the modules, and when the transmission module state is a fault state, even if other associated modules (such as processing modules) have errors, the power supply fault in the transmission module is prioritized to be located, so that the fault range can be gradually narrowed by the sectional isolation method, and only the core power supply is retained and the peripheral circuit interference is excluded.
[0072] For example, the management module can include a firmware management module and a software management module. The log of the management module can be directly read through the interface of the out-of-band fault processor, avoiding dependence on the operating system of the fault; an error processing callback can be registered in the bus link device driver, and the out-of-band fault processor alarm can be triggered when a link fault is detected.
[0073] For example, for the connection module, the dislocation type of the connection module can be located through a protocol analyzer, if the connection module reports an error and the driver does not respond, whether a physical layer fault exists can be confirmed based on cross-layer association verification, in combination with data obtained by the out-of-band fault processor sensor.
[0074] According to an embodiment of the present application, by fusing multi-source information such as a transmission module, a management module and a connection module, the out-of-band fault processor can be used to bypass the software layer to directly perform reset on the faulty hardware, meeting the requirement of zero dependence, millisecond level and high reliability of hardware fault self-healing through an out-of-band recovery path when the driver or operating system crashes, which can shorten the system downtime from minutes to seconds, and block the fault cascade diffusion.
[0075] According to an embodiment of the present application, the fault information further includes module identifiers of the plurality of modules and time information of generation of the fault; and the sending of the fault information to the out-of-band fault processor includes: encapsulating the module identifiers, the time information and the fault type to obtain encapsulation information; and sending the encapsulation information to the out-of-band fault processor based on a preset transmission strategy, wherein the preset transmission strategy includes a protocol conversion strategy, a zero processing strategy and a mapping strategy.
[0076] In an embodiment of the present application, the identifier of the module can include a module identifier and a corresponding error code. The time information can represent a timestamp of generation of the fault. The encapsulation information can be encapsulation information including the module identifier of the faulty module, the timestamp of generation of the fault and the fault type information. The protocol conversion strategy can represent a strategy of dynamic encapsulation of data by the in-band fault processor according to the type of the faulty device. The zero processing strategy can represent a strategy of supporting direct passage of critical data and hardware queue by design through an out-of-band channel (for example, an intelligent platform management interface). The mapping strategy can be a strategy of decoupling from a physical address by assigning a virtual identifier to each data source by the in-band fault processor; the out-of-band fault processor can maintain a mapping table in real time, and dynamically update the association relationship. The encapsulation information header can embed the mapping table, for fast analysis and positioning of the faulty module by the out-of-band processor.
[0077] For example, the fault identification module in the out-of-band fault processor can deliver the graphics card fault information to the fault processing module through an intelligent platform management interface (IPMI) protocol, and the information processing module in the fault processing module can trigger a reset signal to the out-of-band fault processor through a serial communication bus (I2C) according to the received fault type, and the out-of-band fault processor can send a corresponding reset command to the component to be reset according to the specific fault information.
[0078] After obtaining the encapsulation information, the out-of-band fault processor can parse the encapsulation information, which can include the following steps: after obtaining the encapsulation information, the header mark can be parsed, and when the header mark indicates a zero processing strategy, it can be directly transmitted to the target module; if it is a protocol conversion or mapping strategy, the mapping reference area can be read, the local mapping table can be queried, and the physical address and protocol driver can be obtained; the corresponding protocol stack can be called to send data to the target module.
[0079] According to the embodiments of the present application, the out-of-band channel can ensure that the fault information bypasses the operating system, and even if the system crashes or kernel deadlocks, the fault signal can still be transmitted through the hardware level channel, avoiding the delay or failure risk of traditional dependence on the driver. By triggering a hardware reset signal directly through a serial communication bus without software intervention, nanosecond-level fault response (such as instantaneous reset when a graphics card hangs) can be achieved, thereby reducing the average repair time.
[0080] According to the embodiments of the present application, the state information includes at least one of a running state, a process state, and a usage state; the above method further includes at least one of the following: in the case that the temperature state, the power consumption information, the read-write state, and the exception identifier in the running state meet the preset conditions, obtaining exception information, wherein the preset conditions include at least one of the following: the temperature state indicates that the temperature value of the processing module is greater than the preset temperature threshold, the power consumption information indicates that the power consumption value of the processing module is greater than the preset power consumption threshold, the read-write state indicates that there is a data abnormality in the read-write process of the processing module, and the exception identifier is that the out-of-band fault processor receives information indicating that the processing module has an exception; in the case that the process update frequency in the process state is less than a first frequency threshold, obtaining exception information; in the case that the usage frequency in the usage state is less than a second frequency threshold, or the fluctuation value of the usage frequency is greater than a fluctuation threshold, obtaining exception information.
[0081] In embodiments of the present application, the running state can represent a plurality of states of the processing module, including but not limited to temperature state, power consumption information, read-write state and exception identification. The process state can be the process identifier of the processing module identified by the monitoring tool or the custom script, and then the process information of the processing module is obtained in real time according to the monitoring information. The use state can be the utilization rate index of the processing module. The update frequency can be the heartbeat signal of a certain process.
[0082] For example, the detection program of the first fault identification module in the in-band fault processor can monitor whether the GPU computing task is normally performed in real time. The detection program can use an automated monitoring system. The automated monitoring system can include a collection module, a management module and a visualization module. The collection module can be used to manage and monitor the data center GPU in the cluster environment. The functions can include active running state monitoring, comprehensive diagnosis, system alarm and governance strategy. The hardware module can include power supply and clock management. By deploying the collection module container, the GPU indicators, GPU workload behavior or monitored GPU in the cluster can be collected in real time. The alarm rules are set in the visualization module, and the GPU process fault and GPU state exception are alarmed.
[0083] For example, the GPU state can be monitored based on the automated monitoring system, including whether the GPU temperature is too high, error alarm, etc. The monitoring tool can also be used to monitor the GPU health state indicators. For example, the GPU temperature can be represented as "DCGM_FI_DEV_GPU_TEMP", the GPU power consumption can be represented as "DCGM_FI_DEV_POWER_USAGE", the number of video memory errors can be represented as "DCGM_FI_DEV_ECC_SBE_VOL_TOTAL", and the error code can be represented as "DCGM_FI_DEV_XID_ERRORS".
[0084] For example, whether the GPU process is faulty can be determined by monitoring the running state of the GPU process or by the GPU usage. The monitoring tool can provide the GPU usage of each process. Assuming that there is a process and the process identifier of the process is known, the GPU usage of the process can be monitored, if the GPU usage of the process should be greater than 0 continuously under normal circumstances, when it is 0 continuously for a period of time, it can mean that the process is suspended, an alarm rule for the process can be created in the visualization module, the process state monitoring index can be represented as "CGM_FI_PROC_{PID}_GPU_UTIL", if DCGM_FI_PROC_GPU_UTIL{pid="GPU_pid"} =0 for 5 minutes, an alarm is given, the first fault identification module triggers a reset signal to the first fault processing module, and the first fault processing module performs a reset operation according to the error type.
[0085] According to an embodiment of the application, the plurality of modules comprises a power supply module; the fault information processing method further comprises: performing conversion processing on ripple information of the power supply module to obtain spectrum information; and determining state information of the processing module based on the spectrum information and an evaluation strategy, wherein the evaluation strategy indicates a mapping relationship between the spectrum information and a power supply fault type of the power supply module.
[0086] In an embodiment of the application, the power supply module can be a power supply rail of the processing module. The ripple information can comprise ripple characteristics and frequency components of the power supply rail. The evaluation strategy can be a power consumption spectrum analysis strategy, which is used to determine the mapping relationship between the spectrum information and the power supply fault type of the power supply module. The physical signal characteristics can be analyzed to break through the bottleneck of hardware monitoring.
[0087] For example, the obtained ripple information of the power supply module is decomposed into components to obtain switching ripple, high-frequency noise and transient drop characteristics; then, the switching ripple, the high-frequency noise and the transient drop characteristics are compared with respective threshold values to obtain ripple amplitude, frequency anomaly and transient response results; then, a time-domain waveform of every millisecond is identified, and the waveform is subjected to Fourier transform to obtain spectrum characteristics, and the fault characteristics are analyzed based on the spectrum characteristics, for example, if a 2.5 GHz frequency band attenuation is greater than 20%, it can be determined that the link layer signal integrity is degraded.
[0088] According to an embodiment of the present application, the plurality of modules comprises a temperature control module; the fault information processing method further comprises: performing feature extraction on the acoustic wave signal of the temperature control module to obtain initial features indicating the energy distribution of the target audio in the target frequency band; performing conversion and dimensionality reduction processing on the initial energy features to obtain intermediate features of the acoustic wave signal, the intermediate features having corresponding time frames; performing state evaluation on the temperature control module based on time sequence features obtained by sequentially weighting the intermediate features using the time frames to obtain an evaluation result, and determining the state information of the processing module based on the evaluation result.
[0089] In an embodiment of the present application, the temperature control module can be a fan or a heat sink in the device to be detected. The initial features can be a plurality of frequency features obtained through feature extraction and optimization processing. The intermediate features can be features after fault-sensitive feature enhancement on the initial features. The time sequence features can be short-time local features obtained after weighting the intermediate features. The evaluation result can be a large piece of information indicating the state of the processing module obtained through an evaluation model or a preset evaluation algorithm.
[0090] Considering that GPU fan bearing wear or heat sink dust accumulation can cause mechanical noise of a specific frequency (e.g., 4kHz-8kHz high-frequency whistling), the sound wave signal can be collected by deploying a microphone array, and initial signal enhancement and ambient noise suppression processing is performed on the sound wave signal, including beamforming on the array signal to focus on the sound source in the direction of the GPU fan, or collecting background noise when there is no GPU load, and subtracting the frequency band energy proportion from the real-time signal to obtain initial features; thereby realizing feature extraction and optimization through mel frequency cepstral coefficients.
[0091] For example, the mel band energy entropy can be calculated by using spectral entropy mel product to enhance the fault-sensitive features to obtain enhancement results (intermediate features); and then the enhancement results are classified and evaluated, and in the case of multiple types of enhancement results, the enhancement results can be weighted based on a multi-index fusion strategy to obtain an evaluation result, which can include the bearing wear degree and the heat sink dust accumulation degree of the temperature control module.
[0092] For example, the microphone array can be integrated into the in-band fault processor case, the evaluation model is associated with the out-of-band fault processor, non-intrusive monitoring is realized through voiceprint features, especially for predictive maintenance of high-density GPU clusters, online monitoring is realized without stopping and disassembling, and zero interference to actual business processing tasks is realized.
[0093] In a feasible embodiment, electromagnetic radiation of the GPU core can be captured by using a near-field probe, and whether the to-be-detected device has a fault can be judged based on the fact that normal computing units present a stable radiation mode under load, and memory faults or clock jitter will cause the radiation spectrum to spread. The sensor can be deployed on the server backplane, a baseline model of radiation is constructed, and deviation threshold alarm is realized.
[0094] For example, a sensor near-field probe array can be fixed on the server backplane, the radiation spectrum of the GPU core under standard load is collected, a ''stable radiation baseline model'' is generated, and is stored as a template library; the near-field probe continuously captures the near-field electromagnetic signal of the GPU at a preset sampling rate (for example, greater than or equal to 1 GHz), and outputs a real-time spectrum frame; the real-time spectrum is compared with the baseline model point by point, and the power density difference is calculated; if the deviation of any frequency point exceeds a preset threshold (for example, 3 dB), it can be determined as an abnormality; and then according to the abnormal frequency point distribution mode, it is mapped to the memory fault or clock jitter category, an alarm log is generated and reported to the operation and maintenance system.
[0095] According to the embodiments of the present application, through the analysis of the abnormal frequency point distribution mode, the fault source can be located more accurately, the misjudgment and missed judgment are reduced, and the operation and maintenance efficiency is improved; without any physical modification or disassembly of the GPU hardware, the monitoring can be completed only by the externally deployed sensor, which does not interfere with the normal operation of the GPU and does not affect the system performance and stability, and is suitable for online real-time monitoring.
[0096] According to the embodiments of the present application, the above method further includes: in the case that the second operation result is an abnormal result, updating the target module based on a preset update strategy to obtain an updated module, so as to execute the to-be-processed task by using the updated module.
[0097] In the embodiments of the present application, the preset update strategy can be a strategy for adaptively optimizing fault processing according to real-time system state and data characteristics, and can include a dynamic weighted aggregation strategy, an adaptive sparsification and lag compensation asynchronous strategy.
[0098] For example, when a single GPU in a certain supernode domain (for example, an 8-GPU node) in the to-be-detected device fails, the node is downgraded from a tensor parallel group to a complete data node. This can include: performing fault detection and resource isolation on the faulty tensor parallel group, and splitting and recombining the original tensor parallel group to obtain an updated normal subgroup and a downgraded node, and processing the task in the form of an independent copy, and reconstructing the communication group, so that the downgraded node participates in gradient synchronization, avoiding non-interlayer tensor communication; and then through a differential gradient synchronization and lag compensation mechanism, the normal tensor group synchronizes the gradient according to the original strategy, the downgraded node can be regarded as an independent copy, participates in global synchronization through weighted gradient aggregation, and the downgraded node can fall behind the normal tensor group due to completeness, and can avoid blocking through gradient buffering and asynchronous overlapping execution strategy.
[0099] It can be understood that, while the target module is subjected to the updating operation, the node where the target module is located can also be subjected to the updating operation.
[0100] According to the embodiment of the present application, by dynamically adjusting the gradient aggregation strategy, the low-throughput mode continues to participate in the training, avoiding the whole node to be discarded. Through the four-step strategy of fault isolation, topology degradation, dynamic gradient weighting, and asynchronous compensation, the fault node can be seamlessly converted from the tensor group to the degraded node, avoiding the whole node to be discarded, and the throughput loss is compressed to within 5%. The hardware failure is converted into soft degradation, and through the non-uniform parallel and resource elastic scheduling, the flexible coexistence of GPU failure is realized.
[0101] Figure 3 A flowchart of another fault information processing method according to an embodiment of the present application is shown.
[0102] As Figure 3 shown, the fault information processing method of this embodiment is applied to an out-of-band fault processor and includes operation S310.
[0103] In operation S310, in response to receiving the fault information sent by the in-band fault processor, a second updating operation is performed on the target module based on the fault information, and a second operation result is obtained. The fault information is obtained according to the fault information processing method applied to the in-band fault processor described above.
[0104] In the embodiment of the present application, the fault information can include temperature too high, power consumption too high, read / write error, and error code. For example, after the in-band fault processor performs the first updating operation on the target process, the first operation result is abnormal, and the module state of each module of the processing module is in an abnormal state, the fault information is sent to the out-of-band processing module through the out-of-band channel; the out-of-band processing module can perform a second updating operation on the target module according to the mapping relationship between the fault information and the updating operation corresponding to the target module, and obtain a second operation result.
[0105] Figure 4 A block diagram of a fault information processing system according to an embodiment of the present application is shown.
[0106] As Figure 4 As shown, according to different reset modes, the fault information processing system 400 can include an in-band fault processor 41, an out-of-band fault processor 42, and a device to be detected 43. The in-band fault processor 41 can include a first module 411 and a second module 412, the first module 411 can include an in-band identification module 411a and an in-band processing module 411b, and the second module 412 can include an information management module 412a and an information processing module 412b. The out-of-band fault processor 42 can include an out-of-band identification module 421 and an out-of-band processing module 422. The in-band fault processor 41 can be a processor associated with the device to be detected 43, and the out-of-band fault processor 42 can be a processor independent of the device to be detected 43.
[0107] Taking the device to be detected as a graphics card as an example, when the graphics card performs a large-scale cluster training / inference task, the in-band identification module 411a can detect, analyze and judge the graphics card fault type and send a reset signal to the in-band processing module 411b, and the in-band processing module 411b can reset the related registers according to the fault type. When the fault type cannot be repaired by the image processor reset function, the second module 412 of the model reports to the out-of-band fault processor 42 at this time, and the fault type is displayed in the system event log (System Event Log, SEL) or intelligent diagnostic log (Intelligent Diagnostic Log, IDL) and the solution is informed, for example, try to perform system restart, power-off restart, uninterrupted power supply restart or replace related operations on a certain slot.
[0108] The in-band identification module 411a can be composed of related detection programs. When the in-band identification module 411a detects a hardware fault, an unrecoverable error, a driver crash, an application error, a display server crash, a GPU hard lock, or a poor heat dissipation causing the GPU to trigger a self-protection mechanism, the in-band identification module 411a triggers an in-band reset signal to the in-band processing module 411b, and the in-band processing module 411b executes the reset command to try to restore the graphics card to a normal working state.
[0109] The in-band identification module 411a can classify the graphics card fault type into three categories, mainly including hardware failure, driver program exception, and application error. The hardware failure can include entering an unrecoverable error state due to power fluctuations, graphics card overheating (even if the thermal throttling threshold has not been reached but instability has been caused), graphics memory error, bus transient error, etc. The driver program exception can be that a defect in the driver program code causes the kernel to crash or fall into a state where it cannot continue to operate safely. The application error can be that the application causes the GPU to hang.
[0110] For example, when an application performs an illegal operation (e.g., illegal memory access, execution for an excessively long time, deadlock, resource exhaustion, etc.) during deep learning training, scientific computing, or running an unstable graphics program, the GPU context crashes and cannot be recovered internally by the driver.
[0111] The second module 412 is mainly composed of a graphics card hardware register, and the module complies with a hardware register protocol to allow the system baseboard management controller to communicate with the graphics card; when the in-band processing module 411b cannot process the graphics card fault, the information processing module 412b can be triggered to process, that is, when the system management interface command itself can run but cannot manage the target graphics card or when all the system management interface commands cannot run, an out-of-band reset mode can be used to recover the faulty graphics card.
[0112] The information management module 412a can be used to record and store graphics card fault information, and classify graphics card fault types according to components to be reset, and the fault types can include GPU fault, graphics memory fault, fan fault, input / output interface fault, external power supply module fault, and bus fault; the information processing module 412b triggers a reset signal to the out-of-band fault processor according to different fault types received from the information management module 412a, and the out-of-band fault processor sends a corresponding set command to the target module 431 through a bus.
[0113] According to an embodiment of the present application, the fault information includes fault type information; based on the fault information, a second update operation is performed on the target module, including: based on the fault type information and a mapping relationship between the fault type information and the update operation, an update instruction is generated; and based on the update instruction, the second update operation is performed on the target module.
[0114] In an embodiment of the present application, the out-of-band fault processor can internally store a mapping relationship table between fault type information and fault optimization strategies. The fault types can include transient faults and persistent faults. The fault optimization strategies can include a delayed reset strategy for transient faults and an immediate hard reset strategy for persistent faults. For firmware-level faults, the associated components (such as a GPU on-board management processor) can be reset first, and then the firmware image is re-flashed through a bus.
[0115] For example, for transient faults, in the case of sensor false alarms (such as transient jumps in fan speed) and communication noise interference, a 60-second observation window can be triggered, and if the fault is automatically recovered during the observation window, the reset is cancelled, otherwise a soft reset is performed.
[0116] For example, for persistent faults, in the case of abnormal power module output, complete fan stop, and GPU core overheating, a cold reset command (Cold Reset) can be directly sent to force the target component to power off and restart.
[0117] According to the embodiment of the present application, through the accurate mapping of the fault type and the reset strategy, the out-of-band fault processor can realize the optimization from the indiscriminate restart to the accurate reset, significantly reduce the average fault recovery time, and further improve the fault repair efficiency.
[0118] Based on the above fault information processing method, the present application further provides a fault information processing device. The following will be combined with the Figure 5 The device will be described in detail.
[0119] Figure 5 The structure block diagram of the fault information processing device according to the embodiment of the present application is shown.
[0120] As Figure 5 shown, the fault information processing device 500 of the embodiment is applied to an in-band fault processor, and includes a first updating module 510 and a sending module 520.
[0121] The first updating module 510 is configured to perform a first updating operation on a target process in the processing module to obtain a first operation result in response to detecting that the state information of the processing module in the to-be-detected device is abnormal information. In an embodiment, the first updating module 510 can be configured to perform the operation S210 described above, and details are not repeated here.
[0122] The sending module 520 is configured to send fault information to an out-of-band fault processor in a case that the first operation result is abnormal, and a module state of a plurality of modules associated with the processing module satisfies a preset state condition based on the detection instruction, so that the out-of-band fault processor performs a second updating operation on a target module in the plurality of modules based on the fault information to obtain a second operation result. The in-band fault processor is a processor associated with the to-be-detected device, the out-of-band fault processor is a processor independent of the to-be-detected device, and the preset state condition includes any one of the following: an instruction execution result of the detection instruction indicates that the processing module or the plurality of modules is abnormal; and the detection instruction is abnormal. In an embodiment, the sending module 520 can be configured to perform the operation S220 described above, and details are not repeated here.
[0123] According to the embodiment of the present application, through the first updating module 510 and the sending module 520 in the fault information processing device 500, the in-band processing by the in-band fault processor and the out-of-band processing by the out-of-band fault processor are cooperated to form a fault information processing architecture, and a double-channel and zero-blind-zone fault detection is formed, so that the fault information is sent to the out-of-band fault processor for processing under the condition that the in-band processor cannot successfully process the fault information based on the multi-dimensional target condition comprehensive judgment. Since the fault information is judged based on the multi-dimensional target condition, the target module to be updated can be accurately and quickly located, and then the out-of-band fault processor can accurately perform the second updating operation on the target module with hardware exception based on the fault information, thereby improving the fault information processing efficiency, realizing the joint analysis of the fault information and the hardware level, forming the in-band and out-of-band parallel fault information processing mechanism, avoiding the blind area of the traditional single-channel detection, realizing the comprehensive detection and processing of the fault information, meeting the fault fine-grained precision positioning, and further improving the reliability and usability of the system.
[0124] According to the embodiment of the present application, the processing module includes a plurality of sub-processing modules, the abnormal information includes a first identifier of an abnormal sub-processing module in the plurality of sub-processing modules and a second identifier of the target process; and the first updating module 510 includes: a process determination sub-module, configured to determine the target process from a plurality of processes of the processing module based on the first identifier and the second identifier, so as to stop the target process.
[0125] According to the embodiment of the present application, the device further includes: an instruction generation module and a process updating module. The instruction generation module is configured to generate an updating instruction in the case that the first operation result is abnormal and the state of the target process meets a preset state condition; and the process updating module is configured to update the plurality of processes based on the updating instruction to obtain an updating result.
[0126] According to the embodiment of the present application, the fault information includes fault type information of each of the plurality of modules; and the device further includes: a target module determination module, configured to determine the target module from the plurality of modules based on the fault type information, a module state of each of the plurality of modules, and association information of a module associated with the plurality of modules.
[0127] According to an embodiment of the present application, the plurality of modules comprises a transmission module, a management module and a connection module; the target module determining module comprises a transmission processing submodule, a management determining submodule and a connection determining submodule. The transmission processing submodule is configured to determine a target transmission module from the plurality of modules based on a first encoding type for the transmission module in the fault type information, a transmission state of the transmission module, and first associated information for the transmission module; the management determining submodule is configured to determine a target management module from the plurality of modules based on a second encoding type for the management module in the fault type information, a management state of the management module, and second associated information for the management module, wherein the second associated information comprises at least one of sensor information, interface state and performance state; and the connection determining submodule is configured to determine a target connection module from the plurality of modules based on protocol fault type information for the connection module in the fault type information, a connection state of the connection module, and third associated information for the connection module.
[0128] According to an embodiment of the present application, the fault information further comprises module identification of the plurality of modules and time information of the fault generation; and the sending module 520 comprises an encapsulation submodule and an information sending submodule. The encapsulation submodule is configured to encapsulate the module identification, the time information and the fault type to obtain encapsulation information; and the information sending submodule is configured to send the encapsulation information to the out-of-band fault processor based on a preset transmission strategy, wherein the preset transmission strategy comprises a protocol conversion strategy, a zero processing strategy and a mapping strategy.
[0129] According to an embodiment of the present application, the state information comprises at least one of a running state, a process state and a usage state; and the device further comprises at least one of a first information determining module, a second information determining module and a third information determining module. The first information determining module is configured to obtain abnormal information in a case where temperature state, power consumption information, read-write state and abnormality identification in the running state satisfy a preset condition, wherein the preset condition comprises at least one of the following: the temperature state indicates that a temperature value of the processing module is greater than a preset temperature threshold, the power consumption information indicates that a power consumption value of the processing module is greater than a preset power consumption threshold, the read-write state indicates that there is a data abnormality in a read-write process of the processing module, and the abnormality identification is that the in-band fault processor receives information indicating that the processing module has an abnormality; the second information determining module is configured to obtain abnormal information in a case where a process update frequency in the process state is less than a first frequency threshold; and the third information determining module is configured to obtain abnormal information in a case where a usage frequency in the usage state is less than a second frequency threshold or a fluctuation value of the usage frequency is greater than a fluctuation threshold.
[0130] According to an embodiment of the present application, the plurality of modules include a power supply module; the device further includes a conversion module and a state information determination module. The conversion module is configured to perform conversion processing on the ripple information of the power supply module to obtain spectrum information; the state information determination module is configured to determine state information of the processing module based on the spectrum information and an evaluation strategy, wherein the evaluation strategy indicates a mapping relationship between the spectrum information and a power supply fault type of the power supply module.
[0131] According to an embodiment of the present application, the plurality of modules include a temperature control module; the device further includes a feature extraction module, a dimension reduction module and a weighting module. The feature extraction module is configured to perform feature extraction on the acoustic wave signal of the temperature control module to obtain initial features indicating energy distribution of the target audio in the target frequency band; the dimension reduction module is configured to perform conversion and dimension reduction processing on the initial energy features to obtain intermediate features of the acoustic wave signal, the intermediate features having corresponding time frames; the weighting module is configured to perform state evaluation on the temperature control module based on time sequence features obtained by sequentially weighting the intermediate features using the time frames to obtain an evaluation result, and determine state information of the processing module based on the evaluation result.
[0132] According to an embodiment of the present application, the device further includes a target module updating module configured to, in a case where the second operation result is an abnormal result, update the target module based on a preset updating strategy to obtain an updated module, and execute the to-be-processed task by using the updated module.
[0133] Figure 6 A structural block diagram of another fault information processing device according to an embodiment of the present application is shown.
[0134] As Figure 6 shown, the fault information processing device 600 of this embodiment is applied to an out-of-band fault processor and includes a second updating module 610.
[0135] The second updating module 610 is configured to, in response to receiving fault information sent by the in-band fault processor, perform a second updating operation on the target module based on the fault information to obtain a second operation result, wherein the fault information is obtained according to the method described above. In an embodiment, the second updating module 610 can be configured to perform the operation S310 described above, and details are not described herein again.
[0136] According to an embodiment of the present application, the fault information includes fault type information; the second updating module 610 includes an updating instruction generation submodule and an execution submodule. The updating instruction generation submodule is configured to generate an updating instruction based on the fault type information and a mapping relationship between the fault type information and the updating operation; the execution submodule is configured to perform the second updating operation on the target module based on the updating instruction.
[0137] According to an embodiment of the present application, any of the first updating module 510 and the sending module 520, or the second updating module 610 can be combined in one module, or any of them can be split into multiple modules. Alternatively, at least part of the function of one or more of these modules can be combined with at least part of the function of other modules, and implemented in one module. According to an embodiment of the present application, at least one of the first updating module 510 and the sending module 520, or the second updating module 610 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on board, a system on package, an application specific integrated circuit (ASIC), or any other reasonable way of integrating or packaging a circuit, etc. hardware or firmware, or in any one of the three implementation ways of software, hardware and firmware, or in any appropriate combination of any of them. Alternatively, at least one of the first updating module 510 and the sending module 520, or the second updating module 610 can be at least partially implemented as a computer program module which, when executed, can perform the corresponding function.
[0138] Figure 7 A block diagram of an electronic device suitable for implementing the fault information processing method according to an embodiment of the present application is shown.
[0139] As shown in Figure 7 The electronic device 700 includes a memory 710 and a processor 720 configured to execute the fault information processing method according to the instructions and data stored in the memory.
[0140] In an embodiment of the present application, the processor 720 can include an in-band fault processor 41 and an out-of-band fault processor 42. The stored instructions and data include but are not limited to first updating operation instructions, second updating operation instructions, state information of processing modules, first operation results and second operation results.
[0141] Figure 8 A block diagram of another electronic device suitable for implementing the fault information processing method according to an embodiment of the present application is shown.
[0142] As shown in Figure 8 As shown, the electronic device 800 according to an embodiment of the present application includes a processor 801 which can perform various appropriate actions and processes in accordance with a program stored in a read only memory (ROM) 802 or a program loaded into a random access memory (RAM) 803 from a storage section 808. The processor 801 can include, for example, a general purpose microprocessor (e.g., a CPU), an instruction set processor, and / or a dedicated microprocessor (e.g., an application specific integrated circuit (ASIC)), and so on. The processor 801 can also include an on-board memory for cache use. The processor 801 can include a single processing unit or multiple processing units for performing the various actions of the method processes according to embodiments of the present application.
[0143] In the RAM 803, various programs and data required for the operation of the electronic device 800 are stored. The processor 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. The processor 801 performs various operations of the method processes according to embodiments of the present application by executing the programs in the ROM 802 and / or the RAM 803. Note that the programs can also be stored in one or more memories other than the ROM 802 and the RAM 803. The processor 801 can also perform various operations of the method processes according to embodiments of the present application by executing the programs stored in the one or more memories.
[0144] According to an embodiment of the present application, the electronic device 800 can further include an input / output (I / O) interface 805 which is also connected to the bus 804. The electronic device 800 can further include one or more of the following components connected to the input / output (I / O) interface 805: an input section 806 including a keyboard, a mouse, etc.; an output section 807 including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a LAN card, a modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the input / output (I / O) interface 805 as necessary. A removable recording medium 811 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is attached to the drive 810 as necessary, so that a computer program read out therefrom is installed in the storage section 808 as necessary.
[0145] The present application also provides a computer readable storage medium, which can be included in the device / apparatus / system described in the above embodiments; or can exist separately without being assembled into the device / apparatus / system. The above computer readable storage medium carries one or more programs, which when executed, implement the method according to the embodiments of the present application.
[0146] According to an embodiment of the present application, the computer readable storage medium can be a non-transitory computer readable storage medium, for example, can include but not limited to: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the present application, a computer readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present application, the computer readable storage medium can include the ROM 802 and / or the RAM 803 described above and / or one or more memory other than the ROM 802 and the RAM 803.
[0147] Embodiments of the present application also include a computer program product, which includes a computer program containing program codes for executing the method shown in the flow chart. When the computer program product is run in a computer system, the program codes are used to make the computer system implement the fault information processing method provided by the embodiments of the present application.
[0148] The above functions defined in the system / device of the embodiments of the present application are performed when the computer program is executed by the processor 801. According to an embodiment of the present application, the system, device, module, unit, etc. described above can be implemented by computer program modules.
[0149] In one embodiment, the computer program can rely on tangible storage media such as optical storage media, magnetic storage media, etc. In another embodiment, the computer program can also be transmitted, distributed, and downloaded in the form of signals on a network medium, and be downloaded and installed through the communication part 809, and / or installed from the detachable medium 811. The program codes contained in the computer program can be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the foregoing.
[0150] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 809, and / or installed from the detachable medium 811. When the computer program is executed by the processor 801, the above functions defined in the system of the embodiments of the present application are performed. According to an embodiment of the present application, the system, device, apparatus, module, unit, etc. described above can be implemented by computer program modules.
[0151] According to embodiments of the present application, program code for implementing the computer programs provided by embodiments of the present application can be written in any combination of one or more programming languages, and can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Program code can be executed entirely on a user's computing device, partially on a user's computing device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device through any kind of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., through an Internet service provider to the Internet).
[0152] The flow and block diagrams in the drawings represent possible architectural, functional, and operational architectures of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flow and block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may
[0153] Those skilled in the art will understand that features recited in the various embodiments of the present application can be combined and / or integrated in a variety of ways, even if such combinations or integrations are not expressly noted in the present application. In particular, features recited in the various embodiments of the present application can be combined and / or integrated in a variety of ways without departing from the spirit and teachings of the present application. All such combinations and / or integrations are within the scope of the present application.
[0154] The embodiments of the present application have been described above. However, these embodiments are merely for the purpose of illustration and are not intended to limit the scope of the present application. Although each of the embodiments has been described above, this does not mean that measures in each of the embodiments cannot be used advantageously in combination. Those skilled in the art can make various substitutions and modifications without departing from the scope of the present application, and these substitutions and modifications should fall within the scope of the present application.< / pid> < / pid>
Claims
1. A fault information processing method, characterized in that, Applied to an in-band fault processor, the method includes: In response to detecting that the status information of the processing module in the device under test is abnormal, a first update operation is performed on the target process in the processing module to obtain a first operation result. The device under test is a graphics card. The first update operation includes stopping the target process. If the first operation result, the fault type of the device under test, and the detection result based on the detection instruction meet the target conditions, fault information is sent to an out-of-band fault processor. This allows the out-of-band fault processor to perform a second update operation on the target module among multiple modules based on the fault information, obtaining a second operation result. The in-band fault processor is embedded in the device under test and is used to perform in-band fault processing on the software faults of the device under test. The out-of-band fault processor is a processor independent of the device under test and is used to perform in-band fault processing on the device under test if the first operation result, the fault type of the device under test, and the detection result based on the detection instruction meet the target conditions. Hardware faults in the detection equipment are handled out-of-band. The fault information includes fault type information, module identifier, and fault occurrence time information for each of the multiple modules. The fault information is sent to the out-of-band fault processor, which includes: encapsulating the module identifier, the time information, and the fault type information to obtain encapsulated information; and sending the encapsulated information to the out-of-band fault processor based on a preset transmission strategy. The preset transmission strategy includes a protocol conversion strategy for dynamically encapsulating data according to the fault type information, a zero-processing strategy that achieves key data pass-through and hardware queue support through out-of-band channel design, and a mapping strategy that allocates virtual identifiers to the data source to achieve decoupling from the physical address. The method further includes: when the result of the second operation is an abnormal result, updating the target module based on a preset update strategy to obtain an updated module, so as to use the updated module to execute the task to be processed, wherein the preset update strategy includes a dynamic weighted aggregation strategy, an adaptive sparsity strategy and a lag compensation asynchronous strategy. Wherein, the first operation result, the fault type of the device under test, and the detection result based on the detection command satisfy the target condition including at least one of the following: The result of the first operation indicates that the first update operation failed; The fault type indicates that the fault type of the device under test is a hardware fault type; The execution result of the detection command indicates that the processing module or the multiple modules are abnormal; The detection instruction is executed abnormally. The detection instruction is used to detect the module status of multiple modules associated with the processing module. The target condition is a collaborative judgment based on the first operation result, the fault type, and the detection instruction, to distinguish between software recoverable faults and hardware unrecoverable faults, thereby achieving automatic isolation of fault domains.
2. The method according to claim 1, characterized in that, The processing module includes multiple sub-processing modules, and the exception information includes a first identifier of the exception sub-processing module and a second identifier of the target process. Perform a first update operation on the target process in the processing module to obtain a first operation result, including: Based on the first identifier and the second identifier, the target process is determined from multiple processes of the processing module to stop the target process.
3. The method according to claim 2, characterized in that, The method further includes: If the result of the first operation is abnormal and the state of the target process meets the preset state conditions, an update instruction is generated. The multiple processes are updated based on the update instruction to obtain the update result.
4. The method according to claim 1, characterized in that, The method further includes: Based on the fault type information, the module status of each of the multiple modules, and the association information of the modules associated with the multiple modules, the target module is determined from the multiple modules.
5. The method according to claim 4, characterized in that, The multiple modules include a transmission module, a management module, and a connection module; Based on the fault type information, the module status of each of the multiple modules, and the association information of modules associated with the multiple modules, the target module is determined from the multiple modules, including: Based on the first encoding type used for the transmission module in the fault type information, the transmission status of the transmission module, and the first association information used for the transmission module, the target transmission module is determined from the plurality of modules; Based on the second coding type for the management module in the fault type information, the management status of the management module, and the second association information for the management module, a target management module is determined from the plurality of modules, wherein the second association information includes at least one of sensor information, interface status, and performance status; Based on the protocol fault type information for the connection module, the connection status of the connection module, and the third association information for the connection module, the target connection module is determined from the plurality of modules.
6. The method according to claim 1, characterized in that, The status information includes at least one of running status, process status, and usage status; The method further includes at least one of the following: When the temperature status, power consumption information, read / write status, and anomaly identifier in the operating state meet preset conditions, the anomaly information is obtained. The preset conditions include at least one of the following: the temperature status indicates that the temperature value of the processing module is greater than a preset temperature threshold; the power consumption information indicates that the power consumption value of the processing module is greater than a preset power consumption threshold; the read / write status indicates that there is a data anomaly in the processing module during the read / write process; and the anomaly identifier is that the in-band fault processor receives information indicating that there is an anomaly in the processing module. The abnormal information is obtained when the process update frequency in the process state is less than a first frequency threshold; The abnormal information is obtained when the usage frequency in the usage state is less than the second frequency threshold, or when the fluctuation value of the usage frequency is greater than the fluctuation threshold.
7. The method according to claim 1, characterized in that, The plurality of modules includes a power supply module; The method further includes: The ripple information of the power supply module is converted to obtain the spectrum information; Based on the spectrum information and the evaluation strategy, the status information of the processing module is determined, wherein the evaluation strategy indicates the mapping relationship between the spectrum information and the power supply fault type of the power supply module.
8. The method according to claim 1, characterized in that, The multiple modules include a temperature control module; The method further includes: Feature extraction is performed on the acoustic signal of the temperature control module to obtain initial features indicating the energy distribution of the target audio within the target frequency band; The initial energy features are transformed and dimensionality reduced to obtain intermediate features of the acoustic signal, and the intermediate features have corresponding time frames; Based on the time-series features obtained by weighting the intermediate features sequentially using the time frames, the temperature control module is evaluated to obtain an evaluation result, and the status information of the processing module is determined based on the evaluation result.
9. A fault information processing method, characterized in that, Applied to an out-of-band fault processor, the method includes: In response to receiving fault information sent by an in-band fault processor, a second update operation is performed on the target module based on the fault information to obtain a second operation result, wherein the fault information is obtained by the method according to any one of claims 1 to 8.
10. The method according to claim 9, characterized in that, The fault information includes fault type information; Based on the fault information, a second update operation is performed on the target module, including: Based on the fault type information and the mapping relationship between the fault type information and the update operation, an update instruction is generated; The second update operation is performed on the target module based on the update instruction.
11. An electronic device, characterized in that, include: Memory; A processor configured to execute the method according to any one of claims 1 to 10, based on instructions and data stored in the memory.
12. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 10.
13. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Server fault processing method and device, storage medium and electronic equipment
CN111694719A
Method and system for automatically repairing faults of battery swap station
CN115329878A
Fault processing system, fault processing method and related equipment
CN118312339A