Device Fault Handling Method, Electronic Device, Storage Medium, and Product

By analyzing the initial configuration of the processor and the connection method of the faulty equipment, different fault handling mechanisms are adopted to solve the problem of operation interruption caused by PCIe equipment failure, improve the accuracy and granularity of fault handling, and ensure the stability of the equipment.

CN120011127BActive Publication Date: 2025-07-25INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510489529.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-25
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

In the prior art, when the PCIe device fails, the processor uses a unified EDPC technology to handle the operation of the normal PCIe device, which lacks granularity and accuracy.

Method used

According to the fault handling mechanism supported by the processor, the processor is initialized and configured. Combined with the connection method between the fault device and the processor, different fault handling mechanisms are used for fault handling, including shielding or triggering port containment processing, generating fault reporting information, and using the message header information log to verify the source of the fault.

Benefits of technology

It improves the granularity and accuracy of fault handling, avoids the operation interruption of non-failed equipment, and ensures the stable operation of normal PCIe equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011127B_ABST
    Figure CN120011127B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method for processing device failures, an electronic device, a storage medium, and a product. The method includes: initializing and configuring a processor based on a failure handling mechanism supported by the processor; determining a connection mode between a failed device and the processor in response to failure information of the failed device; and performing failure handling on the failed device based on the connection mode and / or the initialization configuration. The method of the present disclosure configures different failure handling mechanisms for processing according to the support capabilities of the processor, and performs failure handling on the failed device using different failure handling mechanisms according to the connection relationship between the failed device and the processor, thereby avoiding the interruption of the operation of normal PCIe devices caused by using the EDPC technology for failure handling of different types of failure reports, and improving the granularity and accuracy of failure handling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computers, and in particular, to a method for processing device failures, an electronic device, a storage medium, and a product. Background Art

[0002] With the development of computer technology, Peripheral Component Interconnect Express (PCIe) devices have become an indispensable part of current computer systems. However, with the evolution of PCIe technology, the operating speed of PCIe devices is getting faster and faster, and their failures are becoming more and more frequent.

[0003] When a PCIe device reports a failure, the processor usually uses Enhanced Downstream Port Containment (EDPC) technology to handle different types of failure reports. This causes all PCIe devices to interrupt their connections to the processor when one or more of the PCIe devices connected to the processor fail, affecting the operation of normal PCIe devices. Summary of the Invention

[0004] The present disclosure provides a method for processing device failures, an electronic device, a storage medium, and a product. The method includes: initializing the configuration of the processor based on the failure handling mechanism supported by the processor; determining the connection mode between the failed device and the processor in response to the failure information of the failed device; and handling the failure of the failed device based on the connection mode and / or the initialization configuration. By configuring different failure handling mechanisms for the processor according to the support capabilities of the processor and using different failure handling mechanisms to handle the failures of failed devices according to the connection relationship between the failed devices and the processor, the present disclosure avoids the interruption of the operation of normal PCIe devices caused by using EDPC technology to handle different types of failure reports, and improves the processing granularity and accuracy of failure handling.

[0005] An embodiment of the first aspect of the present disclosure provides a method for processing device failures, including: initializing the configuration of the processor based on the failure handling mechanism supported by the processor; determining the connection mode between the failed device and the processor in response to the failure information of the failed device; and handling the failure of the failed device based on the connection mode and / or the initialization configuration.

[0006] In some embodiments of the present disclosure, based on the fault handling mechanism supported by the processor, the initialization configuration of the processor includes: when the processor supports root port programmable input / output, masking the first fault handling mechanism of the processor and configuring the second fault handling mechanism and / or the third fault handling mechanism for the processor; when the processor does not support root port programmable input / output, configuring the first fault handling mechanism for the processor.

[0007] In some embodiments of the present disclosure, the first fault handling mechanism includes: when a faulty device has a fault, triggering port containment processing; the second fault handling mechanism includes: when the fault of the faulty device is a non-support for a configuration space request, not triggering port containment processing, and when the fault of the faulty device is a fault other than a non-support for a configuration space request, triggering port containment processing; the third fault handling mechanism includes: when the fault of the faulty device is a completion timeout error, not triggering port containment processing, and when the fault of the faulty device is an error other than a completion timeout error, triggering port containment processing.

[0008] In some embodiments of the present disclosure, the connection mode between the faulty device and the processor includes: the faulty device is directly connected to the processor, and the faulty device is indirectly connected to the processor through a switch.

[0009] In some embodiments of the present disclosure, based on the connection mode and / or the initialization configuration, the fault handling of the faulty device includes: when the first fault handling mechanism is initialized and configured, triggering port containment processing; when the second fault handling mechanism or the third fault handling mechanism is initialized and configured, the fault handling of the faulty device is performed based on the connection mode.

[0010] In some embodiments of the present disclosure, when the second fault handling mechanism or the third fault handling mechanism is initialized and configured, the fault handling of the faulty device based on the connection mode includes: when the faulty device is directly connected to the processor, using the second fault handling mechanism to perform fault handling on the faulty device; when the faulty device is indirectly connected to the processor through a switch, using the third fault handling mechanism to perform fault handling on the faulty device.

[0011] In some embodiments of the present disclosure, when the faulty device is directly connected to the processor, using the second fault handling mechanism to perform fault handling on the faulty device includes: based on the fault information, determining whether the fault of the faulty device is a non-support for a configuration space request; based on the determination result, performing fault handling on the faulty device.

[0012] In some embodiments of the present disclosure, fault handling of a faulty device based on the result of the determination includes: when the result of the determination is that the fault of the faulty device is a non - supported configuration space request, generating first reporting information of the fault and not triggering port containment processing; when the result of the determination is that the fault of the faulty device is not a non - supported configuration space request, triggering port containment processing.

[0013] In some embodiments of the present disclosure, when the faulty device is indirectly connected to the processor through a switch, using a third fault handling mechanism, fault handling of the faulty device includes: based on the fault information, determining whether the fault of the faulty device is a completion timeout error; based on the result of the determination, performing fault handling on the faulty device.

[0014] In some embodiments of the present disclosure, fault handling of a faulty device based on the result of the determination includes: when the result of the determination is that the fault of the faulty device is not a completion timeout error, triggering port containment processing; when the result of the determination is that the fault of the faulty device is a completion timeout error, generating second reporting information of the fault and obtaining the message header information log corresponding to the fault, so as to determine whether to perform port containment processing based on the message header information log.

[0015] In some embodiments of the present disclosure, determining whether to perform port containment processing based on the message header information log includes: based on the type of the message header information log, determining the space address indicated by the message header information log; determining whether the device corresponding to the space address is the faulty device to determine whether to perform port containment processing.

[0016] In some embodiments of the present disclosure, determining the space address indicated by the message header information log based on the type of the message header information log includes: when the type of the message header information log is a memory read - write type or an input - output read - write type, based on the information address of the message header information log, determining the space address indicated by the message header information log, so as to determine a first information set based on the space address; when the type of the message header information log is a configuration space read - write type, based on the information address offset of the message header information log, determining the first information set.

[0017] In some embodiments of the present disclosure, determining whether to perform port containment processing by determining whether the device corresponding to the space address is the faulty device includes: when it is determined that the device corresponding to the first information set is the faulty device, or when it is determined that the device corresponding to the first information set is the switch connected to the faulty device, not performing port containment processing; when it is determined that the device corresponding to the first information set is not the faulty device and the device corresponding to the first information set is not the switch connected to the fault, performing port containment processing.

[0018] In some embodiments of the present disclosure, the method further includes: determining the type of the message header information log based on the memory - mapped space applied for by the faulty device.

[0019] In some embodiments of the present disclosure, the first information set includes at least one of the following: bus information of the faulty device, device information of the faulty device, and function information of the faulty device.

[0020] In some embodiments of the present disclosure, the method further includes: determining the device identity information of the faulty device and the cause of the fault of the faulty device based on the first reported information or the second reported information; and performing fault handling on the faulty device by using a preset solution based on the device identity information and the cause of the fault.

[0021] In some embodiments of the present disclosure, the port containment process includes: uninstalling the driver of the faulty device and removing the device label of the faulty device; releasing the link state between the faulty device and the processor; enumerating the faulty device, and re-enabling the driver of the faulty device.

[0022] An embodiment of the second aspect of the present disclosure provides an electronic device, including: a processor and a memory for storing a computer program that can run on the processor, wherein the processor is configured to execute the method described in the embodiment of the first aspect of the present disclosure when running the computer program.

[0023] An embodiment of the third aspect of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method described in the embodiment of the first aspect of the present disclosure.

[0024] An embodiment of the fourth aspect of the present disclosure provides a computer program product, including a computer program that implements the method described in the embodiment of the first aspect of the present disclosure when executed by a processor.

[0025] In summary, according to a device fault handling method provided by the present disclosure, it includes: initializing and configuring the processor based on a fault handling mechanism supported by the processor; determining a connection mode between the faulty device and the processor in response to fault information of the faulty device; and performing fault handling on the faulty device based on the connection mode and / or the initialization configuration. The method of the present disclosure configures different fault handling mechanisms for the processor according to the support capabilities of the processor, and uses different fault handling mechanisms to perform fault handling on the faulty device according to the connection relationship between the faulty device and the processor, avoiding the interruption of the operation of normal PCIe devices caused by using the EDPC technology for fault handling of different types of fault reports, and improving the handling granularity and accuracy of fault handling.

[0026] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The accompanying drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0028] Figure 1 is a schematic flowchart of a device fault handling method provided by an embodiment of the present disclosure;

[0029] Figure 2 is a schematic diagram of a scenario of a device fault handling method provided by an embodiment of the present disclosure;

[0030] Figure 3 is a schematic diagram of a scenario of another device fault handling method provided by an embodiment of the present disclosure;

[0031] Figure 4 is a schematic flowchart of a device fault handling method provided by an embodiment of the present disclosure;

[0032] Figure 5 is a schematic flowchart of another device fault handling method provided by an embodiment of the present disclosure;

[0033] Figure 6 is a schematic flowchart of yet another device fault handling method provided by an embodiment of the present disclosure;

[0034] Figure 7 is a schematic flowchart of still another device fault handling method provided by an embodiment of the present disclosure;

[0035] Figure 8 is a schematic flowchart of another device fault handling method provided by an embodiment of the present disclosure;

[0036] Figure 9 is a schematic diagram of a process example of yet another device fault handling method provided by an embodiment of the present disclosure;

[0037] Figure 10 is a schematic diagram of a process example of a device fault handling method provided by an embodiment of the present disclosure;

[0038] Figure 11 is a schematic structural diagram of another device fault handling device provided by an embodiment of the present disclosure;

[0039] Figure 12 is a schematic diagram of the hardware composition structure of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners

[0040] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted in the following description for clarity and conciseness.

[0041] With the development of computer technology, PCIe devices have become an indispensable part of current computer systems. However, with the evolution of PCIe technology, the operating speed of PCIe devices is getting faster and faster, and their failures are becoming more and more frequent. For some fault problems, the corresponding operating system of the PCIe device can be automatically repaired, but most problems cannot be automatically repaired by the operating system, resulting in machine downtime or restart, so that the operating system cannot run normally.

[0042] To solve the system downtime caused by the fatal faults of PCIe devices, the EDPC technology has emerged in the industry. The EDPC technology is an error isolation and recovery technology for the PCIe bus. When a link error (such as a transaction layer data packet with a format error, an unexpected halt, etc.) is detected, it disables a single PCIe link and forcibly terminates the outstanding requests, thereby avoiding error propagation and protecting the operating system from potential bad data effects.

[0043] Specifically, in the related art, the EDPC is usually triggered by the Advanced Error Reporting (AER) mechanism of PCIe. AER is a mechanism for detecting and reporting errors that occur in PCle devices. It allows PCle devices to detect and report various types of errors, such as non-fatal, recoverable, and severe errors. AER implements a set of registers and corresponding error notification mechanisms on PCle devices, and information about errors can be obtained by reading these registers. Using AER, the system can better monitor and handle the error conditions of PCle devices to improve data integrity and reliability.

[0044] Specifically, the cooperation of EDPC and AER can implement the isolation and repair functions of PCIe devices, such as Figure 1 The method includes the following steps as shown:

[0045] ① The processor detects an uncorrectable error in the device (i.e., an error that the operating system cannot automatically repair). ② The processor sends a signal to the operating system for error handling. ③ The operating system battery management module notifies the device driver to uninstall the driver of the device connected to the processor; and removes the device label of the device to prevent subsequent access to the Memory-Mapped Input / Output (MMIO) space. ④ The operating system battery management module restores the device through the following operations: releases the (Link) link state of the device, re-enumerates the device, and re-enables the device driver of the device.

[0046] The above related technologies can accurately isolate and repair a faulty device when the PCIe root port of the processor corresponds to a single device. However, when the PCIe root port corresponds to multiple devices, that is, when the device corresponding to the PCIe root port is a PCIe expansion (PCIe Switch) device (such as Figure 2 shown), when one of the multiple devices fails (for example, device 1 fails), the PCIe switch 1 triggers a DPC, and it will cause the PCIe root port to accompany a timeout fault, thereby causing the processor to trigger a DPC, resulting in the interruption of the normal operation of the device.

[0047] And during the enumeration process of using DPC to handle device faults, since the device has not been initialized yet, some PCIE configuration (configuration, cfg) instructions cause UnsupportedRequest (UR) errors to occur. In related technologies, the UR errors are usually masked to avoid false alarms. However, when the operating system completes the enumeration of the device and runs normally, due to masking the PCIE UR errors, more serious errors will occur in the processor.

[0048] The following explains the technical terms related to the present disclosure:

[0049] Root Port Input / Output (Root Port PIO, RP PIO) is a mechanism in the PCIe architecture for managing errors encountered when the root port (Root Port) sends non-posted requests. It provides fine control over uncorrectable errors and advisory errors. The RP PIO error control register provides fine error management capabilities for non-POSTED requests, allowing flexible configuration according to the request type (configuration, input / output, memory) and error type.

[0050] To solve the technical problems existing in the related art, embodiments of the present disclosure provide a method for processing device failures.

[0051] First, the application scenarios related to the present disclosure will be exemplarily described below:

[0052] As Figure 3 shown, the processor is directly connected to Device 1 through a PCIe port, Device 2 is connected to the PCIe port of the processor through PCIe Switch 1, and Device 3 is connected to the PCIe port of the processor through PCIe Switch 2.

[0053] It should be understood that Figure 3 only for illustrative purposes, the application scenarios of the present disclosure may include some or all of the devices as Figure 3 shown, and the present disclosure does not limit the number of devices directly connected to the processor, nor the number of devices connected to the processor through a PCIe switch. For example, the application scenario of the present disclosure may only include one or more devices directly connected to the processor; or only include one or more devices indirectly connected to the processor through a PCIe switch; or include one or more devices directly connected to the processor and one or more devices indirectly connected to the processor.

[0054] Next, the embodiments of the present disclosure will be introduced in detail.

[0055] As Figure 4 shown, embodiments of the present disclosure provide a method for processing device failures, including the following steps:

[0056] Step 101, perform initialization configuration on the processor based on the fault handling mechanism supported by the processor.

[0057] In some embodiments, different initialization configurations can be performed on the processor according to whether the processor supports Root Port Programmable I / O (RP PIO).

[0058] In some embodiments, when the processor does not support Root Port Programmable I / O, the initialization configuration of the processor can be to trigger port containment processing when a faulty device has a fault, that is, regardless of the fault type of the faulty device, the port containment processing method is uniformly used to process the faulty device.

[0059] In some embodiments, when the processor supports Root Port Programmable I / O, the initialization configuration of the processor can adopt port containment processing for some fault types and only report faults without adopting port containment processing for some fault types, so as to implement different fault handling methods for different fault types to avoid affecting the operation of normal devices.

[0060] In some embodiments, the port containment process can be an EDPC process, a downstream port containment (DPC) process, or an improved processing method based on the DPC process, etc. The present disclosure does not limit this.

[0061] Step 102: In response to the fault information of the faulty device, determine the connection mode between the faulty device and the processor.

[0062] In some embodiments, the fault information is used to indicate that a device connected to the processor has failed.

[0063] In some embodiments, when the fault information is received, the processor can query the connection information of the faulty device to determine the connection mode between the faulty device and the processor, and then, in combination with the initialization configuration of the processor, perform targeted fault handling on the faulty device. However, this is not limited thereto, and the present disclosure does not limit the technical means used to determine the connection mode.

[0064] In some embodiments, the connection modes between the faulty device and the processor include: the faulty device is directly connected to the processor, and the faulty device is indirectly connected to the processor through a switch. The switch can be a PCIe switch.

[0065] Specifically, the direct connection between the faulty device and the processor can be a direct connection between the faulty device and the PCIe root port of the processor; the indirect connection between the faulty device and the processor through a PCIe switch can be that the faulty device is connected to the PCIe root port of the processor through a PCIe switch.

[0066] Step 103: Perform fault handling on the faulty device based on the connection mode and / or initialization configuration.

[0067] In some embodiments, the initialization configuration is that when a faulty device has a fault, port containment processing is triggered. At this time, regardless of whether the connection mode between the faulty device and the processor is a direct connection or an indirect connection, the port containment processing method is used to perform fault handling on the faulty device, that is, directly perform fault handling on the faulty device based on the initialization configuration.

[0068] In some embodiments, the initialization configuration of the processor is that some fault types use port containment processing, and some fault types do not use port containment processing but only perform fault reporting. Since the fault types that require port containment processing are different for different connection modes, it is necessary to combine the connection mode between the faulty device and the processor to determine whether the fault type of the faulty device requires port containment processing to perform fault handling on the faulty device.

[0069] In some embodiments, the port containment process may include unloading the driver of the faulty device and removing the device label of the faulty device; releasing the link state between the faulty device and the processor; enumerating the faulty device, and re-enabling the driver of the faulty device, so as to isolate and repair the faulty device.

[0070] In summary, the device fault handling method proposed according to the present disclosure includes: initializing and configuring the processor based on the fault handling mechanism supported by the processor; determining the connection mode between the faulty device and the processor in response to the fault information of the faulty device; and performing fault handling on the faulty device based on the connection mode and / or the initialization configuration. The method of the present disclosure configures different fault handling mechanisms for the processor according to the support capabilities of the processor, and performs fault handling on the faulty device using different fault handling mechanisms according to the connection relationship between the faulty device and the processor, avoiding the interruption of the operation of normal PCIe devices caused by using the EDPC technology for fault handling of different types of fault reports, and improving the processing granularity and accuracy of fault handling.

[0071] Figure 5 Further shown is a flowchart of a device fault handling method proposed by the present disclosure. Based on Figure 1 the embodiments shown are further explained, Figure 5 it may include the following steps.

[0072] Step 201, when the processor supports root port programmable input / output, mask the first fault handling mechanism of the processor, and configure the second fault handling mechanism and / or the third fault handling mechanism for the processor.

[0073] In some embodiments, the first fault handling mechanism includes: when a faulty device has a fault, triggering the port containment process.

[0074] In some embodiments, the second fault handling mechanism includes: when the fault of the faulty device is a non-support configuration space request, not triggering the port containment process, and when the fault of the faulty device is a fault other than the non-support configuration space request, triggering the port containment process.

[0075] In some embodiments, the third fault handling mechanism includes: when the fault of the faulty device is a completion timeout error, not triggering the port containment process, and when the fault of the faulty device is an error other than the completion timeout error, triggering the port containment process.

[0076] In some embodiments, the second fault handling mechanism is used when the faulty device is directly connected to the processor, and the third fault handling mechanism is used when the faulty device is indirectly connected to the processor through a PCIe switch.

[0077] In some embodiments, a non-supported configuration space request may be an error in the UR of the configuration space request, but is not limited thereto.

[0078] In some embodiments, a Completion Timeout error may be a timeout of a configuration space request, a timeout of an input / output request, a timeout of a memory request, etc., and the present disclosure does not limit this.

[0079] In some alternative embodiments, when all devices connected to the processor are directly connected, only a second fault handling mechanism may be configured for the processor; when all devices connected to the processor are indirectly connected, only a third fault handling mechanism may be configured for the processor; when there are both direct connections and indirect connections among the devices connected to the processor, both a second fault handling mechanism and a third fault handling mechanism may be configured for the processor.

[0080] Step 202, when the processor does not support root port programmable input / output, configure a first fault handling mechanism for the processor.

[0081] In some embodiments, when the processor does not support root port programmable input / output, that is, when the processor does not support different fault handling methods according to the fault type, a first fault handling mechanism may be directly configured for the processor so that fault handling can be performed when a device connected to the processor fails.

[0082] In summary, the device fault handling method proposed according to the present disclosure includes: when the processor supports root port programmable input / output, shielding the first fault handling mechanism of the processor and configuring a second fault handling mechanism and / or a third fault handling mechanism for the processor; when the processor does not support root port programmable input / output, configuring a first fault handling mechanism for the processor. The method of the present disclosure configures different fault handling mechanisms for the processor according to the capabilities of the processor, so as to improve the granularity of processor fault handling and the accuracy of fault handling.

[0083] Figure 6 Further shows a flowchart of a device fault handling method proposed by the present disclosure. Based on Figure 1 and Figure 2 The embodiments shown are further explained. Figure 6 It may include the following steps.

[0084] Step 301, when the processor does not support root port programmable input / output, configure a first fault handling mechanism for the processor.

[0085] In some embodiments, the principle of step 301 is the same as that of step 201, and reference may be made to the embodiments and related descriptions in step 201, which will not be elaborated here.

[0086] Step 302, in response to the fault information of the faulty device, trigger port containment processing.

[0087] In some embodiments, since the initialization configuration of the processor is the first fault handling mechanism, when the fault information is received, the processor is not capable of adopting different fault handling methods according to different fault types. Therefore, the processor can directly perform port containment processing on the faulty device to handle the fault.

[0088] In summary, the device fault handling method proposed according to the present disclosure includes: when the processor does not support root port programmable input / output, configure the first fault handling mechanism for the processor; in response to the fault information of the faulty device, trigger port containment processing. By configuring the first processing mechanism for the processor that does not support root port programmable input / output, the method of the present disclosure enables fault handling to be achieved by using port containment processing when a device directly or indirectly connected to the processor fails, thereby improving the applicable scope of the present disclosure.

[0089] Figure 7 Further shown is a flowchart of a device fault handling method proposed by the present disclosure. Based on Figure 1 and Figure 2 the embodiments shown, further explained Figure 7 it may include the following steps:

[0090] Step 401, when the processor supports root port programmable input / output, mask the first fault handling mechanism of the processor and configure the second fault handling mechanism for the processor, or configure the second fault handling mechanism and the third fault handling mechanism for the processor.

[0091] In some embodiments, the principle of step 401 is the same as that of step 201. The embodiments and related descriptions in step 201 can be referred to, and will not be elaborated here.

[0092] Step 402, in response to the fault information for the faulty device, determine that the faulty device is directly connected to the processor.

[0093] In some embodiments, when the fault information is received, the processor can query the connection information of the faulty device to determine that the faulty device is directly connected to the processor, but it is not limited thereto. The present disclosure does not limit the technical means used to determine the connection method.

[0094] Step 403, use the second fault handling mechanism to handle the fault of the faulty device.

[0095] In some embodiments, the fault information can also indicate the fault type of the faulty device. Therefore, it is possible to determine whether the fault of the faulty device is a non-support for configuration space requests according to the fault information, and then based on the judgment result, handle the fault of the faulty device.

[0096] Specifically, when it is determined according to the fault information that the fault of the faulty device is an unsupported configuration space request, the processing method for the faulty device at this time is: generating the first reporting information of the fault and not triggering port containment processing.

[0097] When it is determined according to the fault information that the fault of the faulty device is not an unsupported configuration space request (i.e., other faults other than the unsupported configuration space request), the processing method for the faulty device at this time is: triggering port containment processing.

[0098] Specifically, when the device is directly connected to the processor and the processor enumerates the device, since a normal device that has not been initialized completely will cause a fault of an unsupported configuration space request to occur (enumeration is used for the processor to read or write data from or to the device, and since the device has not been initialized completely, the processor cannot read or write data, so a fault of an unsupported configuration space request will occur), and at this time the normal device has not malfunctioned, but just has not been initialized completely, so this normal device does not need to perform port containment processing.

[0099] Furthermore, in order to avoid that the fault of an unsupported configuration space request is not caused by a normal device not being initialized completely, that is, the device actually malfunctions and causes an unsupported configuration space request, it is necessary to generate the first reporting information to analyze the specific cause of the unsupported configuration space request, so as to determine whether to perform corresponding fault handling.

[0100] In other words, according to the first reporting information, the cause of the fault of an unsupported configuration space request can be analyzed, so that when the unsupported configuration space request is not caused by a normal device not being initialized completely, the preset fault handling method is used to perform fault handling on the faulty device, so as to realize the isolation and repair of the faulty device.

[0101] In some embodiments, the first reporting information may be fault advisory reporting (Advisory Errorreport) information, but is not limited thereto.

[0102] In summary, the device fault handling method proposed according to the present disclosure includes: when the processor supports root port programmable input / output, shielding the first fault handling mechanism of the processor, and configuring a second fault handling mechanism for the processor, or configuring a second fault handling mechanism and a third fault handling mechanism for the processor; in response to fault information for a faulty device, determining that the faulty device is directly connected to the processor; and using the second fault handling mechanism to handle the fault of the faulty device. The method of the present disclosure, by configuring a second processing mechanism for the processor when the processor supports root port programmable input / output, enables different types of faults to be handled differently when a device directly connected to the processor fails. Thus, when the fault is a configuration space request that is not supported, only reporting information is generated, and port containment processing is not triggered, avoiding normal devices that have not completed initialization from triggering fault handling and ensuring the normal operation of the device.

[0103] Figure 8 Further, a flowchart showing a device fault handling method proposed by the present disclosure is presented. Based on Figure 1 and Figure 2 the embodiments shown, it is further explained that Figure 8 it may include the following steps:

[0104] Step 501: When the processor supports root port programmable input / output, shield the first fault handling mechanism of the processor, and configure a third fault handling mechanism for the processor, or configure a second fault handling mechanism and a third fault handling mechanism for the processor.

[0105] In some embodiments, the principle of step 501 is the same as that of step 201. Reference can be made to the embodiments and related descriptions in step 201, and details are not repeated here.

[0106] Step 502: In response to fault information for a faulty device, determine that the faulty device is indirectly connected to the processor.

[0107] In some embodiments, when receiving the fault information, the processor can determine that the faulty device is indirectly connected to the processor by querying the connection information of the faulty device, but this is not limited thereto. The present disclosure does not limit the technical means used to determine the connection method.

[0108] Step 503: Use the third fault handling mechanism to handle the fault of the faulty device.

[0109] In some embodiments, the fault information may further indicate the fault type of the faulty device. Therefore, it is possible to determine whether the fault of the faulty device is a completion timeout error based on the fault information, and then handle the fault of the faulty device based on the judgment result.

[0110] Specifically, when it is determined according to the fault information that the fault of the faulty device is not a completion timeout error, the processing method for the faulty device at this time is: trigger port containment processing.

[0111] When it is determined according to the fault information that the fault of the faulty device is a completion timeout error, generate a second report message of the fault, and obtain the message header information log corresponding to the fault, so as to determine whether to perform port containment processing based on the message header information log (RP PIO Header Log).

[0112] Furthermore, the first information set indicated by the message header information log can be determined according to the type of the message header information log; then, it is judged whether the device corresponding to the first information set is the faulty device, so as to determine whether to perform port containment processing.

[0113] Specifically, when a faulty device indirectly connected to the processor fails, the PCIe switch connected to the faulty device will first trigger port containment processing to handle the fault of the faulty device. At this time, the PCIe switch will send a fault message to the processor. Since the PCIe switch has triggered port containment processing and the faulty device is currently being fault processed, in order to prevent the processor from performing port containment processing again because of this fault message, it is necessary to compare the source of the fault message to determine whether the device that sent the fault message is the device that has triggered port containment processing.

[0114] Specifically, when the type of the message header information log is a memory read / write type or an input / output read / write type, based on the information address of the message header information log, determine the spatial address indicated by the message header information log, so as to determine the first information set based on the spatial address; when the type of the message header information log is a configuration space read / write type, determine the first information set based on the information address offset of the message header information log.

[0115] In some embodiments, the type of the message header information log can be determined based on the memory mapping space applied by the faulty device, but this is not limited thereto. The present disclosure does not limit the method for determining the type of the message header information log.

[0116] In some embodiments, when the type of the message header information log is a memory read / write type or an input / output read / write type, the message header information log can directly indicate the spatial address of the device where the fault occurs, and then determine the device where the fault occurs according to the spatial address, so as to obtain the first information set.

[0117] In some embodiments, when the type of the message header information log is a configuration space read / write type, the data at different positions of the message header information can indicate the first information set of the device where the fault occurs.

[0118] In other words, when the type of the packet header information log is of the memory read / write type or the input / output read / write type, it is necessary to determine the specific faulty device according to the spatial address indicated by the packet header information log, so as to determine the first information set; while when the type of the packet header information log is of the configuration space read / write type, the packet header information log can directly indicate the first information set, so the packet header information log can be directly analyzed to determine the first information set.

[0119] For example, taking the packet header information log as the 9-bit memory read type information as an example, the first 3 bits can be used to indicate the bus information, the middle 3 bits can indicate the device information, and the last 3 bits can indicate the function information.

[0120] Furthermore, the first information set includes at least one of the following: the bus information of the faulty device, the device information of the faulty device, and the function information of the faulty device.

[0121] In some alternative embodiments, it is also possible to determine whether the device corresponding to the judgment spatial address is a faulty device, so as to determine whether the fault information is generated due to the port containment process triggered by the faulty device, so as to avoid the fault information generated due to the port containment process from triggering the port containment process again.

[0122] Furthermore, when it is determined that the device corresponding to the first information set is a faulty device, or it is determined that the device corresponding to the first information set is a PCIe switch connected to the faulty device, the port containment process is not executed; when it is determined that the device corresponding to the first information is not a faulty device, and the device corresponding to the first information set is not a PCIe switch connected to the faulty device, the port containment process is executed.

[0123] In other words, when the device corresponding to the first information set is a faulty device, or it is determined that the device corresponding to the first information set is a PCIe switch connected to the faulty device, it means that the above-mentioned fault information is generated by the PCIe switch or the faulty device when the faulty device is being fault-processed by using the port containment process. At this time, the faulty device is already being fault-processed, and there is no need for the processor to perform the port containment process again, so the port containment process is not executed.

[0124] In other words, when the device corresponding to the first information set is not a faulty device, or it is determined that the device corresponding to the first information set is not a PCIe switch connected to the faulty device, it means that the above-mentioned fault information is not generated when the faulty device is being fault-processed by using the port containment process, that is, there is a faulty device that needs to be fault-processed currently, so the port containment process is executed.

[0125] For example, Figure 3Taking the devices included as an example, when device 2 sends a fault message to the processor, the register in the processor parses the message header information log. When it is determined that the fault source indicated by the message header information log is device 2 or PCIe switch 1, port containment processing is not performed, and this fault handling ends.

[0126] Exemplarily, taking Figure 3 the devices included as an example, when device 2 sends a fault message to the processor, the register in the processor parses the message header information log. When it is determined that the fault source indicated by the message header information log is not device 2 or PCIe switch 1 (such as device 1, device 3, PCIe switch, etc.), port containment processing is performed.

[0127] In some embodiments, the second reported information may be fault advisory reported information, but is not limited thereto.

[0128] In some alternative embodiments, the device identity information of the faulty device and the fault cause of the faulty device may also be determined based on the first reported information or the second reported information; and then, based on the device identity information and the fault cause, a preset solution is used for fault handling.

[0129] Specifically, the first reported information or the second reported information may include information such as the device model and fault cause of the faulty device. Then, according to a preset mapping relation table, the preset solution corresponding to the information such as the device model and fault cause in the first reported information or the second reported information is found by looking up the table, so as to use the corresponding preset solution for fault handling.

[0130] In summary, the device fault handling method proposed according to the present disclosure includes: when the processor supports root port programmable input / output, shielding the first fault handling mechanism of the processor, and configuring a third fault handling mechanism for the processor, or configuring a second fault handling mechanism and a third fault handling mechanism for the processor; in response to the fault information for the faulty device, determining that the faulty device is directly connected to the processor; using the third fault handling mechanism to perform fault handling on the faulty device. The method of the present disclosure, by configuring a third processing mechanism for the processor when the processor supports root port programmable input / output, enables different types of faults to be handled differently when a device directly connected to the processor fails. Thus, when the fault is a timeout completion error, only a reported information is generated, and further, based on the message header information log, it is determined whether to trigger port containment processing, verifying whether the fault information is generated by the faulty device triggering port containment processing, thereby preventing the faulty device that has already triggered port containment processing from triggering port containment processing again, and ensuring the normal operation of non-faulty devices.

[0131] In summary, the present disclosure has the following beneficial effects:

[0132] 1. By configuring different fault handling mechanisms for processing according to the support capabilities of the processor, and using different fault handling mechanisms to perform fault handling for faulty devices based on the connection relationship between the faulty device and the processing, it avoids the interruption of the operation of normal PCIe devices caused by using the EDPC technology for fault handling of different types of fault reports, and improves the granularity and accuracy of fault handling.

[0133] 2. When the processor supports programmable input / output of the root port, by configuring a second processing mechanism for the processor, so that when a device directly connected to the processor fails, different types of faults are handled differently. Thus, when the fault is a non - supported configuration space request, only a reporting message is generated and the port containment processing is not triggered, avoiding the triggering of fault handling for normal devices that have not completed initialization and ensuring the normal operation of non - faulty devices.

[0134] 3. When the processor supports programmable input / output of the root port, by configuring a third processing mechanism for the processor, so that when a device directly connected to the processor fails, different types of faults are handled differently. Thus, when the fault is a timeout completion error, only a reporting message is generated, and further based on the message header information log, it is determined whether to trigger the port containment processing, verifying whether the fault information is generated by the faulty device triggering the port containment processing, thereby avoiding the faulty device that has already triggered the port containment processing from triggering the port containment processing again, and ensuring the normal operation of non - faulty devices.

[0135] The following is an exemplary description of the present disclosure:

[0136] Embodiment 1. When a fatal error occurs during the enumeration process of Device 1 as shown in Figure 3 perform fault handling using the method shown in Figure 9 including the following steps:

[0137] After Device 1 fails, the triggering method of the PCIe root port DPC can be configured as RP PIO (i.e., the above - mentioned second fault handling mechanism). At this time, by configuring the RP PIO - related registers, the DPC function of the root port is triggered for some error types, and the DPC function of the root port is not triggered for other error types, and only a fault advisory report is made.

[0138] 1. Check whether the PCIe root port supports the DPC function of the RP PIO mechanism.

[0139] 2. When the PCIe root port supports the DPC function of the RP PIO mechanism, the PCIe root port is initialized as follows: the AER mechanism is masked, the RP PIO mechanism is configured, and it is set that the UR error of the configuration space request does not trigger DPC, and the remaining errors are set to trigger DPC (that is, the above-mentioned second fault handling mechanism).

[0140] 3. When the PCIe root port does not support the DPC function of the RP PIO mechanism, the PCIe root port is initialized: the traditional AER mechanism is configured to trigger DPC (that is, the above-mentioned first fault handling mechanism).

[0141] 4. The PCIe root port, as the requester, determines that device 1 has a fatal failure (that is, the above-mentioned processor receives the fault information of the faulty device).

[0142] 5. The PCIe root port determines whether the processor supports the RP PIO mechanism.

[0143] 6. When the PCIe root port does not support the RP PIO mechanism, the fatal error uses the traditional AER mechanism to trigger DPC.

[0144] 7. When the PCIe root port supports the RP PIO mechanism, it is further determined whether the fatal error is a UR error of the configuration space request.

[0145] 8. If the fatal error is a UR error of the configuration space request, the fatal error does not trigger DPC, and only a recommendatory fault report is made.

[0146] 9. If the fatal error is not a UR error of the configuration space request, the fatal error uses the RP PIO mechanism to trigger DPC.

[0147] Embodiment 2. When a fatal error occurs during the enumeration process of device 2 as shown in Figure 3 the method shown in Figure 10 is executed for fault handling, including the following steps:

[0148] If DPC is triggered at the downstream port of PCIe switch 1 (i.e., device 2), and this may be accompanied by one or more completion timeout errors occurring in the PCIe root device. The present disclosure filters out the CTO faults by using the RP PIO mechanism, accurately locates the source of the CTO faults to distinguish whether it is the CTO fault accompanied by the DPC triggered at the downstream port of PCIe switch 1, so as to accurately handle the faults, avoid the DPC caused by the RootPort of the accompanied CTO faults, and allow the PCIe root device to continue to operate normally with other PCIe switches (PCIe switch 2) and their downstream ports (device 3).

[0149] 1. During the device initialization phase, check whether the PCIe root port supports the DPC function of the RP PIO mechanism.

[0150] 2. When the PCIe root port supports the DPC function of the RP PIO mechanism, initialize the PCIe root port: mask the AER mechanism, configure the RP PIO mechanism, and set not to trigger DPC when a completion timeout error (Cfg CTO, I / O CTO, Mem CTO errors) occurs, and set the remaining errors to trigger DPC (i.e., the above third fault handling mechanism).

[0151] 3. When the PCIe root port does not support the DPC function of the RP PIO mechanism, initialize the PCIe root port: configure the traditional AER mechanism to trigger DPC.

[0152] 4. The PCIe root port, as a requester, discovers a fatal fault.

[0153] 5. When the PCIe root port does not support the RP PIO mechanism, the fatal error uses the traditional AER mechanism to trigger DPC.

[0154] 6. When the PCIe switch 1 supports the RP PIO mechanism, further determine whether the fatal error is a Cfg CTO, I / O CTO, or Mem CTO error.

[0155] 7. If the fatal error is at least one of the Cfg CTO, I / O CTO, and Mem CTO errors, then temporarily do not trigger DPC, and only report the fault suggestively. The suggestive fault report can indicate the identity information and fault cause of the faulty device, etc., for subsequent analysis and processing.

[0156] 8. If the fatal error is not at least one of the Cfg CTO, I / O CTO, and Mem CTO errors, then use the RPPIO mechanism to trigger DPC.

[0157] 9. Further, when the user software receives the Cfg CTO, I / O CTO, and MemCTO errors discovered by the PCIe as a requester, it is necessary to use the processor of the root port to parse the message header information log (RP PIO Header Log) to determine the source of the CTO error.

[0158] 10. Further, judge the type of the message header information log. If the parsing type is the memory space read / write type, then determine which device's space address the information address belongs to according to the information address indicated by the message header information log, and record the bus information, device information, and function information (bus, devices, function) of the device. Among them, the type of the information log can be judged by collecting the memory mapping (MMIO) space applied by the device.

[0159] 11. If the parsing type is the configuration space read / write type, the bus information, device information, and function information of the device indicated by the message header information log are determined based on the information address offset of the message header information log.

[0160] 12. Determine whether the device indicated by the message header information log is Device 2.

[0161] 13. If the device indicated by the message header information log is Device 2, no processing is performed and the fault handling ends.

[0162] 14. If the device indicated by the message header information log, the PCIe root port performs DPC.

[0163] To implement the device fault handling method provided by the embodiments of the present disclosure, the embodiments of the present disclosure further provide a device fault handling apparatus, as Figure 11 shown, the device fault handling apparatus 1100 includes:

[0164] An initialization unit 1101, configured to perform initialization configuration on the processor based on the fault handling mechanism supported by the processor;

[0165] A determination unit 1102, configured to determine the connection mode between the faulty device and the processor in response to the fault information of the faulty device;

[0166] A processing unit 1103, configured to perform fault handling on the faulty device based on the connection mode and / or initialization configuration.

[0167] In some embodiments, the initialization unit 1101 is further configured to, when the processor supports root port programmable input / output, mask the first fault handling mechanism of the processor and configure the second fault handling mechanism and / or the third fault handling mechanism for the processor; when the processor does not support root port programmable input / output, configure the first fault handling mechanism for the processor.

[0168] In some embodiments, the first fault handling mechanism includes: when a fault occurs in the faulty device, triggering port containment processing; the second fault handling mechanism includes: when the fault of the faulty device is a non-support for a configuration space request, not triggering port containment processing, and when the fault of the faulty device is a fault other than a non-support for a configuration space request, triggering port containment processing; the third fault handling mechanism includes: when the fault of the faulty device is a completion timeout error, not triggering port containment processing, and when the fault of the faulty device is an error other than a completion timeout error, triggering port containment processing.

[0169] In some embodiments, the connection mode between the faulty device and the processor includes: the faulty device is directly connected to the processor, and the faulty device is indirectly connected to the processor through a switch.

[0170] In some embodiments, the processing unit 1103 is further configured to trigger port containment processing when initializing and configuring the first fault handling mechanism; and perform fault handling on the faulty device based on the connection method when initializing and configuring the second fault handling mechanism or the third fault handling mechanism.

[0171] In some embodiments, the processing unit 1103 is further configured to perform fault handling on the faulty device by using the second fault handling mechanism when the faulty device is directly connected to the processor; and perform fault handling on the faulty device by using the third fault handling mechanism when the faulty device is indirectly connected to the processor through a switch.

[0172] In some embodiments, the processing unit 1103 is further configured to determine, based on the fault information, whether the fault of the faulty device is an unsupported configuration space request; and perform fault handling on the faulty device based on the determination result.

[0173] In some embodiments, the processing unit 1103 is further configured to perform fault handling on the faulty device based on the determination result, including: when the determination result indicates that the fault of the faulty device is an unsupported configuration space request, generating first reporting information about the fault and not triggering port containment processing; and when the determination result indicates that the fault of the faulty device is not an unsupported configuration space request, triggering port containment processing.

[0174] In some embodiments, the processing unit 1103 is further configured to determine, based on the fault information, whether the fault of the faulty device is a completion timeout error; and perform fault handling on the faulty device based on the determination result.

[0175] In some embodiments, the processing unit 1103 is further configured to trigger port containment processing when the determination result indicates that the fault of the faulty device is not a completion timeout error; and when the determination result indicates that the fault of the faulty device is a completion timeout error, generating second reporting information about the fault and obtaining the message header information log corresponding to the fault, so as to determine whether to perform port containment processing based on the message header information log.

[0176] In some embodiments, the processing unit 1103 is further configured to determine the space address indicated by the message header information log based on the type of the message header information log, so as to determine a first information set based on the space address; and when the type of the message header information log is a configuration space read / write type, determining the first information set based on the information address offset of the message header information log.

[0177] In some embodiments, the processing unit 1103 is further configured to determine the spatial address indicated by the packet header information log based on the type of the packet header information log, including: when the type of the packet header information log is a memory read / write type or an input / output read / write type, determining the spatial address based on the information address of the packet header information log; when the type of the packet header information log is a configuration space read / write type, determining the spatial address based on the information address offset of the packet header information log.

[0178] In some embodiments, the processing unit 1103 is further configured not to perform port containment processing when it is determined that the device corresponding to the first information set is a faulty device or when it is determined that the device corresponding to the first information set is a switch connected to the faulty device; and to perform port containment processing when it is determined that the device corresponding to the first information set is not a faulty device and the device corresponding to the first information set is not a switch connected to the faulty device.

[0179] In some embodiments, the processing unit 1103 is further configured to determine the type of the packet header information log based on the memory mapping space requested by the faulty device.

[0180] In some embodiments, the processing unit 1103 is further configured to determine the device identity information of the faulty device and the cause of the fault of the faulty device based on the first reported information or the second reported information; and to perform fault handling on the faulty device using a preset solution based on the device identity information and the cause of the fault.

[0181] In some embodiments, the processing unit 1103 is further configured to unload the driver program of the faulty device, remove the device label of the faulty device; release the link state between the faulty device and the processor; enumerate the faulty device, and re-enable the driver program of the faulty device.

[0182] In summary, the device fault handling apparatus proposed according to the present disclosure includes an initialization unit configured to perform initialization configuration on the processor based on the fault handling mechanism supported by the processor; a determination unit configured to determine the connection mode between the faulty device and the processor in response to the fault information of the faulty device; and a processing unit configured to perform fault handling on the faulty device based on the connection mode and / or the initialization configuration. The apparatus of the present disclosure configures different fault handling mechanisms for the processor according to the support capabilities of the processor, and performs fault handling on the faulty device using different fault handling mechanisms according to the connection relationship between the faulty device and the processor, avoiding the interruption of the operation of normal PCIe devices caused by using the EDPC technology for fault handling of different types of fault reports, and improving the granularity and accuracy of fault handling.

[0183] It should be noted that: When the device failure handling apparatus provided in the above embodiments performs device failure handling, only the division of the above program modules is used for illustration. In actual applications, the above processing can be allocated to different program modules according to needs, that is, the internal structure of the device failure handling apparatus is divided into different program modules to complete all or part of the processing described above. In addition, the device failure handling apparatus provided in the above embodiments and the method embodiments of the device failure handling method provided in the embodiments of the present disclosure belong to the same concept. For the specific implementation process, please refer to the method embodiments, which will not be elaborated here.

[0184] Figure 12 is a schematic diagram of the hardware composition structure of the electronic device provided in the embodiments of the present disclosure. As Figure 12 shown, the electronic device 1200 includes at least one processor 1202; and a memory 1201 communicatively connected to the at least one processor 1202; wherein, the memory 1201 stores instructions executable by the at least one processor 1202, and the instructions are executed by the at least one processor 1202 to implement the steps of the device failure handling method described in the embodiments of the present disclosure.

[0185] Optionally, the electronic device may specifically be the device failure handling apparatus of the embodiments of the present application, and the electronic device can implement the corresponding processes implemented by the device failure handling apparatus in the various methods of the embodiments of the present application. For the sake of brevity, it will not be elaborated here.

[0186] It can be understood that the electronic device further includes a communication interface 1203. Each component in the electronic device is coupled together through a bus system 1204. It can be understood that the bus system 1204 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 1204 further includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 12 all kinds of buses are labeled as the bus system 1204.

[0187] It can be understood that the memory 1201 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM, Read Only Memory), a programmable read-only memory (PROM, Programmable Read-Only Memory), an erasable programmable read-only memory (EPROM, Erasable Programmable Read-Only Memory), an electrically erasable programmable read-only memory (EEPROM, Electrically Erasable Programmable Read-Only Memory), a ferromagnetic random access memory (FRAM, ferromagnetic random access memory), a flash memory (Flash Memory), a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM, Compact Disc Read-Only Memory); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM, Random Access Memory), which is used as an external cache. This method is described by way of example but not limitation, and many forms of RAM are available, such as a static random access memory (SRAM, Static Random Access Memory), a synchronous static random access memory (SSRAM, Synchronous Static Random Access Memory), a dynamic random access memory (DRAM, Dynamic Random Access Memory), a synchronous dynamic random access memory (SDRAM, Synchronous Dynamic Random Access Memory), a double data rate synchronous dynamic random access memory (DDR SDRAM, Double Data Rate Synchronous Dynamic Random Access Memory), an enhanced synchronous dynamic random access memory (ESDRAM, Enhanced Synchronous Dynamic Random Access Memory), a sync link dynamic random access memory (SLDRAM, SyncLink Dynamic Random Access Memory), a direct rambus random access memory (DRRAM, Direct Rambus Random Access Memory).The memory 1201 described in the embodiments of the present invention is intended to include but not limited to these and any other suitable types of memories.

[0188] The method disclosed in the above embodiments of the present disclosure can be applied to or implemented by the processor 1202. The processor 1202 has the ability to process signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit in hardware or the command in software form in the processor 1202. The above processor 1202 can be a general-purpose processor, DSP, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 1202 can implement or execute each method, step, and logic block diagram disclosed in the embodiments of the present invention. The general-purpose processor can be a microprocessor or any conventional processor, etc. Combining with the steps of the method disclosed in the embodiments of the present invention, it can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by the combination of the hardware and software modules in the decoding processor. The software module can be located in the storage medium, and this storage medium is located in the memory 1201. The processor 1202 reads the information in the memory 1201 and combines its hardware to complete the steps of the foregoing method.

[0189] In an exemplary embodiment, the electronic device can be implemented by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), FPGAs, general-purpose processors, controllers, MCUs, microprocessors, or other electronic components for executing the foregoing method.

[0190] The embodiments of the present disclosure also provide a non-transitory computer-readable storage medium storing computer commands, and the computer commands are used to cause the computer to execute the steps of the device fault handling method described in the embodiments of the present disclosure when executed.

[0191] The embodiments of the present disclosure also provide a computer program product, including a computer program, and the computer program realizes the steps of the device fault handling method described in the embodiments of the present disclosure when executed by the processor.

[0192] Optionally, the computer-readable storage medium can be applied to the device fault handling device in the embodiments of the present application, and the computer commands cause the computer to execute the corresponding processes implemented by the device fault handling device in each method of the embodiments of the present application. For the sake of brevity, it will not be elaborated here.

[0193] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the couplings, direct couplings, or communication connections between the various components shown or discussed can be through some interfaces. The indirect couplings or communication connections of devices or units can be electrical, mechanical, or other forms.

[0194] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units; some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0195] In addition, in each embodiment of the present invention, the various functional units can all be integrated in a processing unit, or each unit can be separately used as a unit, or two or more units can be integrated in a unit; the above-mentioned integrated units can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.

[0196] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program commands. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: removable storage devices, ROM, RAM, magnetic disks, or optical disks and other various media that can store program codes.

[0197] Alternatively, if the above-mentioned integrated units of the present invention are implemented in the form of software function modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several commands to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present invention. And the foregoing storage medium includes: removable storage devices, ROM, RAM, magnetic disks, or optical disks and other various media that can store program codes.

[0198] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claimed rights.

Claims

1. A method for processing device failures, characterized in that, The method includes: Based on the fault handling mechanism supported by the processor, initialize and configure the processor; In response to the fault information of the faulty device, determine the connection method between the faulty device and the processor; Based on the connection method and / or the initialization configuration, perform fault handling on the faulty device; Among them, the initializing and configuring the processor based on the fault handling mechanism supported by the processor includes: When the processor supports root port programmable input / output, mask the first fault handling mechanism of the processor, and configure the second fault handling mechanism and / or the third fault handling mechanism for the processor; When the processor does not support the root port programmable input / output, configure the first fault handling mechanism for the processor; among them, the first fault handling mechanism includes: when the faulty device has a fault, trigger port containment processing; the second fault handling mechanism includes: when the fault of the faulty device is a non-support for configuration space request, do not trigger the port containment processing, and when the fault of the faulty device is a fault other than the non-support for configuration space request, trigger the port containment processing; the third fault handling mechanism includes: when the fault of the faulty device is a completion timeout error, do not trigger the port containment processing, and when the fault of the faulty device is an error other than the completion timeout error, trigger the port containment processing; Among them, the port containment processing includes: uninstall the driver of the faulty device, and remove the device label of the faulty device; release the link state between the faulty device and the processor; enumerate the faulty device, and re-enable the driver of the faulty device.

2. The method according to claim 1, wherein The connection method between the faulty device and the processor includes: the faulty device is directly connected to the processor, and the faulty device is indirectly connected to the processor through a switch.

3. The method according to claim 1, wherein The performing fault handling on the faulty device based on the connection method and / or the initialization configuration includes: When the first fault handling mechanism is initialized and configured, trigger port containment processing; When the second fault handling mechanism or the third fault handling mechanism is initialized and configured, perform fault handling on the faulty device based on the connection method; Among them, the performing fault handling on the faulty device based on the connection method when the second fault handling mechanism or the third fault handling mechanism is initialized and configured includes: When the faulty device is directly connected to the processor, use the second fault handling mechanism to perform fault handling on the faulty device; When the faulty device is indirectly connected to the processor through a switch, use the third fault handling mechanism to perform fault handling on the faulty device.

4. The method according to claim 3, wherein The using the second fault handling mechanism to perform fault handling on the faulty device when the faulty device is directly connected to the processor includes: Based on the fault information, determine whether the fault of the faulty device is a non-support for configuration space request; Based on the judgment result, perform fault handling on the faulty device.

5. The method according to claim 4, wherein Performing fault handling on the faulty device based on the judgment result includes: When the judgment result indicates that the fault of the faulty device is the unsupported configuration space request, generating first reporting information for the fault and not triggering port containment processing; When the judgment result indicates that the fault of the faulty device is not the unsupported configuration space request, triggering port containment processing.

6. The method according to claim 3, wherein When the faulty device is indirectly connected to the processor through the switch, performing fault handling on the faulty device using the third fault handling mechanism includes: Based on the fault information, determining whether the fault of the faulty device is a completion timeout error; Performing fault handling on the faulty device based on the judgment result.

7. The method according to claim 6, wherein Performing fault handling on the faulty device based on the judgment result includes: When the judgment result indicates that the fault of the faulty device is not the completion timeout error, triggering port containment processing; When the judgment result indicates that the fault of the faulty device is the completion timeout error, generating second reporting information for the fault and obtaining the message header information log corresponding to the fault, and based on the message header information log, determining whether to perform port containment processing.

8. The method according to claim 7, wherein Determining whether to perform port containment processing based on the message header information log includes: Based on the type of the message header information log, determining the first information set indicated by the message header information log; Determining whether the device corresponding to the first information set is the faulty device to determine whether to perform port containment processing.

9. The method according to claim 8, wherein Determining the space address indicated by the message header information log based on the type of the message header information log includes: When the type of the message header information log is a memory read / write type or an input / output read / write type, based on the information address of the message header information log, determining the space address indicated by the message header information log, and based on the space address, determining the first information set; When the type of the message header information log is a configuration space read / write type, based on the information address offset of the message header information log, determining the first information set.

10. The method according to claim 8, characterized in that, Determining whether the device corresponding to the first information set is the faulty device to determine whether to perform port containment processing includes: When it is determined that the device corresponding to the first information set is the faulty device, or it is determined that the device corresponding to the first information set is the switch connected to the faulty device, not performing port containment processing; When it is determined that the device corresponding to the first information set is not the faulty device and the device corresponding to the first information set is not the switch connected to the faulty device, performing port containment processing.

11. The method according to claim 8, wherein The method further includes: Based on the memory mapping space applied for by the faulty device, determining the type of the message header information log.

12. The method according to claim 11, wherein The first information set includes at least one of the following: the bus information of the faulty device, the device information of the faulty device, and the function information of the faulty device.

13. The method according to any one of claims 5 or 7, characterized in that The method further includes: Based on the first reporting information or the second reporting information, determine the device identity information of the faulty device and the cause of the fault of the faulty device; Based on the device identity information and the cause of the fault, perform fault handling using a preset solution.

14. An electronic device, characterized in that, Comprising: A processor and a memory for storing a computer program capable of running on the processor, wherein, when the processor is used to run the computer program, it executes the device fault handling method according to any one of claims 1-13.

15. A non-transitory computer-readable storage medium storing computer commands, characterized in that, The computer command is used to cause the computer to execute the device fault handling method according to any one of claims 1-13.

16. A computer program product, characterized in that, Comprising a computer program which, when executed by a processor, implements the device fault handling method according to any one of claims 1-13.

Citation Information

Patent Citations

  • Endpoint device management method, device and system

    CN112306913A