Equipment fault processing method, electronic equipment, storage medium and product
By initializing and configuring different fault handling mechanisms on the processor and troubleshooting according to the connection between the faulty equipment and the processor, the problem of fault handling interruption caused by EDPC technology is solved, and the accuracy and granularity of fault handling are improved.
Patent Information
- Application Number
- CN202510489529.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-18
AI Technical Summary
When a PCIe device fails and errors occur, the processor uses EDPC technology to deal with the fault, causing all PCIe devices to interrupt the connection with the processor, affecting the operation of normal PCIe devices.
Based on the fault handling mechanism supported by the processor, the processor is initialized and configured. In response to the fault information of the fault device, the connection method between the fault device and the processor is determined, and the fault processing is performed on the fault device based on the connection method and/or the initial configuration.
By configuring different fault handling mechanisms according to the processor's support capabilities and troubleshooting according to the connection relationship between the fault device and the processor, the operation interruption of normal PCIe equipment caused by EDPC technology is avoided, and the processing granularity and accuracy of fault handling is improved.
Smart Images

Figure CN120011127A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the computer field, and in particular to a device failure processing method, an electronic device, a storage medium and a product. Background Art
[0002] With the development of computer technology, high-speed serial computer expansion bus (Peripheral Component Interconnect Express, PCIe) devices have become an indispensable part of current computer systems. However, with the evolution of PCIe technology, the running speed of PCIe devices is getting faster and faster, and their failures are becoming more and more frequent.
[0003] When a PCIe device fails and reports an error, the processor usually uses Enhanced Downstream Port Containment (EDPC) technology to handle different types of fault errors. This means that when the processor is connected to multiple PCIe devices, if one or more of the PCIe devices fails, using EDPC technology to handle the fault will cause all PCIe devices to lose connection with the processor, affecting the normal operation of the PCIe devices. Summary of the invention
[0004] The present disclosure provides a device fault handling method, electronic device, storage medium and product, wherein the method includes: initializing and configuring the processor based on a fault handling mechanism supported by the processor; determining the connection mode between the faulty device and the processor in response to the fault information of the faulty device; and performing fault handling on the faulty device based on the connection mode and / or the initialization configuration. The present disclosure configures different fault handling mechanisms for the processing according to the support capability of the processor, and performs fault handling for the faulty device using different fault handling mechanisms according to the connection relationship between the faulty device and the processing, thereby avoiding the interruption of normal PCIe device operation caused by using EDPC technology for fault handling of different types of fault reports, and improving the processing granularity and accuracy of fault handling.
[0005] The first aspect of the present disclosure provides a method for handling device faults, including: initializing and configuring a processor based on a fault handling mechanism supported by the processor; determining a connection method between the faulty device and the processor in response to fault information of the faulty device; and handling the fault of the faulty device based on the connection method and / or the initialization configuration.
[0006] In some embodiments of the present disclosure, initializing and configuring the processor based on the fault handling mechanisms supported by the processor includes: when the processor supports root port programmable input and output, shielding the first fault handling mechanism of the processor, and configuring the second fault handling mechanism and / or the third fault handling mechanism for the processor; when the processor does not support root port programmable input and output, configuring the first fault handling mechanism for the processor.
[0007] In some embodiments of the present disclosure, a first fault handling mechanism includes: when a fault exists in a faulty device, port containment processing is triggered; a second fault handling mechanism includes: when the fault of the faulty device is a failure to support a configuration space request, port containment processing is not triggered; when the fault of the faulty device is a fault other than failure to support a configuration space request, port containment processing is triggered; a third fault handling mechanism includes: when the fault of the faulty device is a completion timeout error, port containment processing is not triggered; when the fault of the faulty device is an error other than a completion timeout error, port containment processing is triggered.
[0008] In some embodiments of the present disclosure, the connection manner between the faulty device and the processor includes: the faulty device is directly connected to the processor, and the faulty device is indirectly connected to the processor through a switch.
[0009] In some embodiments of the present disclosure, fault handling of a faulty device based on a connection mode and / or initialization configuration includes: when the first fault handling mechanism is initialized for configuration, port containment processing is triggered; when the second fault handling mechanism or the third fault handling mechanism is initialized for configuration, fault handling of the faulty device is performed based on the connection mode.
[0010] In some embodiments of the present disclosure, when the second fault handling mechanism or the third fault handling mechanism is initialized for configuration, fault handling of the faulty device is performed based on the connection mode, including: when the faulty device is directly connected to the processor, the second fault handling mechanism is used to perform fault handling on the faulty device; when the faulty device is indirectly connected to the processor via a switch, the third fault handling mechanism is used to perform fault handling on the faulty device.
[0011] In some embodiments of the present disclosure, when a faulty device is directly connected to a processor, a second fault handling mechanism is used to perform fault handling on the faulty device, including: based on fault information, determining whether the fault of the faulty device is due to failure to support configuration space requests; and based on the determination result, performing fault handling on the faulty device.
[0012] In some embodiments of the present disclosure, based on the result of judgment, fault processing of the faulty device includes: when the result of judgment is that the fault of the faulty device does not support the configuration space request, generating the first reporting information of the fault, and not triggering the port containment processing; when the result of judgment is that the fault of the faulty device does not not support the configuration space request, triggering the port containment processing.
[0013] In some embodiments of the present disclosure, when a faulty device is indirectly connected to a processor via a switch, a third fault handling mechanism is used to perform fault handling on the faulty device, including: judging whether the fault of the faulty device is a completion timeout error based on fault information; and performing fault handling on the faulty device based on the judgment result.
[0014] In some embodiments of the present disclosure, based on the result of judgment, fault processing of the faulty device includes: when the result of judgment is that the fault of the faulty device is not a completion timeout error, triggering port containment processing; when the result of judgment is that the fault of the faulty device is a completion timeout error, generating second reporting information of the fault, and obtaining the message header information log corresponding to the fault, so as to determine whether to perform port containment processing based on the message header information log.
[0015] In some embodiments of the present disclosure, determining whether to perform port containment processing based on the packet header information log includes: determining the space address indicated by the packet header information log based on the type of the packet header information log; and determining whether the device corresponding to the space address is a faulty device to determine whether to perform port containment processing.
[0016] In some embodiments of the present disclosure, determining the space address indicated by the message header information log based on the type of the message header information log includes: when the type of the message header information log is a memory read-write type or an input-output read-write type, determining the space address indicated by the message header information log based on the information address of the message header information log, so as to determine the first information set based on the space address; when the type of the message header information log is a configuration space read-write type, determining the first information set based on the information address offset of the message header information log.
[0017] In some embodiments of the present disclosure, judging whether the device corresponding to the space address is a faulty device to determine whether to perform port containment processing includes: when it is judged that the device corresponding to the first information set is a faulty device, or when it is judged that the device corresponding to the first information set is a switch connected to the faulty device, port containment processing is not performed; when it is judged that the device corresponding to the first information set is not a faulty device, and the device corresponding to the first information set is not a switch connected to the fault, port containment processing is performed.
[0018] In some embodiments of the present disclosure, the method further includes: determining the type of the message header information log based on the memory mapping space requested by the faulty device.
[0019] In some embodiments of the present disclosure, the first information set includes at least one of the following: bus information of the faulty device, device information of the faulty device, and function information of the faulty device.
[0020] In some embodiments of the present disclosure, the method further includes: determining the device identity information of the faulty device and the fault cause of the faulty device based on the first reporting information or the second reporting information; and performing fault handling using a preset solution based on the device identity information and the fault cause.
[0021] In some embodiments of the present disclosure, the port containment process includes: uninstalling the driver of the faulty device and removing the device tag of the faulty device; releasing the link state between the faulty device and the processor; enumerating the faulty device and re-enabling the driver of the faulty device.
[0022] The second aspect embodiment of the present disclosure proposes an electronic device, comprising: a processor and a memory for storing a computer program that can be run on the processor, wherein the processor, when being used to run the computer program, executes the method described in the first aspect embodiment of the present disclosure.
[0023] The third aspect embodiment of the present disclosure proposes a non-transitory computer-readable storage medium storing computer commands, wherein the computer commands are used to enable a computer to execute the method described in the first aspect embodiment of the present disclosure.
[0024] The fourth aspect embodiment of the present disclosure provides a computer program product, including a computer program, which implements the method described in the first aspect embodiment of the present disclosure when executed by a processor.
[0025] In summary, a device fault handling method provided in the present disclosure includes: initializing and configuring the processor based on a fault handling mechanism supported by the processor; determining the connection mode between the faulty device and the processor in response to the fault information of the faulty device; and performing fault handling on the faulty device based on the connection mode and / or the initialization configuration. The method disclosed in the present disclosure configures different fault handling mechanisms for the processing according to the support capabilities of the processor, and performs fault handling for the faulty device using different fault handling mechanisms according to the connection relationship between the faulty device and the processing, thereby avoiding the interruption of normal PCIe device operation caused by the use of EDPC technology for fault handling for different types of fault reports, and improving the processing granularity and accuracy of fault handling.
[0026] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure. Figure 1 A schematic diagram of a process flow of a device failure handling method provided by an embodiment of the present disclosure; Figure 2 A schematic diagram of a scenario of a method for handling a device failure provided by an embodiment of the present disclosure; Figure 3 A schematic diagram of another method for handling device failures provided by an embodiment of the present disclosure; Figure 4 A schematic diagram of a process flow of a device failure handling method provided by an embodiment of the present disclosure; Figure 5 A flowchart of another device failure processing method provided by an embodiment of the present disclosure; Figure 6 A flowchart of another device failure processing method provided by an embodiment of the present disclosure; Figure 7 A flowchart of another device failure processing method provided by an embodiment of the present disclosure; Figure 8 A flowchart of another device failure processing method provided by an embodiment of the present disclosure; Fig. 9 A flowchart of another method for handling device failures provided by an embodiment of the present disclosure; Fig.10 An example flow chart of a method for handling equipment failures provided in an embodiment of the present disclosure; Fig.11 A schematic diagram of the structure of another device failure handling apparatus provided by an embodiment of the present disclosure; Fig.12 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0028] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0029] With the development of computer technology, PCIe devices have become an indispensable part of current computer systems. However, with the evolution of PCIe technology, the running speed of PCIe devices is getting faster and faster, and their failures are becoming more and more frequent. For some fault problems, the operating system corresponding to the PCIe device can automatically repair them, but for most problems, the operating system cannot automatically repair them, resulting in machine downtime or restart, so that the operating system cannot run normally.
[0030] In order to solve the system downtime caused by fatal failure of PCIe devices, EDPC technology has emerged in the industry. EDPC technology is an error isolation and recovery technology for PCIe bus. When a link error (such as malformed transaction layer data message, unexpected downtime, etc.) is detected, it disables a single PCIe link and forcibly terminates the unfinished request, thereby avoiding error propagation and protecting the operating system from potential bad data.
[0031] Specifically, the related art usually uses the Advanced Error Reporting (AER) mechanism of PCIe to trigger EDPC. AER is a mechanism for detecting and reporting errors occurring in PCle devices, which allows PCle devices to detect and report various types of errors, such as non-fatal, recoverable, and serious errors. AER implements a set of registers and corresponding error notification mechanisms on PCle devices, and information about errors can be obtained by reading these registers. Using AER, the system can better monitor and handle error conditions of PCle devices to improve data integrity and reliability.
[0032] Specifically, EDPC and AER can work together to implement the isolation and repair functions of PCIe devices, such as Figure 1 The method shown includes the following steps: ① The processor detects an uncorrectable error in the device (i.e., an error that the operating system cannot automatically fix). ② The processor sends a signal to the operating system for error handling. ③ The operating system battery management module notifies the device driver to uninstall the driver of the device connected to the processor; and removes the device tag of the device to prevent subsequent access to the memory-mapped input / output (MMIO) space. ④ The operating system battery management module recovers the device by releasing the device's (Link) link state, re-enumerating the device, and re-enabling the device driver of the device.
[0033] The above-mentioned related technologies can accurately isolate and repair the faulty device when the PCIe root port of the processor corresponds to a single device. However, when the PCIe root port corresponds to multiple devices, that is, the device corresponding to the PCIe root port is a PCIe expansion (PCIe Switch) device (such as Figure 2 As shown in FIG. 1 , when one of the multiple devices fails (for example, device 1 fails), PCIe switch 1 triggers DPC, and causes a timeout failure at the PCIe root port, thereby causing the processor to trigger DPC, resulting in interruption of normal device operation.
[0034] In addition, during the enumeration process of using DPC to handle device faults, since the device has not yet been initialized, some PCIE configuration (cfg) instructions cause PCIE Unsupported Request (UR) errors. In the related art, the UR error is usually shielded to avoid false alarms. However, when the operating system completes the enumeration of the device and runs normally, the PCIE UR error is shielded, resulting in a more serious processor error.
[0035] The following is an explanation of the technical terms involved in this disclosure: Root Port PIO (RP PIO) is a mechanism in the PCIe architecture for managing errors encountered when the Root Port sends non-posted requests. It provides fine-grained control over uncorrectable errors and advisory errors. The RP PIO error control register provides fine-grained error management capabilities for non-posted requests, allowing flexible configuration based on request type (configuration, input output, memory) and error type.
[0036] In order to solve the technical problems existing in the related art, an embodiment of the present disclosure provides a method for handling device failure.
[0037] The following first describes the application scenarios involved in the present disclosure by way of example: like Figure 3 As shown, the processor is directly connected to device 1 through a PCIe port, device 2 is connected to the PCIe port of the processor through a PCIe switch 1, and device 3 is connected to the PCIe port of the processor through a PCIe switch 2.
[0038] It should be understood that Figure 3 For example only, the application scenarios of the present disclosure may include: Figure 3Some or all of the devices shown, and the present disclosure does not limit the number of devices directly connected to the processor, and does not limit the number of devices connected to the processor through a PCIe switch. For example, the application scenario of the present disclosure may include only one or more devices directly connected to the processor; or only one or more devices indirectly connected to the processor through a PCIe switch; or one or more devices directly connected to the processor, and one or more devices indirectly connected to the processor.
[0039] The embodiments of the present disclosure will be described in detail below.
[0040] like Figure 4 As shown, an embodiment of the present disclosure provides a method for handling a device failure, comprising the following steps: Step 101 : Initialize and configure the processor based on a fault handling mechanism supported by the processor.
[0041] In some embodiments, different initialization configurations may be performed for the processor depending on whether the processor supports root port programmable I / O (RP PIO).
[0042] In some embodiments, when the processor does not support root port programmable input and output, the processor's initialization configuration can be to trigger port containment processing when a faulty device fails, that is, regardless of the fault type of the faulty device, the faulty device is uniformly processed using port containment processing.
[0043] In some embodiments, when the processing supports root port programmable input and output, the processor's initialization configuration can use port containment processing for some fault types, and not use port containment processing for some fault types but only report the fault, thereby implementing different fault handling methods for different fault types to avoid affecting the normal operation of the device.
[0044] In some embodiments, the port containment processing may be EDPC processing, or downstream port containment (DPC) processing, or an improved processing method based on DPC processing, etc., which is not limited in the present disclosure.
[0045] Step 102: In response to the fault information of the faulty device, determine the connection mode between the faulty device and the processor.
[0046] In some embodiments, the fault information is used to indicate that a device connected to the processor has failed.
[0047] In some embodiments, when fault information is received, the processor can determine the connection method between the faulty device and the processor by querying the connection information of the faulty device, and then perform targeted fault handling methods on the faulty device in combination with the initialization configuration of the processor, but is not limited to this. The present disclosure does not limit the technical means used to determine the connection method.
[0048] In some embodiments, the connection mode between the faulty device and the processor includes: the faulty device is directly connected to the processor, and the faulty device is indirectly connected to the processor via a switch. The switch may be a PCIe switch.
[0049] Specifically, the faulty device being directly connected to the processor may be directly connected to a PCIe root port of the processor; the faulty device being indirectly connected to the processor via a PCIe switch may be connected to a PCIe root port of the processor via a PCIe switch.
[0050] Step 103: Based on the connection mode and / or initialization configuration, perform fault processing on the faulty device.
[0051] In some embodiments, the initialization configuration is to trigger port containment processing when a faulty device is faulty. At this time, regardless of whether the faulty device is directly or indirectly connected to the processor, the faulty device is processed by port containment processing, that is, the faulty device is processed directly based on the initialization configuration.
[0052] In some embodiments, the processor is initialized and configured to use port containment for some fault types, and not use port containment for some fault types but only perform fault reporting. Since different connection methods correspond to different fault types that require port containment, it is necessary to combine the connection method between the faulty device and the processor to determine whether the fault type of the faulty device requires port containment so as to perform fault processing on the faulty device.
[0053] In some embodiments, the port containment process may include uninstalling the driver of the faulty device and removing the device tag of the faulty device; releasing the link state between the faulty device and the processor; enumerating the faulty device and re-enabling the driver of the faulty device, thereby isolating and repairing the faulty device.
[0054] In summary, the device fault handling method proposed in the present disclosure includes: initializing and configuring the processor based on the fault handling mechanism supported by the processor; determining the connection mode between the faulty device and the processor in response to the fault information of the faulty device; and performing fault handling on the faulty device based on the connection mode and / or the initialization configuration. The method disclosed in the present disclosure configures different fault handling mechanisms for the processing according to the support capabilities of the processor, and performs fault handling for the faulty device using different fault handling mechanisms according to the connection relationship between the faulty device and the processing, thereby avoiding the interruption of normal PCIe device operation caused by the use of EDPC technology for fault handling for different types of fault reports, and improving the processing granularity and accuracy of fault handling.
[0055] Figure 5 A flow chart of a device failure processing method proposed in the present disclosure is further shown. Figure 1 The illustrated embodiment further explains, Figure 5 The following steps may be included.
[0056] Step 201 : when the processor supports root port programmable input and output, shield the first fault handling mechanism of the processor, and configure the second fault handling mechanism and / or the third fault handling mechanism for the processor.
[0057] In some embodiments, the first fault handling mechanism includes: triggering a port throttling process when a faulty device is faulty.
[0058] In some embodiments, the second fault handling mechanism includes: when the fault of the faulty device is failure to support configuration space request, port containment processing is not triggered; when the fault of the faulty device is a fault other than failure to support configuration space request, port containment processing is triggered.
[0059] In some embodiments, the third fault processing mechanism includes: when the fault of the faulty device is a completion timeout error, the port containment process is not triggered; when the fault of the faulty device is an error other than the completion timeout error, the port containment process is triggered.
[0060] In some embodiments, the second fault handling mechanism is used when the faulty device is directly connected to the processor, and the third fault handling mechanism is used when the faulty device is indirectly connected to the processor via a PCIe switch.
[0061] In some embodiments, not supporting the configuration space request may be that an error occurs in the UR of the configuration space request, but is not limited thereto.
[0062] In some embodiments, the completion timeout error may be a timeout for allocation space request, a timeout for input / output request, a timeout for memory request, etc., which is not limited in the present disclosure.
[0063] In some optional embodiments, when all devices connected to the processor are directly connected, only the second fault handling mechanism may be configured for the processor; when all devices connected to the processor are indirectly connected, only the third fault handling mechanism may be configured for the processor; when the devices connected to the processor are both directly connected and indirectly connected, both the second fault handling mechanism and the third fault handling mechanism may be configured for the processor at the same time.
[0064] Step 202: When the processor does not support root port programmable input and output, configure a first fault handling mechanism for the processor.
[0065] In some embodiments, when the processor does not support root port programmable input and output, that is, when the processor does not support different fault handling methods according to the fault type, the first fault handling mechanism can be directly configured for the processor to enable fault handling when a device connected to the processor fails.
[0066] In summary, the device fault handling method proposed in the present disclosure includes: when the processor supports the root port programmable input and output, shielding the first fault handling mechanism of the processor, and configuring the second fault handling mechanism and / or the third fault handling mechanism for the processor; when the processor does not support the root port programmable input and output, configuring the first fault handling mechanism for the processor. The method disclosed in the present disclosure improves the granularity of the processor fault handling and improves the accuracy of fault handling by configuring different fault handling mechanisms for the processor according to the processor's capabilities.
[0067] Figure 6 A flow chart of a device failure processing method proposed in the present disclosure is further shown. Figure 1 and Figure 2 The illustrated embodiment further explains, Figure 6 The following steps may be included.
[0068] Step 301: When the processor does not support root port programmable input and output, configure a first fault handling mechanism for the processor.
[0069] In some embodiments, the principle of step 301 is the same as that of step 201 , and reference may be made to the embodiment and related description in step 201 , which will not be repeated here.
[0070] Step 302: In response to the fault information of the faulty device, triggering port containment processing.
[0071] In some embodiments, since the processor is initially configured as the first fault handling mechanism, when receiving fault information, the processor is unable to adopt different fault handling methods according to different fault types. Therefore, the processor can directly perform port containment processing on the faulty device to perform fault handling.
[0072] In summary, the device fault handling method proposed in the present disclosure includes: when the processor does not support the root port programmable input and output, configuring a first fault handling mechanism for the processor; and triggering port containment processing in response to fault information of the faulty device. The method of the present disclosure configures the first processing mechanism for the processor that does not support the root port programmable input and output, so that when a device directly or indirectly connected to the processor fails, the port containment processing can be used to implement fault handling, thereby increasing the scope of application of the present disclosure.
[0073] Figure 7 A flow chart of a device failure processing method proposed in the present disclosure is further shown. Figure 1 and Figure 2 The illustrated embodiment further explains, Figure 7 The following steps may be included: Step 401 : when the processor supports the root port programmable input and output, shield the first fault handling mechanism of the processor, and configure the second fault handling mechanism for the processor, or configure the second fault handling mechanism and the third fault handling mechanism for the processor.
[0074] In some embodiments, the principle of step 401 is the same as that of step 201 , and reference may be made to the embodiment and related description in step 201 , which will not be repeated here.
[0075] Step 402, in response to the fault information for the faulty device, determining that the faulty device is directly connected to the processor.
[0076] In some embodiments, when fault information is received, the processor may determine whether the faulty device is directly connected to the processor by querying the connection information of the faulty device, but the present disclosure is not limited to this. The technical means used to determine the connection method are not limited.
[0077] Step 403: Use the second fault handling mechanism to handle the fault of the faulty device.
[0078] In some embodiments, the fault information may also indicate the fault type of the faulty device. Therefore, it is possible to determine whether the fault of the faulty device is due to not supporting the configuration space request based on the fault information, and then perform fault processing on the faulty device based on the determination result.
[0079] Specifically, when it is determined according to the fault information that the fault of the faulty device is that the configuration space request is not supported, the faulty device is processed as follows: first reporting information of the fault is generated, and port containment processing is not triggered.
[0080] When it is determined according to the fault information that the fault of the faulty device is not a failure to support the configuration space request (that is, a fault other than failure to support the configuration space request), the processing method for the faulty device is: triggering port containment processing.
[0081] Specifically, when the device is directly connected to the processor and the processor enumerates the device, a normal device that has not completed initialization will cause a fault of not supporting the configuration space request (enumeration is used by the processor to read or write data to the device, and because the device has not completed initialization, the processor cannot read or write data, so a fault of not supporting the configuration space request will occur). At this time, the normal device has not failed, but has only failed to complete initialization, so the normal device does not need to perform port containment processing.
[0082] Furthermore, in order to avoid the situation where the failure of not supporting the configuration space request is not caused by the failure of normal equipment to complete initialization, that is, the equipment actually fails and causes the failure to support the configuration space request, it is necessary to generate the first reporting information to analyze the specific cause of not supporting the configuration space request, so as to determine whether to perform corresponding fault processing.
[0083] In other words, based on the first reported information, the fault cause of not supporting the configuration space request can be analyzed, so that when the failure to support the configuration space request is not caused by the failure of the normal device to complete the initialization, the preset fault handling method can be used to handle the faulty device, thereby achieving isolation and repair of the faulty device.
[0084] In some embodiments, the first reporting information may be advisory error report information, but is not limited thereto.
[0085] In summary, the device fault handling method proposed in the present disclosure includes: when the processor supports the root port programmable input and output, shielding the first fault handling mechanism of the processor, and configuring the second fault handling mechanism for the processor, or configuring the second fault handling mechanism and the third fault handling mechanism for the processor; responding to the fault information for the faulty device, determining that the faulty device is directly connected to the processor; and performing fault handling on the faulty device using the second fault handling mechanism. The method disclosed in the present disclosure configures the second processing mechanism for the processor when the processor supports the root port programmable input and output, so that when a device directly connected to the processor fails, different fault handling is used for different types of faults, so that when the fault is not supporting the configuration space request, only reporting information is generated, and port containment processing is not triggered, thereby avoiding normal devices that have not completed initialization from triggering fault handling and ensuring the normal operation of the device.
[0086] Figure 8 A flow chart of a device failure processing method proposed in the present disclosure is further shown. Figure 1 and Figure 2 The illustrated embodiment further explains, Figure 8 The following steps may be included: Step 501: When the processor supports the root port programmable input and output, shield the first fault handling mechanism of the processor, and configure the third fault handling mechanism for the processor, or configure the second fault handling mechanism and the third fault handling mechanism for the processor.
[0087] In some embodiments, the principle of step 501 is the same as that of step 201 , and reference may be made to the embodiment and related description in step 201 , which will not be repeated here.
[0088] Step 502, in response to the fault information for the faulty device, determining that the faulty device is indirectly connected to the processor.
[0089] In some embodiments, when fault information is received, the processor may determine that the faulty device is indirectly connected to the processor by querying the connection information of the faulty device, but the present disclosure is not limited to this. The technical means used to determine the connection method are not limited.
[0090] Step 503: Use the third fault handling mechanism to handle the fault of the faulty device.
[0091] In some embodiments, the fault information may also indicate the fault type of the faulty device. Therefore, it is possible to determine whether the fault of the faulty device is a completion timeout error based on the fault information, and then perform fault processing on the faulty device based on the determination result.
[0092] Specifically, when it is determined according to the fault information that the fault of the faulty device is not a completion timeout error, the processing method for the faulty device is: triggering port containment processing.
[0093] When it is determined according to the fault information that the fault of the faulty device is a completion timeout error, second fault reporting information is generated, and a message header information log corresponding to the fault is obtained to determine whether to perform port containment processing based on the message header information log (RP PIO Header Log).
[0094] Furthermore, the first information set indicated by the message header information log may be determined according to the type of the message header information log; and then it is determined whether the device corresponding to the first information set is a faulty device to determine whether to perform port containment processing.
[0095] Specifically, when a faulty device indirectly connected to the processor fails, the PCIe switch connected to the faulty device will first trigger port containment processing to perform fault processing on the faulty device. At this time, the PCIe switch will send a fault message to the processor. Since the PCIe switch has already triggered the port containment processing, the faulty device is currently undergoing fault processing. Therefore, in order to prevent the processor from performing port containment processing again due to the fault information, it is necessary to compare the source of the fault information to determine whether the device that sent the fault information is the device that has already triggered the port containment processing.
[0096] Specifically, when the type of the message header information log is a memory read-write type or an input-output read-write type, the space address indicated by the message header information log is determined based on the information address of the message header information log, so as to determine the first information set based on the space address; when the type of the message header information log is a configuration space read-write type, the first information set is determined based on the information address offset of the message header information log.
[0097] In some embodiments, the type of the message header information log may be determined based on the memory mapping space requested by the faulty device, but is not limited thereto. The present disclosure does not limit the manner of determining the type of the message header information log.
[0098] In some embodiments, when the type of the message header information log is a memory read-write type or an input-output read-write type, the message header information log can directly indicate the spatial address of the failed device, and then determine the failed device based on the spatial address, thereby obtaining the first information set.
[0099] In some embodiments, when the type of the message header information log is a configuration space read-write type, data at different positions of the message header information may indicate a first information set of a failed device.
[0100] In other words, when the type of the message header information log is a memory read-write type or an input-output read-write type, it is necessary to determine the specific device with the fault based on the space address indicated by the message header information log, thereby determining the first information set; and when the type of the message header information log is a configuration space read-write type, the message header information log can directly indicate the first information set, so the message header information log can be directly analyzed to determine the first information set.
[0101] For example, taking the 9-bit storage read / write type information in the message header information log as an example, the first 3 bits can be used to indicate bus information, the middle 3 bits can indicate device information, and the last 3 bits can indicate function information.
[0102] Furthermore, the first information set includes at least one of the following: bus information of the faulty device, device information of the faulty device, and function information of the faulty device.
[0103] In some optional embodiments, it is also possible to determine whether the device corresponding to the spatial address is a faulty device to determine whether the fault information is generated due to the faulty device triggering the port containment process, so as to avoid the fault information generated by the port containment process triggering the port containment process again.
[0104] Furthermore, when it is determined that the device corresponding to the first information set is a faulty device, or when it is determined that the device corresponding to the first information set is a PCIe switch connected to the faulty device, port containment processing is not performed; when it is determined that the device corresponding to the first information is not a faulty device, and the device corresponding to the first information set is not a PCIe switch connected to the fault, port containment processing is performed.
[0105] In other words, when the device corresponding to the first information set is a faulty device, or it is determined that the device corresponding to the first information set is a PCIe switch connected to the faulty device, it indicates that the above fault information is generated by the PCIe switch or the faulty device when the faulty device is being processed using port containment processing. At this time, the faulty device is already undergoing fault processing, and there is no need for the processor to perform port containment processing again, so port containment processing is not performed.
[0106] In other words, when the device corresponding to the first information set is not a faulty device, or when it is determined that the device corresponding to the first information set is not a PCIe switch connected to the faulty device, it means that the above fault information is not generated when the faulty device is handled using port containment processing, that is, there is currently a faulty device that requires fault handling, so port containment processing is performed.
[0107] For example, Figure 3 Taking the devices included as an example, when device 2 sends fault information to the processor, the register in the processor parses the message header information log. When it is determined that the fault source indicated by the message header information log is device 2 or PCIe switch 1, port containment processing is not performed and the fault processing is ended.
[0108] For example, Figure 3 Taking the devices included as an example, when device 2 sends fault information to the processor, the register in the processor parses the message header information log, and when it is determined that the fault source indicated by the message header information log is not device 2 or PCIe switch 1 (such as device 1, device 3, PCIe switch, etc.), port containment processing is performed.
[0109] In some embodiments, the second reporting information may be fault advisory reporting information, but is not limited thereto.
[0110] In some optional embodiments, the device identity information of the faulty device and the fault cause of the faulty device may be determined based on the first reported information or the second reported information; and then based on the device identity information and the fault cause, the fault may be processed using a preset solution.
[0111] Specifically, the first reporting information or the second reporting information may include information such as the device model and fault cause of the faulty device, and then, according to a preset mapping relationship table, the preset solution corresponding to the device model, fault cause and other information in the first reporting information or the second reporting information is searched by table lookup, so as to use the corresponding preset solution to handle the fault.
[0112] In summary, the device fault handling method proposed in the present disclosure includes: when the processor supports the root port programmable input and output, shielding the first fault handling mechanism of the processor, and configuring the third fault handling mechanism for the processor, or configuring the second fault handling mechanism and the third fault handling mechanism for the processor; responding to the fault information for the faulty device, determining that the faulty device is directly connected to the processor; using the third fault handling mechanism to perform fault handling on the faulty device. The method of the present disclosure configures the third processing mechanism for the processor when the processor supports the root port programmable input and output, so that when a device directly connected to the processor fails, different fault handling is adopted for different types of faults, so that when the fault is a timeout completion error, only reporting information is generated, and further based on the message header information log, it is determined whether to trigger the port containment processing, and it is verified whether the fault information is generated by the faulty device triggering the port containment processing, thereby avoiding the faulty device that has triggered the port containment processing from triggering the port containment processing again, and ensuring the normal operation of non-faulty devices.
[0113] In summary, the present disclosure has the following beneficial effects: 1. By configuring different fault handling mechanisms for processing according to the support capabilities of the processor, and using different fault handling mechanisms to handle faulty devices according to the connection relationship between the faulty device and the processing, it is avoided that different types of fault reports are all handled by EDPC technology, which leads to the interruption of normal PCIe device operation, and the processing granularity and accuracy of fault handling are improved.
[0114] 2. By configuring a second processing mechanism for the processor when the processor supports programmable input and output of the root port, different fault handling methods are used for different types of faults when a fault occurs in a device directly connected to the processor. Thus, when the fault is a failure to support a configuration space request, only reporting information is generated and port containment processing is not triggered, thereby avoiding normal devices that have not completed initialization from triggering fault processing and ensuring the normal operation of non-faulty devices.
[0115] 3. By configuring the third processing mechanism for the processor when the processor supports programmable input and output of the root port, when a device directly connected to the processor fails, different fault processing is adopted for different types of faults, so that when the fault is a timeout completion error, only reporting information is generated, and further based on the message header information log, it is determined whether to trigger the port containment processing, and it is verified whether the fault information is generated by the faulty device triggering the port containment processing, thereby avoiding the faulty device that has triggered the port containment processing from triggering the port containment processing again, and ensuring the normal operation of non-faulty devices.
[0116] The following is an exemplary description of the present disclosure: Embodiment 1, with Figure 3 When a fatal error occurs during the enumeration of device 1, execute Fig. 9 The method shown in the figure is used to handle the fault, including the following steps: When device 1 fails, the triggering mode of the PCIe root port DPC can be configured to be RP PIO (that is, the second fault handling mechanism mentioned above). At this time, by configuring the RP PIO related registers, some error types can trigger the DPC function of the root port, while other error types do not trigger the DPC function of the root port, and only fault advisory reporting is performed.
[0117] 1. Check whether the PCIe root port supports the DPC function of the RP PIO mechanism.
[0118] 2. When the PCIe root port supports the DPC function of the RP PIO mechanism, the PCIe root port is initialized as follows: the AER mechanism is disabled, the RP PIO mechanism is configured, and the UR error of the configuration space request is set not to trigger the DPC, and the other errors are set to trigger the DPC (i.e., the second fault handling mechanism mentioned above).
[0119] 3. When the PCIe root port does not support the DPC function of the RP PIO mechanism, the PCIe root port is initialized: the traditional AER mechanism is configured to trigger DPC (that is, the first fault handling mechanism mentioned above).
[0120] 4. The PCIe root port as the requester determines that device 1 has a fatal fault (ie, the above processor receives fault information of the faulty device).
[0121] 5. The PCIe root port determines whether the processor supports the RP PIO mechanism.
[0122] 6. When the PCIe root port does not support the RP PIO mechanism, the fatal error triggers the DPC using the legacy AER mechanism.
[0123] 7. When the PCIe root port supports the RP PIO mechanism, it is further determined whether the fatal error is a UR error of the configuration space request.
[0124] 8. If the fatal error is a UR error in the configuration space request, the fatal error does not trigger the DPC, and only a advisory fault report is made.
[0125] 9. If the fatal error is not a UR error of the configuration space request, the fatal error triggers the DPC using the RP PIO mechanism.
[0126] Embodiment 2, with Figure 3 When a fatal error occurs during the enumeration of device 2, execute Fig.10 The method shown in the figure is used to handle the fault, including the following steps: If DPC is triggered at the downstream port of PCIe switch 1 (i.e., device 2), and this may be accompanied by one or more completion timeout errors in the PCIe root device. The present disclosure uses the RP PIO mechanism to filter out CTO faults, and accurately locates the source of the CTO fault to distinguish whether the downstream port of PCIe switch 1 triggers the DPC and the accompanying CTO fault, so as to accurately handle the fault error, avoid DPC caused by the accompanying CTO fault RootPort, and allow the PCIe root device to continue to operate normally with other PCIe switches (PCIe switch 2) and its downstream port (device 3).
[0127] 1. During the device initialization phase, check whether the PCIe root port supports the DPC function of the RP PIO mechanism.
[0128] 2. When the PCIe root port supports the DPC function of the RP PIO mechanism, the PCIe root port is initialized: the AER mechanism is disabled, the RP PIO mechanism is configured, and it is set not to trigger DPC when a completion timeout error (Cfg CTO, I / O CTO, Mem CTO error) occurs. The other errors are set to trigger DPC (that is, the third fault handling mechanism mentioned above).
[0129] 3. When the PCIe root port does not support the DPC function of the RP PIO mechanism, the PCIe root port is initialized: the traditional AER mechanism is configured to trigger DPC.
[0130] 4. The PCIe root port detects a fatal fault as a requester.
[0131] 5. When the PCIe root port does not support the RP PIO mechanism, the fatal error triggers the DPC using the legacy AER mechanism.
[0132] 6. When the PCIe switch 1 supports the RP PIO mechanism, it is further determined whether the fatal error is a Cfg CTO, I / O CTO, or Mem CTO error.
[0133] 7. If the fatal error is at least one of the Cfg CTO, I / O CTO, and Mem CTO errors, DPC will not be triggered temporarily, and only a suggestive fault report will be made. The suggestive fault report can indicate the identity information of the faulty device, the cause of the fault, etc., to facilitate subsequent analysis and processing.
[0134] 8. If the fatal error is not at least one of the Cfg CTO, I / O CTO, and Mem CTO errors, the RPPIO mechanism is used to trigger the DPC.
[0135] 9. Further, when the user software receives Cfg CTO, I / O CTO, MemCTO errors discovered by PCIe as a requester, it is necessary to use the root port processor to parse the message header information log (RP PIO Header Log) to determine the source of the CTO error.
[0136] 10. Further, determine the type of the message header information log. If the parsing type is a memory space read-write type, determine the device space address to which the information address belongs based on the information address indicated by the message header information log, and record the bus information, device information and function information (bus, devices, function) of the device. The type of information log can be determined by collecting the memory mapping (MMIO) space requested by the device.
[0137] 11. If the parsing type is the configuration space read-write type, the bus information, device information and function information of the device indicated by the message header information log are determined according to the information address offset of the message header information log.
[0138] 12. Determine whether the device indicated by the message header information log is device 2.
[0139] 13. If the device indicated by the message header information log is device 2, no processing is performed and the fault handling ends.
[0140] 14. If the device indicated by the message header information log, the PCIe root port performs DPC.
[0141] In order to implement the device fault processing method provided by the embodiment of the present disclosure, the embodiment of the present disclosure also provides a device fault processing device, such as Fig.11 As shown, the device failure processing device 1100 includes: An initialization unit 1101 is used to initialize and configure the processor based on a fault handling mechanism supported by the processor; A determining unit 1102, configured to determine a connection mode between the faulty device and the processor in response to the fault information of the faulty device; The processing unit 1103 is used to perform fault processing on the faulty device based on the connection mode and / or initialization configuration.
[0142] In some embodiments, the initialization unit 1101 is also used to shield the first fault handling mechanism of the processor when the processor supports the root port programmable input and output, and configure the second fault handling mechanism and / or the third fault handling mechanism for the processor; when the processor does not support the root port programmable input and output, configure the first fault handling mechanism for the processor.
[0143] In some embodiments, the first fault handling mechanism includes: when a fault exists in a faulty device, port containment processing is triggered; the second fault handling mechanism includes: when the fault of the faulty device is a failure to support configuration space requests, port containment processing is not triggered; when the fault of the faulty device is a fault other than failure to support configuration space requests, port containment processing is triggered; the third fault handling mechanism includes: when the fault of the faulty device is a completion timeout error, port containment processing is not triggered; when the fault of the faulty device is an error other than completion timeout error, port containment processing is triggered.
[0144] In some embodiments, the connection mode between the faulty device and the processor includes: the faulty device is directly connected to the processor, and the faulty device is indirectly connected to the processor via a switch.
[0145] In some embodiments, the processing unit 1103 is further configured to trigger port containment processing when the first fault handling mechanism is initialized for configuration; and to perform fault handling on the faulty device based on the connection mode when the second fault handling mechanism or the third fault handling mechanism is initialized for configuration.
[0146] In some embodiments, the processing unit 1103 is also used to perform fault processing on the faulty device using the second fault processing mechanism when the faulty device is directly connected to the processor; and to perform fault processing on the faulty device using the third fault processing mechanism when the faulty device is indirectly connected to the processor via a switch.
[0147] In some embodiments, the processing unit 1103 is further configured to determine, based on the fault information, whether the fault of the faulty device is that it does not support the configuration space request; and perform fault processing on the faulty device based on the result of the determination.
[0148] In some embodiments, the processing unit 1103 is also used to perform fault processing on the faulty device based on the judgment result, including: when the judgment result is that the fault of the faulty device does not support the configuration space request, generating the first fault reporting information, and not triggering the port containment processing; when the judgment result is that the fault of the faulty device does not not support the configuration space request, triggering the port containment processing.
[0149] In some embodiments, the processing unit 1103 is further configured to determine, based on the fault information, whether the fault of the faulty device is a completion timeout error; and perform fault processing on the faulty device based on the result of the determination.
[0150] In some embodiments, the processing unit 1103 is also used to trigger port containment processing when the judgment result is that the fault of the faulty device is not a completion timeout error; when the judgment result is that the fault of the faulty device is a completion timeout error, generate second reporting information of the fault, and obtain the message header information log corresponding to the fault, so as to determine whether to perform port containment processing based on the message header information log.
[0151] In some embodiments, the processing unit 1103 is also used to determine the space address indicated by the message header information log based on the type of the message header information log, so as to determine the first information set based on the space address; when the type of the message header information log is a configuration space read and write type, the first information set is determined based on the information address offset of the message header information log.
[0152] In some embodiments, the processing unit 1103 is also used to determine the space address indicated by the message header information log based on the type of the message header information log, including: when the type of the message header information log is a memory read-write type or an input-output read-write type, the space address is determined based on the information address of the message header information log; when the type of the message header information log is a configuration space read-write type, the space address is determined based on the information address offset of the message header information log.
[0153] In some embodiments, the processing unit 1103 is also used to not perform port containment processing when it is determined that the device corresponding to the first information set is a faulty device, or when it is determined that the device corresponding to the first information set is a switch connected to the faulty device; and to perform port containment processing when it is determined that the device corresponding to the first information set is not a faulty device, and the device corresponding to the first information set is not a switch connected to the fault.
[0154] In some embodiments, the processing unit 1103 is further configured to determine the type of the message header information log based on the memory mapping space requested by the faulty device.
[0155] In some embodiments, the processing unit 1103 is further used to determine the device identity information of the faulty device and the fault cause of the faulty device based on the first reporting information or the second reporting information; and based on the device identity information and the fault cause, perform fault processing using a preset solution.
[0156] In some embodiments, the processing unit 1103 is further used to uninstall the driver of the faulty device and remove the device tag of the faulty device; release the link state between the faulty device and the processor; enumerate the faulty device and re-enable the driver of the faulty device.
[0157] In summary, the device fault handling apparatus proposed in the present disclosure includes an initialization unit for initializing and configuring the processor based on a fault handling mechanism supported by the processor; a determination unit for determining the connection mode between the faulty device and the processor in response to the fault information of the faulty device; and a processing unit for performing fault handling on the faulty device based on the connection mode and / or the initialization configuration. The apparatus of the present disclosure configures different fault handling mechanisms for the processing according to the support capability of the processor, and performs fault handling for the faulty device using different fault handling mechanisms according to the connection relationship between the faulty device and the processing, thereby avoiding the interruption of normal PCIe device operation caused by the use of EDPC technology for fault handling of different types of fault reports, and improving the processing granularity and accuracy of fault handling.
[0158] It should be noted that: the device fault handling device provided in the above embodiment only uses the division of the above program modules as an example when performing device fault handling. In actual applications, the above processing can be assigned to different program modules as needed, that is, the internal structure of the device fault handling device is divided into different program modules to complete all or part of the processing described above. In addition, the device fault handling device provided in the above embodiment and the device fault handling method embodiment provided in the embodiment of the present disclosure belong to the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0159] Fig.12 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present disclosure is shown in FIG. Fig.12 As shown, the electronic device 1200 includes at least one processor 1202; and a memory 1201 communicatively connected to the at least one processor 1202; wherein the memory 1201 stores commands that can be executed by the at least one processor 1202, and the commands are executed by the at least one processor 1202 to implement the steps of the device fault handling method described in the embodiment of the present disclosure.
[0160] Optionally, the electronic device may specifically be a device fault handling device in an embodiment of the present application, and the electronic device may implement the corresponding processes implemented by the device fault handling device in each method in the embodiment of the present application, which will not be described in detail here for the sake of brevity.
[0161] It is understood that the electronic device also includes a communication interface 1203. The various components in the electronic device are coupled together through a bus system 1204. It is understood that the bus system 1204 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 1204 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, Fig.12 Various buses are labeled as bus system 1204.
[0162] It can be understood that the memory 1201 can be a volatile memory or a non-volatile memory, and can also include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disk, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. The present method is described by way of example but not limitation, and many forms of RAM may be used, such as static random access memory (SRAM), synchronous static random access memory (SSRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM, SyncLink Dynamic Random Access Memory), and direct RAM bus random access memory (DRRAM, Direct Rambus Random Access Memory).The memory 1201 described in the embodiments of the present invention is intended to include but is not limited to these and any other suitable types of memories.
[0163] The method disclosed in the above embodiment of the present disclosure can be applied to the processor 1202, or implemented by the processor 1202. The processor 1202 has the ability to process signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 1202 or the command in the form of software. The above processor 1202 can be a general-purpose processor, a DSP, or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The processor 1202 can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor can be a microprocessor or any conventional processor, etc. In combination with the steps of the method disclosed in the embodiment of the present invention, it can be directly embodied as a hardware decoding processor to execute, or it can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium, which is located in the memory 1201, and the processor 1202 reads the information in the memory 1201 and completes the steps of the above method in combination with its hardware.
[0164] In an exemplary embodiment, the electronic device may be implemented by one or more application specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), FPGAs, general purpose processors, controllers, MCUs, microprocessors, or other electronic components to execute the aforementioned method.
[0165] The embodiment of the present disclosure also provides a non-transitory computer-readable storage medium storing computer commands, wherein the computer commands are used to enable the computer to implement the steps of the device failure handling method described in the embodiment of the present disclosure when executed.
[0166] The embodiment of the present disclosure further provides a computer program product, including a computer program, which implements the steps of the device failure handling method described in the embodiment of the present disclosure when executed by a processor.
[0167] Optionally, the computer-readable storage medium can be applied to the equipment fault handling device in the embodiments of the present application, and the computer command enables the computer to execute the corresponding processes implemented by the equipment fault handling device in the various methods of the embodiments of the present application. For the sake of brevity, they will not be repeated here.
[0168] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.
[0169] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0170] In addition, all functional units in the embodiments of the present invention may be integrated into one processing unit, or each unit may be separately used as a unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.
[0171] A person of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiment can be completed by hardware related to program commands, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiment; and the aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROM, RAM, disks or optical disks.
[0172] Alternatively, if the above-mentioned integrated unit of the present invention is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present invention can be essentially or partly reflected in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several commands to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROM, RAM, magnetic disks or optical disks.
[0173] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A method for handling equipment failure, characterized in that: The method comprises: Initializing configuration of the processor based on a fault handling mechanism supported by the processor; In response to the fault information of the faulty device, determining a connection mode between the faulty device and the processor; Based on the connection mode and / or the initialization configuration, fault processing is performed on the faulty device.
2. The method according to claim 1, characterized in that The initialization configuration of the processor based on the fault handling mechanism supported by the processor includes: When the processor supports root port programmable input and output, shielding the first fault handling mechanism of the processor, and configuring the second fault handling mechanism and / or the third fault handling mechanism for the processor; When the processor does not support the root port programmable input and output, the first fault handling mechanism is configured for the processor.
3. The method according to claim 2, characterized in that The first fault handling mechanism includes: triggering port containment processing when the faulty device is faulty; The second fault handling mechanism includes: when the fault of the faulty device is that the configuration space request is not supported, the port containment processing is not triggered, and when the fault of the faulty device is a fault other than the failure to support the configuration space request, the port containment processing is triggered; The third fault processing mechanism includes: when the fault of the faulty device is a completion timeout error, the port containment processing is not triggered; when the fault of the faulty device is an error other than the completion timeout error, the port containment processing is triggered.
4. The method according to claim 1, characterized in that The connection mode between the faulty device and the processor includes: the faulty device is directly connected to the processor, and the faulty device is indirectly connected to the processor via a switch.
5. The method according to claim 3, characterized in that: The performing fault handling on the faulty device based on the connection mode and / or the initialization configuration includes: When the first fault handling mechanism is configured in the initialization, a port containment process is triggered; When the second fault handling mechanism or the third fault handling mechanism is configured by initialization, fault handling is performed on the faulty device based on the connection mode.
6. The method according to claim 5, characterized in that When the second fault handling mechanism or the third fault handling mechanism is initially configured, performing fault handling on the faulty device based on the connection mode includes: When the faulty device is directly connected to the processor, using the second fault handling mechanism to handle the fault of the faulty device; When the faulty device is indirectly connected to the processor via a switch, the third fault handling mechanism is used to perform fault handling on the faulty device.
7. The method according to claim 6, characterized in that When the faulty device is directly connected to the processor, using the second fault handling mechanism to handle the fault of the faulty device includes: Based on the fault information, determining whether the fault of the faulty device is that the configuration space request is not supported; Based on the judgment result, the fault processing is performed on the faulty device.
8. The method according to claim 7, characterized in that The performing fault processing on the faulty device based on the judgment result includes: When the result of the judgment is that the fault of the faulty device is the failure to support the configuration space request, generating first reporting information of the fault, and not triggering the port containment processing; When the result of the determination is that the fault of the faulty device is not the failure to support the configuration space request, the port containment process is triggered.
9. The method according to claim 6, characterized in that When the faulty device is indirectly connected to the processor through the switch, performing fault processing on the faulty device by using the third fault processing mechanism includes: Based on the fault information, determining whether the fault of the faulty device is a completion timeout error; Based on the judgment result, the fault processing is performed on the faulty device.
10. The method according to claim 9, characterized in that The performing fault processing on the faulty device based on the result of the judgment includes: When the result of the judgment is that the fault of the faulty device is not the completion timeout error, triggering the port containment process; When the result of the judgment is that the fault of the faulty device is the completion timeout error, second reporting information of the fault is generated, and a message header information log corresponding to the fault is obtained to determine whether to perform the port containment processing based on the message header information log.
11. The method according to claim 10, characterized in that The determining whether to perform the port containment process based on the message header information log includes: Based on the type of the message header information log, determining a first information set indicated by the message header information log; It is determined whether the device corresponding to the first information set is the faulty device to determine whether to perform the port containment process.
12. The method according to claim 11, characterized in that: The determining, based on the type of the message header information log, the space address indicated by the message header information log comprises: When the type of the message header information log is a memory read-write type or an input-output read-write type, determining a space address indicated by the message header information log based on the information address of the message header information log, so as to determine the first information set based on the space address; When the type of the message header information log is a configuration space read-write type, the first information set is determined based on the information address offset of the message header information log.
13. The method according to claim 11, characterized in that The determining whether the device corresponding to the first information set is the faulty device to determine whether to perform the port containment process includes: When it is determined that the device corresponding to the first information set is the faulty device, or when it is determined that the device corresponding to the first information set is a switch connected to the faulty device, the port containment processing is not performed; When it is determined that the device corresponding to the first information set is not a faulty device, and the device corresponding to the first information set is not a switch connected to the faulty device, the port containment process is performed.
14. The method according to claim 11, characterized in that The method further comprises: Based on the memory mapping space requested by the faulty device, the type of the message header information log is determined.
15. The method according to claim 14, characterized in that The first information set includes at least one of the following: bus information of the faulty device, device information of the faulty device, and function information of the faulty device.
16. The method according to any one of claims 8 or 10, characterized in that The method further comprises: Determine, based on the first reporting information or the second reporting information, device identity information of the faulty device and a fault cause of the faulty device; Based on the device identity information and the cause of the fault, the fault is handled using a preset solution.
17. The method according to any one of claims 3 or 5, characterized in that The port containment process includes: Uninstalling the driver of the faulty device and removing the device mark of the faulty device; releasing the link state between the faulty device and the processor; The faulty device is enumerated and a driver of the faulty device is re-enabled.
18. An electronic device, characterized in that: include: A processor and a memory for storing a computer program that can be run on the processor, wherein the processor is used to execute the device failure processing method according to any one of claims 1-17 when running the computer program.
19. A non-transitory computer-readable storage medium storing computer commands, characterized in that: The computer command is used to enable the computer to execute the device failure processing method according to any one of claims 1-17.
20. A computer program product, characterized in that The invention comprises a computer program, which, when executed by a processor, implements the device failure processing method according to any one of claims 1 to 17.
Citation Information
Patent Citations
Endpoint device management method, device and system
CN112306913A
Failure characterization systems and methods for programmable logic devices
CN112470158A
Dynamic memory allocation method and device for programmable logic controller and electronic equipment
CN119512758A
Equipment fault processing method and device, equipment and medium
CN119621400A
Apparatus and system having PCI root port and direct memory access device functionality
US20110246686A1
Cited By
Dynamic random access memory configuration method and device, electronic equipment and storage medium
CN121034363A