Method for testing wrong disconnection recovery function and related product

By establishing container storage device drivers in the server system, the problems of PCIe device downtime due to uncorrectable errors and device adaptation difficulties in EDR function testing are solved, and efficient automated testing is achieved.

CN121833370APending Publication Date: 2026-04-10西安远图未来科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-24
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

In server systems, PCIe devices can crash due to uncorrectable errors, and the compatibility issues between devices and drivers in EDR functional testing limit testing efficiency.

Method used

By creating a container to store the driver for the device under test and directly calling the relevant driver during testing, automated testing is achieved, thus solving the compatibility problem between the device and the driver.

Benefits of technology

It improves the efficiency of EDR functional testing, ensures device and driver compatibility, realizes automated testing, and avoids the inefficiency and error-prone problems of manual updates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121833370A_ABST
    Figure CN121833370A_ABST
Patent Text Reader

Abstract

The invention discloses a method for testing an error disconnection recovery function and a related product. The method provided by the invention comprises the following steps: establishing a container for storing a driving program related to equipment to be tested; and calling a driving program related to the to-be-tested equipment from the container in response to an error disconnection recovery function test performed on the to-be-tested equipment. According to the method and the device, the container is used for storing the driving program related to the to-be-tested device, and the driving program related to the to-be-tested device can be directly called when EDR function testing is carried out on various devices in a server system, so that the problem that adaptation between the to-be-tested device and the driving program is difficult can be solved; moreover, the driving programs related to the to-be-tested equipment are stored in the container in advance, so that automatic testing is realized, and the efficiency of testing work is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application generally relates to the field of server system technology. More specifically, this application relates to a method and related products for testing fault disconnection recovery functionality. Background Technology

[0002] For server systems with large end-user demands and complex computational operations, PCIe devices inevitably encounter uncorrectable errors during actual operation due to factors such as bit error rate, radiation, and environmental conditions. These errors cannot be handled by the hardware itself or the CPU, ultimately leading to server system crashes. The PCIe specification introduces Downstream Port Containment (DPC) and Error Disconnect Recover (EDR) technologies to handle and recover from these errors. DPC technology suppresses errors by disconnecting the faulty device from propagating the resistance error upwards. After the faulty device is disconnected, EDR technology restores the link using relevant functions in the device driver of the faulty device; the device driver is stored by the operating system.

[0003] For testers of server systems, they need to frequently perform EDR (Electronic Device Response) functional tests on various devices during the operation of the server system. When there are a large number of devices in the server, ensuring compatibility between the devices and drivers becomes a challenge. In addition, the operating system of the server system cannot guarantee that it has all the necessary device drivers, which limits the efficiency of the testing work.

[0004] In view of this, this application proposes a method and related products for fault disconnection recovery function testing, so as to store the device drivers associated with each device in the server system in the server operating system, solve the adaptation problem between the device under test and the driver in EDR testing, and improve the efficiency of EDR function testing. Summary of the Invention

[0005] In order to at least solve one or more of the technical problems mentioned above, this application proposes methods and solutions for fault disconnection recovery function testing in several aspects.

[0006] In a first aspect, this application provides a method for fault disconnection recovery function testing, characterized in that it includes: establishing a container for storing drivers associated with a device under test; and in response to performing fault disconnection recovery function testing on the device under test, calling the drivers associated with the device under test from the container.

[0007] In a second aspect, this application provides a system for fault disconnection recovery function testing, characterized in that it includes: a storage container configured to store drivers related to a device under test; and a test module configured to, in response to performing fault disconnection recovery function testing on the device under test, retrieve the drivers related to the device under test from the storage container.

[0008] In a third aspect, this application provides an electronic device, characterized in that it includes: a processor; and a memory having a program stored thereon for performing error disconnection recovery function tests on a server or switch, wherein when the program is executed by the processor, the electronic device performs the method as described in the first aspect.

[0009] Using the methods and related products for fault disconnection recovery function testing provided above, this application uses containers to store drivers related to the device under test. When performing EDR function testing on various devices in the server system, the drivers related to the device under test can be directly called, thereby solving the problem of difficult compatibility between the device under test and the driver. Moreover, the drivers related to the device under test are all stored in the container in advance, which helps to realize automated testing and improve the efficiency of testing work. Attached Figure Description

[0010] The above and other objects, features, and advantages of exemplary embodiments of this application will become readily apparent upon reading the following detailed description with reference to the accompanying drawings. In the drawings, several embodiments of this application are illustrated by way of example and not limitation, and like or corresponding reference numerals denote like or corresponding parts, wherein: Figure 1 Exemplary schematic diagrams of server systems employing the PCIe bus are shown in some embodiments of this application.

[0011] Figure 2 An exemplary flowchart of a method for EDR and DPC to work together to handle errors is shown in some embodiments of this application.

[0012] Figure 3 An exemplary schematic diagram of a system for testing fault disconnection recovery functionality is shown in some embodiments of this application. Detailed Implementation

[0013] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0014] It should be understood that the terms "comprising" and "including" used in the specification and claims of this application indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.

[0015] It should also be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the application. As used in this specification and claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this specification and claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations.

[0016] As used in this specification and claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if [described condition or event] is detected" may be interpreted, depending on the context, as "once determined," "in response to determination," "once [described condition or event] is detected," or "in response to detection of [described condition or event]."

[0017] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means two or more, unless otherwise expressly specified. "Several" means one or more, unless otherwise expressly specified.

[0018] The specific embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0019] Currently, information technology has permeated numerous industries, with sectors such as manufacturing, healthcare, finance, education, and agriculture all attempting to leverage it for intelligent transformation. As information technology develops, all industries are seeking more efficient ways to process and store information, and server systems are one such option. Generally, a server system includes storage resources, processing resources, input / output devices, and a bus for connecting various hardware components. This bus can be implemented based on the Peripheral Component Interconnect Express (PCIe) standard.

[0020] Figure 1 Exemplary schematic diagrams of server systems employing the PCIe bus in some embodiments of this application are shown. For example... Figure 1 As shown, in some embodiments, server system 100 includes: root complex (RC) 101, central processing unit (CPU) 102, main memory 103, switch 104, and endpoint device 105.

[0021] The root complex 101 is the starting point and root node of the PCIe hierarchy, directly connecting the CPU and main memory 103. Main memory 103 can be volatile memory used to provide the CPU 102 with the instructions and data required for computation in real time. Multiple CPUs 102 can be configured in the server system 100 to provide higher concurrent access and adapt to complex workloads.

[0022] continue Figure 1 The root complex 101, switch 104, and endpoint device 105 are all equipped with PCIe ports for establishing PCIe links. Switch 104 is used to expand PCIe connectivity, allowing one upstream port to be connected to multiple downstream ports. Endpoint device 105 is a leaf node in the PCIe hierarchy; in some embodiments, endpoint device 105 includes storage devices for storing data and instructions, such as high-speed solid-state drives and cache devices; in some embodiments, endpoint device 105 includes network devices for network connectivity, such as Ethernet network cards and wireless network cards; in some embodiments, endpoint device 105 includes input / output devices, such as mice, keyboards, and video displays. Artificial intelligence (AI) technology has demonstrated a powerful influence and application value in fields such as natural language processing and computer vision. Neural network models are the core algorithms of AI technology; the development of AI technology has spawned large models with hundreds of billions of parameters. The training and inference of these neural network models involve massive parallel computing and massive data throughput. In some embodiments, the endpoint device 105 includes an artificial intelligence processor to support AI application scenarios, such as Google's Tensor Processing Unit (TPU), Qualcomm's Cloud AI 100 chip, and NVIDIA's A100, H100, H200 and other graphics processing units (GPUs).

[0023] exist Figure 1In the server system 100, only a portion of the hardware components are shown. It will be apparent to those skilled in the art that the server system 100 may also include components related to… Figure 1 The hardware components shown are different from common hardware components. Furthermore, the server system 100 achieves flexible data manipulation through hardware and software collaboration. In some embodiments, the server system 100 can run operating systems such as Windows or Linux and implement various functions based on instruction sets.

[0024] For server systems with large end-user demands and complex computational operations, PCIe devices inevitably encounter uncorrectable errors during actual operation due to factors such as bit error rate, radiation, and environmental conditions. These errors cannot be handled by the hardware itself or the CPU, ultimately leading to server system crashes. The PCIe specification introduces downstream port containment technology and error disconnection recovery technology to handle and recover from these types of errors.

[0025] Figure 2 An exemplary flowchart of a method 200 in which EDR and DPC work together to handle errors is shown in some embodiments of this application. Figure 2 As shown, method 200 includes: S201: An error has occurred. For example, a PCIe device (such as endpoint device 105, switch 104) has a serious physical or protocol layer error, and the error is determined to be unprocessable by the PCIe device itself or the CPU.

[0026] S202: DPC triggered, disconnecting and isolating the faulty device. In some embodiments, the downstream port of the faulty device is connected to the upstream port of another PCIe device via a PCIe Link. When the faulty device experiences an error that it or the CPU cannot handle, the downstream port of the faulty device triggers the DPC mechanism to disconnect the PCIe Link between the upstream and downstream ports and set the value in the DPC status register to 1.

[0027] S203: Entering the EDR event handler. In some embodiments, the downstream port of the faulty device sends an EDR event to the operating system so that the operating system can start the EDR event handler.

[0028] S204: The EDR event handler calls relevant function calls from the device driver associated with the faulty device to perform recovery operations and rebuild the link. The device driver has complete control and state knowledge of the device. In some embodiments, the EDR event handler calls error handling or recovery callback functions in the device driver to perform recovery operations and rebuild the link. In some embodiments, the recovery operation includes: clearing the DPC trigger status bit and other error registers in the hardware; selectively resetting the device according to the severity of the error; requesting link retraining to rebuild the link; and after successful link rebuilding, the downstream port in the faulty device sets the value in the DPC status register to 0.

[0029] As understood from method 200, DPC technology prevents errors from propagating upwards by disconnecting faulty devices, thus suppressing errors. After the faulty device is disconnected, EDR technology restores the link using relevant function calls in the device driver of the faulty device; the device driver is stored by the operating system. For testers of server systems, they frequently perform EDR functional tests on various devices during server system operation. When there are many devices in a server, ensuring compatibility between devices and drivers becomes a challenge. Furthermore, the server operating system may not have all the necessary device drivers, limiting the efficiency of testing. Therefore, this application proposes a method and related products for fault disconnection recovery function testing, which allows the storage of device drivers associated with various devices in the server system within the server operating system, solving the compatibility problem between the device under test and its driver during EDR testing, thereby improving the efficiency of EDR functional testing.

[0030] In some embodiments, the method for fault disconnection recovery function testing disclosed in this application includes: establishing a container for storing drivers associated with the device under test; and in response to performing fault disconnection recovery function testing on the device under test, calling the drivers associated with the device under test from the container.

[0031] The method disclosed in this application for testing error disconnection recovery function can be used for, for example... Figure 1Each PCIe device in the server system 100 shown undergoes EDR (Electronic Data Reduction) function testing to determine whether the PCIe device's EDR function is functioning correctly. In some embodiments, the server system 100 includes multiple devices under test, such as multiple endpoint devices 105 in the server 100 constituting multiple devices under test, thereby allowing the acquisition of the driver for each device under test and storing the drivers for multiple devices under test in a container. For those skilled in the art, the storage structure of the container can be flexibly chosen, and this application is not limited in this regard. Exemplarily, the container can store the driver for each device under test in a key-value format, where the key corresponds to the name of the device under test and the value corresponds to the driver file.

[0032] In some embodiments, the server system under test uses a Linux operating system. Before testing, access the Linux system of the test machine and use the SCE tool to export the current input / output system configuration file, denoted as configuration file X. In configuration file X, enable the DPC and EDR function switches to generate a new configuration file, denoted as configuration file Y. Use the SCE tool to import configuration file Y into the test machine, reset the test machine, and then start the testing process. This ensures that the server system under test supports DPC and EDR functions, avoiding invalid tests due to configuration not being enabled.

[0033] In some embodiments, when performing EDR (Electronic Demand Reduction) functional testing on the device under test, a PCIe link needs to be established between the device under test and other PCIe devices to facilitate reconnecting the device under test to the PCIe bus. In these embodiments, the driver for the device under test stores function calls for error handling or link recovery, allowing the relevant function calls to be retrieved from the container's driver associated with the device under test, thus enabling the EDR function. If the driver for the device under test is successfully invoked and a PCIe link is established between the device under test and other PCIe devices, the EDR function of the device under test is considered to be functioning correctly; if the device under test fails to establish a PCIe link with other PCIe devices, the EDR function of the device under test is considered to be faulty.

[0034] Understandably, this application uses containers to store drivers related to the device under test. When performing EDR functional tests on various devices in the server system, the drivers related to the device under test can be directly called, thereby solving the problem of difficult compatibility between the device under test and the driver. Moreover, the drivers related to the device under test are all stored in the container in advance, which helps to realize automated testing and improve the efficiency of the testing work.

[0035] In some embodiments, a script for updating the driver is stored in the container.

[0036] It is understandable that during the operation of a server system, devices may be replaced or the same device manufacturer may continuously update drivers. In such cases, incompatibility between the old driver and the device under test can cause the EDR function to malfunction. In these embodiments, the drivers stored in the container are updated automatically using an update script, thereby ensuring compatibility between the driver and the device under test and avoiding test result errors caused by driver incompatibility. Moreover, since the update operation is run by a script, the inefficiency and error-proneness of manual updates are overcome.

[0037] In some embodiments, a script is used to verify whether the version of the driver stored in the container for the device under test is consistent with the official version of the device under test; in response to the verification result being inconsistent, the driver stored in the container for the device under test is updated to the official version.

[0038] To fix bugs or upgrade drivers, device manufacturers release updated driver versions on their official websites. In these embodiments, the driver stored in the container is compared and updated according to the version released on the official website to ensure that the driver is the latest version compatible with the device under test. In some embodiments, the method proposed in this application is used to perform EDR functional testing on a server system. After the server system is connected to the network, the update script accesses the official website of the device under test and verifies the version of the driver stored in the container. If the version of the driver stored in the container is inconsistent with the latest version released on the official website, the driver in the container is updated to the latest version of the driver released on the official website.

[0039] In some embodiments, performing an error disconnection recovery function test on the device under test includes sending a simulated fault signal to the device under test to trigger a downstream port containment mechanism.

[0040] In these embodiments, uncorrectable errors in real-world scenarios are simulated via software to trigger the DCP mechanism. In some embodiments, an in-band error injection module is invoked, and a corresponding script inputs parameter Z into the simulated error injection command to inject an uncorrectable error into the device under test and trigger the DCP mechanism. In some embodiments, if error injection fails, simulated fault signals continue to be sent to the device under test until error injection is confirmed to be successful.

[0041] In some embodiments, the fault disconnection recovery function of the device under test is determined based on the registers related to the downstream port containment mechanism in the device under test; and / or based on whether the operating system receives a fault disconnection recovery event.

[0042] Before DPC is triggered, the downstream port in the device under test is connected to the upstream port in another PCIe device via a PCIe Link. Triggering DPC disconnects the PCIe Link between the upstream and downstream ports, and sets the DPC status register in the downstream port to 1. If the EDR function of the device under test is normal, after DPC is triggered: the operating system receives an EDR event; after calling the driver of the device under test and rebuilding the PCIe Link between the upstream and downstream ports, the value of the DPC status register in the downstream port is changed from 1 to 0. If the EDR function of the device under test is faulty, after DPC is triggered: the operating system will not receive an EDR event, and the value of the DPC status register in the downstream port remains 1. In some embodiments, if the operating system receives an EDR event and the value of the DPC status register in the downstream port of the device under test is set to 0, it is determined that the EDR function of the device under test is normal. In some embodiments, if the operating system does not receive an EDR event and the value of the DPC status register in the downstream port of the device under test remains 1, it is determined that the EDR function of the device under test is faulty.

[0043] In some embodiments, the device under test includes an artificial intelligence processor or a network device. In some embodiments, the device under test includes a tensor processing unit or a graphics processing unit. In some embodiments, the device under test includes an Ethernet network interface card (NIC) or a wireless network interface card (NIC).

[0044] This application also discloses a system for testing error disconnection recovery functionality.

[0045] Figure 3 An exemplary schematic diagram of a system 300 for fault disconnection recovery function testing according to some embodiments of this application is shown. Figure 3 As shown, in some embodiments, system 300 includes: a storage container 301 configured to store drivers associated with the device under test; and a test module 302 configured to retrieve drivers associated with the device under test from the storage container in response to performing an error disconnection recovery function test on the device under test.

[0046] Corresponding to the method for fault disconnection recovery function testing disclosed above, this application discloses the following embodiments for a system for fault disconnection recovery function testing: In some embodiments, the system includes an update module configured to update drivers stored in a storage container.

[0047] In some embodiments, the update module is further configured to: verify whether the version of the driver stored in the storage container for the device under test is consistent with the official version of the device under test; and update the driver stored in the storage container for the device under test to the official version in response to the verification result being inconsistent.

[0048] In some embodiments, the test module is further configured to: send a simulated fault signal to the device under test in order to trigger a downstream port containment mechanism.

[0049] In some embodiments, the test module is further configured to: determine whether the fault disconnection recovery function of the device under test is normal based on the registers related to the downstream port containment mechanism in the device under test; and / or determine whether the fault disconnection recovery function of the device under test is normal based on whether the operating system receives a fault disconnection recovery event.

[0050] In some embodiments, the device under test includes an artificial intelligence processor or a network device.

[0051] This application also discloses an electronic device, including: a processor; and a memory storing a program for performing error disconnection recovery function tests on a server or switch. When the program is executed by the processor, the electronic device implements the method for error disconnection recovery function testing in any of the preceding embodiments.

[0052] In summary, the specific functions implemented by the system 300 and electronic device for fault disconnection recovery function testing provided in this specification can be explained in comparison with the aforementioned embodiments in this specification, and can achieve the technical effects of the aforementioned embodiments. Therefore, they will not be repeated here.

[0053] While numerous embodiments of this application have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many modifications, alterations, and alternatives will arise for those skilled in the art without departing from the spirit and intent of this application. It should be understood that various alternatives to the embodiments of this application described herein may be employed in the practice of this application. The appended claims are intended to define the scope of protection of this application and therefore cover equivalents or alternatives within the scope of these claims.

Claims

1. A method for testing of error disconnect recovery functionality, characterized by, comprising: establishing a container for storing a driver associated with a device under test; in response to performing a link down recovery function test on the device under test, invoking the driver associated with the device under test from the container.

2. The method of claim 1, wherein, comprising: storing a script for updating the driver in the container.

3. The method of claim 2, wherein, comprising: verifying, by the script, whether the version of the driver stored by the container for the device under test is consistent with the official version of the device under test; in response to the result of the verification being inconsistent, updating the driver stored by the container for the device under test to the official version.

4. The method of claim 1, wherein, performing a link down recovery function test on the device under test comprises: sending a simulated fault signal to the device under test to trigger a downstream port containment mechanism.

5. The method of claim 4, wherein, comprising: determining whether the link down recovery function of the device under test is normal according to a register associated with the downstream port containment mechanism in the device under test; and / or determining whether the link down recovery function of the device under test is normal according to whether an operating system receives a link down recovery event.

6. The method of any one of claims 1-5, wherein: the device under test comprises an artificial intelligence processor or a network device.

7. A system for error disconnect recovery function testing, characterized by, comprising: a storage container configured to: store a driver associated with a device under test; a test module configured to: in response to performing a link down recovery function test on the device under test, invoke the driver associated with the device under test from the storage container.

8. The system of claim 7, wherein, comprising: an update module configured to: update the driver stored in the storage container.

9. The system of claim 8, wherein, the update module is further configured to: verify whether the version of the driver stored by the storage container for the device under test is consistent with the official version of the device under test; in response to the result of the verification being inconsistent, update the driver stored by the storage container for the device under test to the official version.

10. An electronic device, comprising: comprising: a processor; and a memory having stored thereon a program for performing a link down recovery function test on a server or a switch, which when executed by the processor, causes the electronic device to implement the method of any one of claims 1-6. ​