Fault recovery method, device, electronic device and storage medium

By clearing the target status register of PCIE devices and injecting errors, detecting and verifying the functions of DPC technology, it solves the problem of difficult to find different types of equipment in the prior art, and improves the stability and comprehensiveness of the system.

CN116244102BActive Publication Date: 2025-06-27INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211652976.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-20
Publication Date
2025-06-27
Estimated Expiration
2042-12-20

AI Technical Summary

Technical Problem

It is difficult to find PCIE devices of different types of failures in the prior art, resulting in the inability to accurately judge the role of DPC technology, and the system detection is incomplete, which affects the stability of the server.

Method used

By clearing the settings of the target status register in the PCIE device, obtaining device information, and injecting target type errors through the target error tool, detecting the DPC function, controlling the operating system to clear the error, and restoring the normal state of the PCIE device.

Benefits of technology

It realizes rapid verification of DPC functions in PCIE devices, avoids the difficulty in finding equipment problems of different types of failures, and improves the comprehensiveness and stability of system detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116244102B_ABST
    Figure CN116244102B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides a fault recovery method, apparatus, electronic device, and storage medium. The method includes: clearing the settings of each item in the target status register of the PCIE device and obtaining the first data transfer rate of the PCIE device and the working status of the downstream port suppression DPC; when it is detected that the first data transfer rate is consistent with the preset rate and the working status is started, injecting an error of the target type into the PCIE device through the target error injection tool and obtaining the setting information of the target status register of the PCIE device and the second data transfer rate of the PCIE device; when it is detected that the target status register is successfully set and the second data transfer rate of the PCIE device is less than the first data transfer rate, controlling the operating system to clear the injected target type error so that the PCIE device returns to normal, thereby ensuring the stability of the system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of servers, and in particular, to a fault recovery method, device, electronic device, and storage medium. Background Art

[0002] With the widespread promotion and application of cloud computing, more and more data centers need to be established. As an important infrastructure in the data center, the stability of servers directly affects the experience and value of cloud services. Peripheral Component Interconnect Express (PCIE) devices are important components of servers, and almost all peripheral hardware uses the PCIE protocol. When an uncorrectable fault occurs in a PCIE device, it will directly affect the operating system (OS) of the server, resulting in the server crashing. Therefore, ensuring the normal operation of PCIE devices is crucial.

[0003] Currently, the Downstream Port Containment (DPC) technology allows stopping PCIE communication below the downstream port after detecting an uncorrectable error at or below the port, avoiding the potential spread of any data corruption, which ensures error control. Therefore, the DPC technology is used in PCIE devices to detect and control errors during the operation of PCIE, and then the software is used to recover the errors.

[0004] However, because the types of errors are complex and diverse, it is only possible to detect whether the DCP can detect and control this type of error at the moment when a fault occurs in the PCIE device. Therefore, currently, a faulty PCIE device is generally used to verify whether the DCP can detect this type of error and perform error control under the Basic Input Output System (BIOS). For different types of errors, it is necessary to find different faulty PCIE devices, and such devices are very difficult to find, resulting in incomplete system detection and thus the stability of the system cannot be guaranteed. Summary of the Invention

[0005] The purpose of the embodiments of the present invention is to provide a fault recovery method, device, electronic device, and storage medium to solve the problem in the prior art that it is difficult to find PCIE devices with different fault types, resulting in the inability to accurately judge the role of the DPC technology in the device, and thus incomplete system detection and the inability to guarantee the stability of the system. The specific technical solutions are as follows:

[0006] In the first aspect of the present invention, a fault recovery method is first provided, and the method includes:

[0007] Clear the settings of each item in the target status register of a Peripheral Component Interconnect Express (PCIE) device in a high-speed serial computer expansion bus standard and obtain the first device information of the PCIE device, where the first device information includes: the first data transfer rate of the PCIE device and the operating status of Downstream Port Clocking (DPC);

[0008] When it is detected that the first data transfer rate is consistent with a preset rate and the operating status is enabled, inject an error of a target type into the PCIE device through a target error injection tool and obtain the second device information of the PCIE device, where the second device information includes: the setting information of the target status register and the second data transfer rate of the PCIE device;

[0009] When it is detected that the target status register is successfully set and the second data transfer rate of the PCIE device is less than the first data transfer rate, control the operating system to clear the injected target type of error so that the PCIE device returns to normal.

[0010] Optionally, the controlling the operating system to clear the injected target type of error so that the PCIE device returns to normal includes:

[0011] Controlling the operating system to clear the injected target type of error, controlling the current data transfer rate of the PCIE device to return to the first data transfer rate, and controlling to clear the setting information of the target status register so that the PCIE device returns to normal.

[0012] Optionally, after controlling the operating system to clear the injected target type of error so that the PCIE device returns to normal when it is detected that the target status register is successfully set and the second data transfer rate of the PCIE device is less than the first data transfer rate, the method further includes:

[0013] Obtaining the log information of the operating system and the script information of the target error injection tool;

[0014] Determining that the PCIE device returns to normal according to the log information and the script information.

[0015] Optionally, before clearing the settings of each item in the target status register of the PCIE device and obtaining the first device information of the PCIE device, the method further includes:

[0016] Configuring a test environment according to the characteristics of the target error injection tool and the DPC so that the target error injection tool and the DPC can function properly.

[0017] Optionally, after configuring the test environment according to the characteristics of the target fault injection tool and the DPC to enable the target fault injection tool and the DPC to function properly, the method further includes:

[0018] Controlling the target status register to stop updating and obtaining third device information of the PCIE device, where the third device information includes: the topology structure of the PCIE device;

[0019] Determining the location information of the target status register according to the topology structure of the PCIE device.

[0020] Optionally, the obtaining the first device information of the PCIE device includes:

[0021] Obtaining the cls value of the first register and the flag bit information of the second register in the PCIE device;

[0022] Determining the first data transfer rate of the PCIE device according to the cls value;

[0023] Determining the status information of the DPC according to the flag bit information.

[0024] Optionally, the injecting a target type of error into the PCIE device by the target fault injection tool includes:

[0025] Obtaining different types of error names in the target fault injection tool;

[0026] Injecting a target type of error into the PCIE device according to the error name.

[0027] In the second aspect of the implementation of the present invention, a fault recovery device is further provided, and the device includes:

[0028] A first module, configured to clear various settings of the target status register in a Peripheral Component Interconnect Express (PCIE) device and obtain first device information of the PCIE device, where the first device information includes: the first data transfer rate of the PCIE device and the working status of the Downstream Port Disable (DPC);

[0029] A second module, configured to, when detecting that the first data transfer rate is consistent with a preset rate and the working status is started, inject a target type of error into the PCIE device by a target fault injection tool and obtain second device information of the PCIE device, where the second device information includes: the set information of the target status register and the second data transfer rate of the PCIE device;

[0030] A third module, configured to control the operating system to clear the injected error of the target type when it is detected that the target status register is successfully set and the second data transfer rate of the PCIE device is lower than the first data transfer rate, so that the PCIE device resumes normal operation.

[0031] In a third aspect of the implementation of the present invention, an electronic device is further provided, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus; the memory is used to store a computer program; the processor is used to implement any of the above-mentioned fault recovery methods executed by the client when executing the program stored on the memory.

[0032] In a fourth aspect of the implementation of the present invention, a computer-readable storage medium is further provided. Instructions are stored in the computer-readable storage medium, and when it runs on a computer, the computer is made to execute any of the above-mentioned fault recovery methods executed by the client.

[0033] In a fifth aspect of the implementation of the present invention, a computer program product containing instructions is further provided. When it runs on a computer, the computer is made to execute any of the above-mentioned fault recovery methods executed by the client.

[0034] The fault recovery method provided by the embodiment of the present invention clears the settings of each item in the target status register of a Peripheral Component Interconnect Express (PCIE) device and obtains the first device information of the PCIE device. The first device information includes: the first data transfer rate of the PCIE device and the working status of Downstream Port Clocking (DPC). When it is detected that the first data transfer rate is consistent with the preset rate and the working status is started, a target type of error is injected into the PCIE device through a target error injection tool. By comparing the first data transfer rate of the PCIE device with the preset rate, it is possible to obtain whether the status of the PCIE device is normal at this time, avoiding interference with subsequent error injection tests. The second device information of the PCIE device is obtained. The second device information includes: the setting information of the target status register and the second data transfer rate of the PCIE device. When it is detected that the target status register is successfully set and the second data transfer rate of the PCIE device is less than the first data transfer rate, the operating system is controlled to clear the injected target type of error, so that the PCIE device returns to normal. Through the setting information in the target status register, it can be determined that the error injection is successful, and through the change of the second data transfer rate of the PCIE device, it can be known that DCP can detect and intercept control for this target type of error, greatly accelerating the verification of the DCP function in the PCIE device. It can be seen that in the embodiment of the present invention, different types of errors are injected into the PCIE device through a target error injection tool to observe the effect of DPC on different types of errors, avoiding the problem that it is difficult to find PCIE devices of different fault types, resulting in incomplete detection of the system and poor system stability. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art.

[0036] Figure 1 The step flow of the fault recovery method provided by the embodiment of the present invention Figure 1 ;

[0037] Figure 2 The step flow of the fault recovery method provided by the embodiment of the present invention Figure 2 ;

[0038] Figure 3 The step flow of the fault recovery method provided by the embodiment of the present invention Figure 3 ;

[0039] Figure 4 The step flow of the fault recovery method provided by the embodiment of the present invention Figure 4 ;

[0040] Figure 5The step flow of the fault recovery method provided by the embodiment of the present invention Figure 5 ;

[0041] Figure 6 The structural schematic diagram of a fault recovery device provided by the embodiment of the present invention;

[0042] Figure 7 The structural schematic diagram of an electronic device provided by the embodiment of the present invention. Detailed implementation manners

[0043] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the following will elaborate on various embodiments of the present invention with reference to the accompanying drawings. However, those of ordinary skill in the art can understand that in various embodiments of the present invention, many technical details are proposed to help readers better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions claimed in the present application can still be implemented. The division of the following embodiments is for convenience of description and should not constitute any limitation on the specific implementation manners of the present invention. The various embodiments can be combined and cross-referenced with each other on the premise of not being contradictory.

[0044] The embodiment of the present invention is applied to the Basic Input Output System (BIOS). The content of the BIOS is integrated on a ROM chip on the server motherboard. Its main functions are to save the most important basic input and output programs of the server system, system information settings, power-on self-test programs, and system startup bootstrap programs, etc. Its main function is to provide the most basic and direct hardware settings and controls for the computer. The embodiment of the present invention verifies the DPC function in the PCIE device through the BIOS.

[0045] Referring to Figure 1 which shows the step flow of the fault recovery method provided by the embodiment of the present invention Figure 1 The method may include:

[0046] Step 101, clear the settings of each item in the target status register of the Peripheral Component Interconnect Express (PCIE) device and obtain the first device information of the PCIE device. The first device information includes: the first data transfer rate of the PCIE device and the working status of the Downstream Port Disable (DPC).

[0047] In the embodiments of the present invention, it is determined whether the target type of error is successfully injected by comparing the parameter changes of the target status register. Therefore, it is first necessary to clear the settings of each item in the target status register of the Peripheral Component Interconnect Express (PCIE) device to observe subsequent changes. Also, since the embodiments of the present invention are for testing the function of the Data Processing Coordination (DPC), and the role of the DPC technology is to stop the PCIE communication below the downstream port after detecting an uncorrectable error at or below the port to avoid potential propagation of any data damage. It can be known that when the DPC detects an error, it will reduce the data transfer rate of the PCIE. Therefore, the function of the DCP can also be verified by the change in the data transfer rate of the PCIE device. In addition, to verify the function of the DPC, it is first necessary to ensure that the DPC is in the startup working state. Therefore, it is also necessary to obtain the current operating state of the DPC.

[0048] Therefore, in this embodiment, the settings in the target status register of the PCIE device are first cleared, and then the current data transfer rate of the PCIE device, that is, the first data transfer rate, and the current operating state of the DPC are obtained.

[0049] It should be noted that since this method is to simulate whether the system has a repair function, and the wrong injection uses the target wrong injection tool CScripts command, it is necessary to modify the options of the operating system to support this test environment. In addition, this function depends on the kernel of the Linux operating system. Therefore, it is necessary to upgrade the kernel. Thus, before clearing the settings of each item in the target status register of the Peripheral Component Interconnect Express (PCIE) device and obtaining the first device information of the PCIE device, it is also necessary to configure the test environment. The specific methods include:

[0050] Configure the test environment according to the characteristics of the target wrong injection tool and the DPC to enable the target wrong injection tool and the DPC to play their roles.

[0051] For example, by modifying the options of the operating system: OS Native AER Support = Disable; IIO EDPC Support = On Fatal and Non-Fatal Errors, the commands of the target wrong injection tool can be edited and implemented in the operating system. For the kernel, it supports the DPC by upgrading. Generally, versions above 5.4.0 can support the DPC.

[0052] Step 102: When it is detected that the first data transfer rate is consistent with the preset rate and the working state is startup, inject the target type of error into the PCIE device through the target wrong injection tool and obtain the second device information of the PCIE device. The second device information includes: the setting information of the target status register and the second data transfer rate of the PCIE device.

[0053] In an embodiment of the present invention, after obtaining the first device information of the PCIE device, it is possible to confirm whether the status of the PCIE device is normal and whether the DPC is started according to the first device information. Therefore, the first data transfer rate is compared with a preset rate. The preset rate is a fixed value. Generally, the data transfer rates of different versions of PCIE devices are certain. For example, the transfer rate of a PCIE Gen1 device is 2.5 Gb / s, the transfer rate of a PCIE Gen2 device is 5.0 Gb / s, and the transfer rate of a PCIE Gen3 device is 8.0 Gb / s. When it is known that the PCIE is Gen3 according to the device model, the preset rate is the transfer rate of the PCIE Gen3 device. When it is detected that the first data transfer rate is consistent with that of PCIE Gen3, it proves that the PCIE device is normal at this time. For the confirmation of the DPC operating status, it can be achieved by checking the setting status of the DPC status register. For example, when it is detected that the setting is 1, it proves that the DPC is started at this time. Of course, it can also be set that when it is detected that the setting is 0, it proves that the DPC is started. The present invention does not make specific limitations here.

[0054] It should be noted that in order to ensure that the DPC detects and controls the errors injected by the target error injection tool, it is necessary to first clear the uncorrectable error mask of the completion timeout in the DPC status register to avoid interfering with the subsequent results.

[0055] After determining that the status of the PCIE device is normal and the DPC has been started in an embodiment of the present invention, the target error injection tool can be used to inject an error of a target type into the PCIE device, and then the second device information of the PCIE device can be obtained to confirm whether the error is successfully injected at this time and whether the DPC can play a role in this type of error. According to the above content, it can be known that the second device information that needs to be obtained at this time includes: the setting information of the target status register and the second data transfer rate of the PCIE device.

[0056] Step 103, when it is detected that the target status register is successfully set and the second data transfer rate of the PCIE device is less than the first data transfer rate, control the operating system to clear the injected target type of error so that the PCIE device returns to normal.

[0057] In an embodiment of the present invention, the second device information of the PCIE device is obtained to confirm whether the error is successfully injected at this time and whether the DPC can play a role in this type of error. In this embodiment, when the error injection is successful, the target status register is set. Therefore, when it is checked that the parameter setting of the target status register shows a successful set, it can be determined that the error of the target type is successfully injected into the BIOS. In addition, since the DPC will reduce the rate of the PCIE after detecting the error, when it is detected that the current second data transfer rate of the PCIE device is less than the first data transfer rate of the PCIE device under normal circumstances, it means that the DCP can play a role in this target type of error at this time. When it is confirmed that the DCP plays a role in this type of error, the kernel of the system can be released to repair the error and clear the injected target type of error, so that the PCIE device returns to normal.

[0058] It should be noted that in order to make the PCIE device return to normal, while controlling the operating system to clear the injected target type of error, it is also necessary to clear the set information of the status register. And because the error has been repaired, the link has been restored and restored to the original speed at this time, that is, the current data transfer rate will be restored to the first data transfer rate. Specifically, it includes: controlling the operating system to clear the injected target type of error, controlling the current data transfer rate of the PCIE device to be restored to the first data transfer rate, and controlling the set information of the target status register to be cleared, so that the PCIE device returns to normal.

[0059] The fault recovery method provided by the embodiment of the present invention clears the settings of each item in the target status register of a Peripheral Component Interconnect Express (PCIE) device and obtains the first device information of the PCIE device. The first device information includes the first data transfer rate of the PCIE device and the working state of Downstream Port Clamping (DPC). When it is detected that the first data transfer rate is consistent with the preset rate and the working state is started, a target type of error is injected into the PCIE device through a target error injection tool. By comparing the first data transfer rate of the PCIE device with the preset rate, it is possible to obtain whether the state of the PCIE device is normal at this time, avoiding interference with subsequent error injection tests. The second device information of the PCIE device is obtained. The second device information includes the setting information of the target status register and the second data transfer rate of the PCIE device. When it is detected that the target status register is successfully set and the second data transfer rate of the PCIE device is less than the first data transfer rate, the operating system is controlled to clear the injected target type of error, so that the PCIE device returns to normal. The success of error injection can be determined through the setting information in the target status register, and it can also be known through the change of the second data transfer rate of the PCIE device that DCP can detect and intercept control for this target type of error, greatly accelerating the verification of the DCP function in the PCIE device. It can be seen that in the embodiment of the present invention, different types of errors are injected into the PCIE device through a target error injection tool to observe the effect of DPC on different types of errors, avoiding the problem that it is difficult to find PCIE devices of different fault types, resulting in incomplete detection of the system and poor system stability.

[0060] Referring to Figure 2 , a step flow of the fault recovery method provided by the embodiment of the present invention is shown Figure 2 , the method may include:

[0061] Step 201, obtain the log information of the operating system and the script information of the target error injection tool.

[0062] In the embodiment of the present invention, the data transfer rate of the PCIE device is used to determine whether DCP plays a role, and it can also be used to determine whether the PCIE device returns to normal. Therefore, in the embodiment of the present invention, the current data transfer rate is viewed by checking the script information of the target error injection tool, and based on this, it is judged whether the PCIE device returns to normal.

[0063] In addition, the status information of error clearing can also be obtained by checking the message log under the operating system. Also, because the injection success of the error is judged by the change of the setting in the target status register, it is also possible to judge whether the PCIE device returns to normal according to whether the target status register is emptied again.

[0064] Step 202: Determine that the PCIe device has returned to normal according to the log information and the script information.

[0065] In the embodiment of the present invention, the log information and the script information are used to determine that the PCIe device has returned to normal. For example, check whether there is a prompt in the message log under the OS: device recovery successful, display the log of successful error repair, and check whether the speed reduction of the PCIe device has recovered through the Cscripts script of the target error annotation tool. When the corresponding information appears, it is determined that the PCIe device has returned to normal.

[0066] The fault recovery method provided by the embodiment of the present invention clears the settings of each item in the target status register of the Peripheral Component Interconnect Express (PCIe) device and obtains the first device information of the PCIe device. The first device information includes: the first data transfer rate of the PCIe device and the working state of the Downstream Port Clocking (DPC). When it is detected that the first data transfer rate is consistent with the preset rate and the working state is started, a target type of error is injected into the PCIe device through the target error annotation tool. By comparing the first data transfer rate of the PCIe device with the preset rate, it is possible to obtain whether the state of the PCIe device is normal at this time, avoiding interference with subsequent error annotation tests. The second device information of the PCIe device is obtained. The second device information includes: the setting information of the target status register and the second data transfer rate of the PCIe device. When it is detected that the target status register is successfully set and the second data transfer rate of the PCIe device is less than the first data transfer rate, the operating system is controlled to clear the injected target type of error, so that the PCIe device returns to normal. The success of error annotation can be clearly determined through the setting information in the target status register, and it can be known from the change of the second data transfer rate of the PCIe device that the DCP for this target type of error can be detected and intercepted and controlled, greatly accelerating the verification of the DCP function in the PCIe device. It can be seen that in the embodiment of the present invention, different types of errors are injected into the PCIe device through the target error annotation tool to observe the effect of DPC on different types of errors, avoiding the problem that it is difficult to find PCIe devices of different fault types, resulting in incomplete detection of the system and poor system stability.

[0067] Refer to Figure 3 which shows the step flow of the fault recovery method provided by the embodiment of the present invention Figure 3 The method may include:

[0068] Step 301: Control the target status register to stop updating and obtain the third device information of the PCIe device. The third device information includes: the topology structure of the PCIe device.

[0069] In the embodiment of the present invention, before clearing the settings of the target status register in the Peripheral Component Interconnect Express (PCIE) device, it is also necessary to control the target status register to stop updating. Since the CPU is operating at a high speed, the PCIE device is constantly changing, and thus the target status register is also continuously updated. To ensure that the settings of the target status register can be cleared, it is necessary to first pause the change of the register status and then perform the clearing.

[0070] In addition, to clear the target status register, it is necessary to first obtain the position of the target status register on the PCIE device and then execute the clear command. At this time, the third device information of the PCIE device is obtained through the pcie.topology() command. The third device information includes the topology structure of the PCIE device. The link information and port information in the PCIE device can also be obtained through pcie.port_map().

[0071] Step 302: Determine the position information of the target status register according to the topology structure of the PCIE device.

[0072] In the embodiment of the present invention, when clearing the target status register, it is necessary to obtain the position of the target status register, which can be obtained according to the topology structure information. Then, when performing the clear command, the corresponding clear operation can be carried out through the obtained position and the name of the target status register. For example, it is carried out through the instruction sv.socket1.uncore.pi5.pxp0.pcieg5.port2.cfg.devsts, where pi5.pxp0.pcieg5.port2 is the slot information of the PCIE device regarding the target status register, and devsts is the name of the target status register.

[0073] The fault recovery method provided by the embodiment of the present invention clears the settings of each item in the target status register of a Peripheral Component Interconnect Express (PCIE) device and obtains the first device information of the PCIE device. The first device information includes: the first data transfer rate of the PCIE device and the working status of Downstream Port Clocking (DPC). When it is detected that the first data transfer rate is consistent with the preset rate and the working status is started, a target type of error is injected into the PCIE device through a target error injection tool. By comparing the first data transfer rate of the PCIE device with the preset rate, it is possible to obtain whether the status of the PCIE device is normal at this time, avoiding interference with subsequent error injection tests. The second device information of the PCIE device is obtained. The second device information includes: the setting information of the target status register and the second data transfer rate of the PCIE device. When it is detected that the target status register is successfully set and the second data transfer rate of the PCIE device is less than the first data transfer rate, the operating system is controlled to clear the injected target type of error, so that the PCIE device returns to normal. The success of error injection can be determined through the setting information in the target status register, and it can be known from the change of the second data transfer rate of the PCIE device that DCP can detect and intercept control for this target type of error, greatly accelerating the verification of the DCP function in the PCIE device. It can be seen that in the embodiment of the present invention, different types of errors are injected into the PCIE device through a target error injection tool to observe the effect of DPC on different types of errors, avoiding the problem that it is difficult to find PCIE devices of different fault types, resulting in incomplete detection of the system and poor system stability.

[0074] Refer to Figure 4 , which shows the step flow of the fault recovery method provided by the embodiment of the present invention Figure 4 , the method may include:

[0075] Step 401, obtain the cls value of the first register and the flag bit information of the second register in the PCIE device.

[0076] In the embodiment of the present invention, the data transfer rate of the PCIE device is obtained through the CLS value in the first register, and the setting information of DPC is obtained through the flag bit information of the second register.

[0077] Step 402, determine the first data transfer rate of the PCIE device according to the cls value.

[0078] In an embodiment of the present invention, it is determined whether the data transfer rate of the PCIE device is consistent with the preset rate by determining whether the cls value of the first register is the same as the preset value. For example, the cls value in the first register is obtained as --0x00000003 through the command sv.socket1.uncore.pi5.pxp0.pcieg5.port2.cfg.linksts.show, and it is determined whether it is consistent with the actual version of the PCIE device through the cls value. Here, 3 corresponds to a gen3 device, and 4 corresponds to a gen4 device. According to this value and the rate version of the PCIE device, if they are consistent, it is considered that the first data transfer rate is consistent with the actual one.

[0079] Step 403: Determine the status information of the DPC according to the flag bit information.

[0080] In an embodiment of the present invention, the status information of the DPC is determined by obtaining the setting situation in the dpc register. Setting refers to setting a certain bit variable. In the register, the setting situation can be determined by the assignment situation of the flag bit. For example, the flag bit information about the DPC setting information in the second register is obtained through the command sv.socket1.uncore.pi5.pxp0.pcieg5.port2.cfg.dpcctl.show, and the status information of the DPC is determined according to the flag bit information. When the flag bit is assigned 1, it proves that the DPC is set and the DPC has been enabled at this time.

[0081] The fault recovery method provided by the embodiment of the present invention clears the settings of each item in the target status register of a Peripheral Component Interconnect Express (PCIE) device and obtains the first device information of the PCIE device. The first device information includes: the first data transfer rate of the PCIE device and the working state of Downstream Port Clocking (DPC). When it is detected that the first data transfer rate is consistent with the preset rate and the working state is started, a target type of error is injected into the PCIE device through a target error injection tool. By comparing the first data transfer rate of the PCIE device with the preset rate, it is possible to obtain whether the status of the PCIE device is normal at this time, avoiding interference with subsequent error injection tests. The second device information of the PCIE device is obtained. The second device information includes: the setting information of the target status register and the second data transfer rate of the PCIE device. When it is detected that the target status register is successfully set and the second data transfer rate of the PCIE device is less than the first data transfer rate, the operating system is controlled to clear the injected target type of error, so that the PCIE device returns to normal. The success of error injection can be determined through the setting information in the target status register, and it can be known from the change in the second data transfer rate of the PCIE device that DCP can detect and intercept control for this target type of error, greatly accelerating the verification of the DCP function in the PCIE device. Thus, it can be seen that in the embodiment of the present invention, different types of errors are injected into the PCIE device through a target error injection tool to observe the effect of DPC on different types of errors, avoiding the problem that it is difficult to find PCIE devices of different fault types, resulting in incomplete detection of the system and poor system stability.

[0082] Referring to Figure 5 , the step flow of the fault recovery method provided by the embodiment of the present invention is shown Figure 5 , the method may include:

[0083] Step 501, obtain the error names of different types in the target error injection tool.

[0084] In the embodiment of the present invention, the target error injection tool includes many different types of errors, and each type of error has its own corresponding error name. For example, software trigger is the name of a type of error.

[0085] Step 502, inject a target type of error into the PCIE device according to the error name.

[0086] In the embodiment of the present invention, when injecting an error of a target type into a PCIE device, it is necessary to select any one of the error type names from different error names in the target error injection tool, or select the name of a specified error type to inject an error into the PCIE device. For example, the command sv.socket1.uncore.pi5.pxp0.pcieg5.port2.cfg.dpcctl.dpcst=1 is used to inject a command into the PCIE device, where "st" in "dpcst" is the abbreviation of "software trigger".

[0087] The fault recovery method provided by the embodiment of the present invention includes clearing the settings of each item in the target status register of a Peripheral Component Interconnect Express (PCIE) device and obtaining the first device information of the PCIE device. The first device information includes the first data transfer rate of the PCIE device and the working status of the Downstream Port Suppression (DPC). When it is detected that the first data transfer rate is consistent with the preset rate and the working status is started, an error of the target type is injected into the PCIE device through the target error injection tool. By comparing the first data transfer rate of the PCIE device with the preset rate, it is possible to obtain whether the status of the PCIE device is normal at this time, avoiding interference with subsequent error injection tests. The second device information of the PCIE device is obtained. The second device information includes the set information of the target status register and the second data transfer rate of the PCIE device. When it is detected that the target status register is successfully set and the second data transfer rate of the PCIE device is less than the first data transfer rate, the operating system is controlled to clear the injected error of the target type so that the PCIE device returns to normal. The success of error injection can be determined through the set information in the target status register, and it can be known from the change of the second data transfer rate of the PCIE device that the DCP for this target type of error can be detected and intercepted and controlled, greatly accelerating the verification of the DCP function in the PCIE device. It can be seen that in the embodiment of the present invention, different types of errors are injected into the PCIE device through the target error injection tool to observe the effect of DPC on different types of errors, avoiding the problem that it is difficult to find PCIE devices of different fault types, resulting in incomplete detection of the system and poor system stability.

[0088] Refer to Figure 6 , which shows a schematic structural diagram of a fault recovery device provided by an embodiment of the present invention. As Figure 6 shown, the device may include:

[0089] The first module 601 is used to clear the settings of each item in the target status register of a Peripheral Component Interconnect Express (PCIE) device and obtain the first device information of the PCIE device. The first device information includes: the first data transfer rate of the PCIE device and the working status of Downstream Port Clocking (DPC).

[0090] The second module 602 is used to, when it is detected that the first data transfer rate is consistent with a preset rate and the working status is started, inject an error of a target type into the PCIE device through a target error injection tool and obtain the second device information of the PCIE device. The second device information includes: the set information of the target status register and the second data transfer rate of the PCIE device.

[0091] The third module 603 is used to, when it is detected that the target status register is set successfully and the second data transfer rate of the PCIE device is less than the first data transfer rate, control the operating system to clear the injected target type of error so that the PCIE device returns to normal.

[0092] Optionally, the third module further includes:

[0093] The first control sub-module is used to control the operating system to clear the injected target type of error, control the current data transfer rate of the PCIE device to return to the first data transfer rate, and control the clearing of the set information of the target status register so that the PCIE device returns to normal.

[0094] Optionally, the fault recovery device further includes:

[0095] An acquisition module is used to acquire the log information of the operating system and the script information of the target error injection tool.

[0096] The first determination module is used to determine that the PCIE device returns to normal according to the log information and the script information.

[0097] A configuration module is used to configure a test environment according to the characteristics of the target error injection tool and DPC so that the target error injection tool and DPC can function properly.

[0098] A control module is used to control the target status register to stop updating and obtain the third device information of the PCIE device. The third device information includes: the topology structure of the PCIE device.

[0099] The second determination module is used to determine the location information of the target status register according to the topology structure of the PCIE device.

[0100] Optionally, the first module further includes:

[0101] The first acquisition sub-module is used to acquire the cls value of the first register and the flag bit information of the second register in the PCIE device.

[0102] The first determination sub-module is configured to determine the first data transfer rate of the PCIE device according to the cls value.

[0103] The second determination sub-module is configured to determine the status information of the DPC according to the flag bit information.

[0104] Optionally, the second module further includes:

[0105] The second acquisition sub-module is configured to acquire different types of error names in the target error injection tool.

[0106] The injection sub-module is configured to inject an error of a target type into the PCIE device according to the error name.

[0107] The fault recovery method provided by the embodiment of the present invention clears the settings of each item in the target status register of the PCIE device and acquires the first device information of the PCIE device. The first device information includes: the first data transfer rate of the PCIE device and the status information of the DPC. By acquiring the status information of the DPC, it is determined whether the DPC is successfully enabled. When it is detected that the first data transfer rate is consistent with the preset rate and the DPC is in the startup state, an error of a target type is injected into the PCIE device through the target error injection tool, and the second device information of the PCIE device is acquired. The second device information includes: the setting information of the target status register and the second data transfer rate of the PCIE device. By acquiring the parameter information of the target status register and the change situation of the data transfer rate of the PCIE device, it can be determined whether the target type of error is successfully injected into the PCIE device at this time. When it is detected that the target status register is successfully set and the second data transfer rate of the PCIE device is less than the first data transfer rate, the operating system is controlled to clear the injected target type of error, so that the PCIE device returns to normal. By the operating system clearing the injected target type of error and making the PCIE device return to normal, the stability of the system is ensured. It can be seen that in the embodiment of the present invention, different types of errors are injected into the PCIE device through the target error injection tool, so as to observe the effect of the DPC on different types of errors, avoiding the problem that it is difficult to find PCIE devices of different fault types, resulting in incomplete detection of the system and poor system stability.

[0108] The embodiment of the present invention also provides an electronic device, as Figure 7 shown, including a processor 701, a communication interface 702, a memory 703, and a communication bus 704. Among them, the processor 701, the communication interface 702, and the memory 703 complete mutual communication through the communication bus 704;

[0109] The memory 703 is used to store a computer program;

[0110] When the processor 701 executes the program stored in the memory 703, the following steps are implemented:

[0111] Clear the settings of each item in the target status register of the Peripheral Component Interconnect Express (PCIE) device and obtain the first device information of the PCIE device. The first device information includes: the first data transfer rate of the PCIE device and the working status of the Downstream Port Clocking (DPC);

[0112] When it is detected that the first data transfer rate is consistent with the preset rate and the working status is started, inject an error of the target type into the PCIE device through the target error injection tool and obtain the second device information of the PCIE device. The second device information includes: the setting information of the target status register and the second data transfer rate of the PCIE device;

[0113] When it is detected that the target status register is successfully set and the second data transfer rate of the PCIE device is less than the first data transfer rate, control the operating system to clear the injected error of the target type so that the PCIE device returns to normal.

[0114] The communication bus mentioned in the above terminal may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is used in the figure to represent it, but it does not mean that there is only one bus or one type of bus.

[0115] The communication interface is used for communication between the above terminal and other devices.

[0116] The memory may include a Random Access Memory (RAM), or may also include a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.

[0117] The above-mentioned processor may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0118] In another embodiment provided by the present invention, there is also provided a computer-readable storage medium storing instructions that, when running on a computer, cause the computer to execute any one of the fault recovery methods executed by the client in the above embodiments.

[0119] In another embodiment provided by the present invention, there is also provided a computer program product containing instructions that, when running on a computer, cause the computer to execute any one of the fault recovery methods executed by the client in the above embodiments.

[0120] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or a data center integrating one or more available media. The available medium may be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).

[0121] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the said element.

[0122] Each embodiment in this specification is described in a related manner. For the same and similar parts between the embodiments, reference can be made to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and reference can be made to the relevant part of the method embodiment for the relevant content.

[0123] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are included within the protection scope of the present invention.

Claims

1. A fault recovery method, characterized in that, The method includes: Clearing the settings of each item in the target status register of a Peripheral Component Interconnect Express (PCIE) device and obtaining first device information of the PCIE device, where the first device information includes: the first data transfer rate of the PCIE device and the working status of Downstream Port Clocking (DPC); When it is detected that the first data transfer rate is consistent with a preset rate and the working status is started, injecting an error of a target type into the PCIE device through a target error injection tool and obtaining second device information of the PCIE device, where the second device information includes: the setting information of the target status register and the second data transfer rate of the PCIE device; When it is detected that the target status register is successfully set and the second data transfer rate of the PCIE device is less than the first data transfer rate, controlling the operating system to clear the injected target type of error, controlling the current data transfer rate of the PCIE device to resume to the first data transfer rate, and controlling to clear the setting information of the target status register, so that the PCIE device resumes normal operation.

2. The method according to claim 1, wherein After controlling the operating system to clear the injected target type of error to make the PCIE device resume normal operation when it is detected that the target status register is successfully set and the second data transfer rate of the PCIE device is less than the first data transfer rate, the method further includes: Obtaining the log information of the operating system and the script information of the target error injection tool; Determining that the PCIE device resumes normal operation according to the log information and the script information.

3. The method according to claim 1, wherein Before clearing the settings of each item in the target status register of the PCIE device and obtaining the first device information of the PCIE device, the method further includes: Configuring a test environment according to the characteristics of the target error injection tool and the DPC, so that the target error injection tool and the DPC can function properly.

4. The method according to claim 3, wherein After configuring the test environment according to the characteristics of the target error injection tool and the DPC so that the target error injection tool and the DPC can function properly, the method further includes: Controlling the target status register to stop updating and obtaining third device information of the PCIE device, where the third device information includes: the topology structure of the PCIE device; Determining the location information of the target status register according to the topology structure of the PCIE device.

5. The method according to claim 1, characterized in that, The obtaining of the first device information of the PCIE device includes: Obtaining the cls value of the first register and the flag bit information of the second register in the PCIE device; Determining the first data transfer rate of the PCIE device according to the cls value; Determining the status information of the DPC according to the flag bit information.

6. The method according to claim 1, characterized in that, The injecting of an error of a target type into the PCIE device through the target error injection tool includes: Obtaining different types of error names in the target error injection tool; Injecting an error of a target type into the PCIE device according to the error name.

7. A fault recovery device, characterized in that, The device includes: The first module is used to clear the settings of each item in the target status register of a Peripheral Component Interconnect Express (PCIE) device in a high-speed serial computer expansion bus standard and obtain the first device information of the PCIE device. The first device information includes: the first data transfer rate of the PCIE device and the working status of Downstream Port Clocking (DPC); The second module is used to, when it is detected that the first data transfer rate is consistent with a preset rate and the working status is started, inject an error of a target type into the PCIE device through a target error injection tool and obtain the second device information of the PCIE device. The second device information includes: the set information of the target status register and the second data transfer rate of the PCIE device; The third module is used to, when it is detected that the target status register is successfully set and the second data transfer rate of the PCIE device is less than the first data transfer rate, control the operating system to clear the injected target type of error, control the current data transfer rate of the PCIE device to resume to the first data transfer rate, and control the clearing of the set information of the target status register, so that the PCIE device returns to normal.

8. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus; The memory is used to store a computer program; The processor is used to implement the method according to any one of claims 1-6 when executing the program stored on the memory.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Error injection method and device for expansion bus of high-speed serial computer

    CN113032199A

  • Server RAS function test method and device, equipment and medium

    CN113407394A