Fault handling method and apparatus for PCI device, and fault handling system
By detecting PCI devices on the PCI bridge link, fault handling methods are determined, including activating the fault handling function of the EP device and setting the fault threshold of the PCI switch device, and adjusting the transmission rate. This solves the problem of inaccurate PCI device fault handling and improves the stability and reliability of the server.
Patent Information
- Application Number
- PCT/CN2025/082995
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-28
- Filing Date
- 2025-03-17
- Publication Date
- 2025-12-04
AI Technical Summary
Existing technologies cannot accurately determine the location and type of PCI device failures, resulting in an inability to effectively handle PCI link failures and affecting the stability and reliability of servers.
By detecting the PCI devices on the PCI bridge link, the fault handling method is determined, including activating the fault handling function of the EP device and setting the fault threshold of the PCI switch device, and adjusting the transmission rate of the PCI bridge link to accurately handle the fault.
It improves the accuracy of PCI device fault handling, ensures server stability and reliability, and reduces the risk of downtime due to faults.
Smart Images

Figure CN2025082995_04122025_PF_FP_ABST
Abstract
Description
PCI device troubleshooting methods and devices, troubleshooting systems
[0001] Cross-references to related applications
[0002] This application claims priority to Chinese Patent Application No. 202410671171.0, filed on May 28, 2024, entitled “Method and Apparatus for Handling Faults of PCI Devices, Fault Handling System”, the entire contents of which are incorporated herein by reference. Technical Field
[0003] This application relates to the server field, and in particular, to a fault handling method, apparatus, and system for Peripheral Component Interconnect (PCI) devices. Background Technology
[0004] All servers, regardless of architecture, support and must support fundamental core modules such as memory and Peripheral Component Interconnect (PCI) devices. The core of using a server is connecting external PCI devices via the PCI bus. Traditional PCI devices include network interface cards (NICs), storage media cards, accelerator cards, and graphics processing units (GPUs). Regardless of the component, it must adhere to the standards of its supported protocols, and PCI devices must also comply with the PCI bus protocol standard. However, due to the high speed of the PCI bus (currently supporting a maximum speed of 32 GT / s, or gigabytes per second), a failure in the PCI bus protocol can lead to serious and fatal problems such as PCI device loss and server system crashes. PCI link failures are categorized into multiple error types, and various devices can be supported on a PCI link. Current technologies cannot accurately pinpoint which device on the PCI link is experiencing a fault, thus hindering accurate fault handling. Summary of the Invention
[0005] This application provides a method, apparatus, and system for handling PCI device faults, which at least solves the problem in related technologies that it is impossible to accurately identify the faulty PCI device, thus making it impossible to determine the accurate fault handling method.
[0006] According to one embodiment of this application, a fault handling method for peripheral component interconnect (PCI) devices is provided, applied to a fault handling system. The fault handling system is installed on the motherboard of a server and connected to a PCI bridge link deployed in the server. The method includes: detecting PCI devices connected to the PCI bridge link, wherein the PCI devices include terminal device (EP) devices and / or peripheral component interconnect (PCI) switch devices, and multiple EP devices can be connected to the PCI switch device; determining a fault handling method for the PCI devices based on the detected PCI devices, and handling the faults of the PCI devices according to the fault handling method; wherein, when the PCI devices connected to the PCI bridge link are only EP devices, activating the fault handling function of the EP devices, the fault handling function being a function set in the EP devices to handle the faults of the EP devices; when the PCI devices connected to the PCI bridge link include PCI switch devices, setting a fault threshold for the PCI switch devices, and adjusting the transmission rate of the PCI bridge link based on the fault threshold, the fault threshold being used to represent a threshold for recording the number of faults of the PCI switch devices, and the fault handling method including the fault handling function and the method of setting the fault threshold for the PCI switch devices.
[0007] In an exemplary embodiment, when the PCI device connected to the PCI bridge link is only an EP device, activating the fault handling function of the EP device includes: setting the Advanced Error Reporting (AER) function of the EP device to an enabled or enabled state to activate the fault handling function of the EP device, wherein the AER function includes the fault handling function.
[0008] In an exemplary embodiment, when the PCI device connected on the PCI bridge link is only an EP device, after activating the fault handling function of the EP device, the method further includes: monitoring a first fault event of the EP device, wherein the first fault event includes at least one of the following: a first fault event that can be repaired, a first fault event that cannot be repaired, a first fault event that causes the EP device to be unable to operate, and a first fault event that allows the EP device to continue operating; and, upon detecting the first fault event, reporting the first fault event to the target operating system and the baseboard management controller in the server.
[0009] The target operating system is configured to perform one of the following operations based on the first fault event: restart the server, perform a self-test, or generate a fault message; the baseboard management controller is configured to perform at least one of the following operations based on the first fault event: record the first fault event, perform a fault troubleshooting operation, or repair the first fault event.
[0010] In an exemplary embodiment, when the PCI device connected to the PCI bridge link includes a PCI switch device, setting a fault threshold for the PCI switch device includes: determining a target register set in the PCI switch device, wherein the target register is set to record the number of failures of the PCI switch device; and setting a fault threshold for the target register according to preset conditions.
[0011] In one exemplary embodiment, the PCI devices connected to the PCI bridge link include PCI switch devices, including one of the following: the PCI devices connected to the PCI bridge link are only PCI switch devices; the PCI devices connected to the PCI bridge link are only PCI switch devices, and one or more EP devices are connected to the PCI switch device; the PCI devices connected to the PCI bridge link include PCI switch devices and EP devices, and the PCI switch device is connected to a first processor in the server, and the EP device is connected to a second processor in the server; the PCI devices connected to the PCI bridge link include PCI switch devices and EP devices, and the PCI switch device is connected to a first processor in the server, and one or more other EP devices are connected to the PCI switch device, and the EP devices are connected to a second processor in the server.
[0012] In one exemplary embodiment, setting a fault threshold for a target register according to preset conditions includes: setting the fault threshold for the target register according to preset conditions and by at least one of the following operations: determining the operating information of a PCI switch device and setting a fault threshold based on the operating information, wherein the operating information includes at least one of the following: operating environment, operating performance, operating security performance, and operating stability performance; determining a first number of PCI switch devices and a second number of EP devices connected in the PCI switch devices, and setting a fault threshold based on the first number and the second number.
[0013] In an exemplary embodiment, after setting the fault threshold of the target register according to preset conditions, the method further includes: detecting the PCI Switch device and one or more EP devices connected to the PCI Switch device; when the PCI Switch device fails or any EP device connected to the PCI Switch device fails, triggering a counter in the target register to perform a fault accumulation operation, wherein the fault accumulation operation is used to accumulate the number of times the PCI Switch device has failed.
[0014] In an exemplary embodiment, when the PCI device connected on the PCI bridge link includes a PCI switch device, before setting the fault threshold of the PCI switch device, the method further includes one of the following: when one or more EP devices are connected to the PCI switch device, enabling the fault handling function of the EP device and receiving the result of the EP device handling the fault according to the fault handling function through the PCI switch device; when the PCI switch device is connected to a first processor in the server and the EP device is connected to a second processor in the server, disabling the fault handling function of the EP device; when the PCI switch device is connected to a first processor in the server, one or more other EP devices are connected to the PCI switch device, and the EP device is connected to a second processor in the server, disabling the fault handling function of the EP device, enabling other fault handling functions of other EP devices, and receiving the result of the other EP devices handling the fault according to other fault handling functions through the PCI switch device.
[0015] In one exemplary embodiment, adjusting the transmission rate of a PCI bridge link based on a fault threshold includes: triggering an adjustment of the transmission rate of the PCI bridge link when the accumulated number of first faults of the PCI bridge link is greater than or equal to the fault threshold, thereby obtaining a first adjusted transmission rate.
[0016] In one exemplary embodiment, if the cumulative number of failures of a PCI bridge link is greater than or equal to a failure threshold, the transmission rate of the PCI bridge link is adjusted to obtain a first adjusted transmission rate, including at least one of the following: if the number of failures is greater than or equal to the failure threshold, the transmission rate of the PCI switch device is adjusted to obtain a first adjusted transmission rate; if the number of failures is greater than or equal to the failure threshold, the transmission rate of one or more EP devices connected in the PCI switch device is adjusted to obtain a first adjusted transmission rate.
[0017] In an exemplary embodiment, if the accumulated number of failures of the PCI bridge link is greater than or equal to a failure threshold, the method triggers an adjustment of the transmission rate of the PCI bridge link to obtain a first adjusted transmission rate. The method further includes: clearing the first failure count; and re-recording the number of failures of the PCI bridge link to obtain a second failure count.
[0018] In an exemplary embodiment, after re-recording the number of PCI bridge link failures to obtain a second failure count, the method further includes: if the second failure count is greater than or equal to a failure threshold, continuing to trigger adjustment of the PCI bridge link's transmission rate to obtain a second adjusted transmission rate; and controlling the operation of the PCI bridge link based on the second adjusted transmission rate.
[0019] In an exemplary embodiment, after controlling the operation of the PCI bridge link based on the second adjusted transmission rate, the method further includes: stopping the adjustment of the transmission rate of the PCI bridge link when the second adjusted transmission rate is equal to the minimum rate of the PCI bridge link; and triggering the adjustment of the network bandwidth of the PCI bridge link to obtain the adjusted network bandwidth when the second adjusted transmission rate is equal to the minimum rate of the PCI bridge link and the number of third faults recorded for the PCI bridge link is greater than or equal to a fault threshold, wherein the number of third faults is the number of faults recorded during the operation of the PCI bridge link according to the second adjusted transmission rate.
[0020] In one exemplary embodiment, when the second adjusted transmission rate is equal to the minimum rate of the PCI bridge link and the recorded third number of PCI bridge link failures is greater than or equal to a failure threshold, the network bandwidth of the PCI bridge link is adjusted. After obtaining the adjusted network bandwidth, the method further includes: generating a second failure event when the adjusted network bandwidth is equal to the minimum network bandwidth of the PCI bridge link, wherein the second failure event includes at least one of the following: a second failure event that can be repaired, a second failure event that cannot be repaired, a second failure event that causes the PCI bridge link to be unable to operate, and a second failure event that allows the PCI bridge link to continue operating; sending the second failure event to a baseboard management controller in a server, wherein the baseboard management controller is configured to record the second failure event and generate failure handling information, wherein the failure handling information is used to instruct the replacement of the faulty PCI switch device, and / or, the replacement of the EP device connected to the faulty PCI switch device.
[0021] In an exemplary embodiment, after adjusting the transmission rate of the PCI bridge link based on a fault threshold, the method further includes: generating target information and sending the target information to a baseboard management controller deployed in a server, wherein the target information includes a target transmission rate, the target transmission rate being the rate after adjusting the transmission rate of the PCI bridge link, and the target transmission rate being less than the transmission rate of the PCI bridge link before adjustment; recording the target transmission rate through the baseboard management controller and generating a prompt message to indicate that the first number of faults in the PCI bridge link exceeds the fault threshold.
[0022] In one exemplary embodiment, detecting PCI devices connected on a PCI bridge link includes: starting a fault handling system when the server is started, detecting N levels on the PCI bridge link, where N is a natural number greater than or equal to 1; and determining the type of PCI device in each level.
[0023] According to another embodiment of this application, a PCI device fault handling apparatus is provided, applied to a fault handling system. The fault handling system is installed on the motherboard of a server and connected to a PCI bridge link deployed in the server. The apparatus includes: a first detection module, configured to detect PCI devices connected to the PCI bridge link, wherein the PCI devices include EP devices and / or PCI switch devices, and multiple EP devices can be connected to the PCI switch device; a first processing module, configured to determine a fault handling method for the PCI devices based on the detected PCI devices, and handle the faults of the PCI devices according to the fault handling method; wherein the first processing module includes: a first processing unit, configured to activate the fault handling function of the EP device when the PCI device connected to the PCI bridge link is only an EP device, the fault handling function being a function set in the EP device for handling the faults of the EP device; a second processing unit, configured to set a fault threshold for the PCI switch device when the PCI device connected to the PCI bridge link includes a PCI switch device, and adjust the transmission rate of the PCI bridge link based on the fault threshold, the fault threshold representing a threshold for recording the number of faults of the PCI switch device, and the fault handling method including the fault handling function and setting the PCI... Methods for setting fault thresholds for switch devices.
[0024] In one exemplary embodiment, the first processing unit includes: a first setting subunit, configured to enable or start the AER function of the EP device when the PCI device connected on the PCI bridge link is only an EP device, so as to start the fault handling function of the EP device, wherein the AER function includes the fault handling function.
[0025] In one exemplary embodiment, the apparatus further includes: a first monitoring module, configured to monitor a first fault event of the EP device after activating the fault handling function of the EP device when the PCI device connected on the PCI bridge link is only an EP device, wherein the first fault event includes at least one of the following: a first fault event that can be repaired, a first fault event that cannot be repaired, a first fault event that causes the EP device to be unable to operate, and a first fault event that allows the EP device to continue operating; and a first reporting module, configured to report the first fault event to the target operating system and the baseboard management controller in the server when the first fault event is detected; wherein the target operating system is configured to perform one of the following operations based on the first fault event: restart the server, perform a self-test operation, and generate a fault prompt message; and the baseboard management controller is configured to perform at least one of the following operations based on the first fault event: record the first fault event, perform a fault troubleshooting operation, and repair the first fault event.
[0026] In one exemplary embodiment, the second processing unit includes: a first determining subunit configured to determine a target register set in a PCI switch device when the PCI device connected on the PCI bridge link includes a PCI switch device, wherein the target register is configured to record the number of failures of the PCI switch device; and a first setting subunit configured to set a failure threshold of the target register according to preset conditions.
[0027] In one exemplary embodiment, the PCI devices connected to the PCI bridge link include PCI switch devices, including one of the following: the PCI devices connected to the PCI bridge link are only PCI switch devices; the PCI devices connected to the PCI bridge link are only PCI switch devices, and one or more EP devices are connected to the PCI switch device; the PCI devices connected to the PCI bridge link include PCI switch devices and EP devices, and the PCI switch device is connected to a first processor in the server, and the EP device is connected to a second processor in the server; the PCI devices connected to the PCI bridge link include PCI switch devices and EP devices, and the PCI switch device is connected to a first processor in the server, and one or more other EP devices are connected to the PCI switch device, and the EP devices are connected to a second processor in the server.
[0028] In an exemplary embodiment, the first setting subunit includes: a first setting submodule, configured to set a fault threshold of a target register according to preset conditions and by at least one of the following operations: determining operating information of a PCI Switch device and setting a fault threshold based on the operating information, wherein the operating information includes at least one of the following: operating environment, operating performance, operating security performance, and operating stability performance; determining a first number of PCI Switch devices and a second number of EP devices connected in the PCI Switch devices, and setting a fault threshold based on the first number and the second number.
[0029] In one exemplary embodiment, the apparatus further includes: a first detection module configured to detect a PCI switch device and one or more EP devices connected to the PCI switch device after setting a fault threshold in the target register according to preset conditions; and a first trigger module configured to trigger a counter in the target register to perform a fault accumulation operation when the PCI switch device or any EP device connected to the PCI switch device fails, wherein the fault accumulation operation is used to accumulate the number of times the PCI switch device has failed.
[0030] In one exemplary embodiment, the above-described apparatus further includes one of the following: a first enabling module, configured to enable the fault handling function of the EP device when, in the case that the PCI device connected on the PCI bridge link includes a PCI Switch device, before setting a fault threshold for the PCI Switch device, and in the case that one or more EP devices are connected to the PCI Switch device, and receive the result of the EP device processing the fault according to the fault handling function through the PCI Switch device; a first disabling module, configured to disable the fault handling function of the EP device when, in the case that the PCI device connected on the PCI bridge link includes a PCI Switch device, before setting a fault threshold for the PCI Switch device, and in the case that the PCI Switch device is connected to a first processor in the server, and the EP device is connected to a second processor in the server; and a second processing module, configured to disable the fault handling function of the EP device, enable other fault handling functions of other EP devices, and receive the result of the other EP devices processing the fault according to other fault handling functions through the PCI Switch device when, in the case that the PCI device connected on the PCI bridge link includes a PCI Switch device, before setting a fault threshold for the PCI Switch device, and in the case that the PCI Switch device is connected to a first processor in the server, one or more other EP devices are connected to the PCI Switch device, and the EP devices are connected to a second processor in the server.
[0031] In one exemplary embodiment, the second processing unit includes: a first triggering subunit, configured to trigger an adjustment of the transmission rate of the PCI bridge link to obtain a first adjusted transmission rate when the number of accumulated failures of the PCI bridge link is greater than or equal to a failure threshold.
[0032] In one exemplary embodiment, the first trigger subunit includes at least one of the following: a first trigger submodule configured to trigger adjustment of the transmission rate of the PCI Switch device to obtain a first adjusted transmission rate when the number of first faults is greater than or equal to a fault threshold; and a second trigger submodule configured to trigger adjustment of the transmission rate of one or more EP devices connected in the PCI Switch device to obtain the first adjusted transmission rate when the number of first faults is greater than or equal to a fault threshold.
[0033] In an exemplary embodiment, the apparatus further includes: a first clearing module, configured to trigger an adjustment of the transmission rate of the PCI bridge link when the accumulated number of failures of the PCI bridge link is greater than or equal to a failure threshold, and after obtaining a first adjusted transmission rate, clear the first failure count; and a first recording module, configured to re-record the number of failures of the PCI bridge link to obtain a second failure count.
[0034] In one exemplary embodiment, the apparatus further includes: a first adjustment module, configured to re-record the number of times the PCI bridge link fails, obtain a second number of failures, and then, if the second number of failures is greater than or equal to a failure threshold, continue to trigger the adjustment of the transmission rate of the PCI bridge link to obtain a second adjusted transmission rate; and a first control module, configured to control the operation of the PCI bridge link based on the second adjusted transmission rate.
[0035] In one exemplary embodiment, the apparatus further includes: a first stop module configured to stop adjusting the transmission rate of the PCI bridge link after controlling the operation of the PCI bridge link based on the second adjusted transmission rate, provided that the second adjusted transmission rate is equal to the minimum rate of the PCI bridge link; and a second trigger module configured to trigger adjusting the network bandwidth of the PCI bridge link to obtain the adjusted network bandwidth when the second adjusted transmission rate is equal to the minimum rate of the PCI bridge link and the number of third faults recorded for the PCI bridge link is greater than or equal to a fault threshold, wherein the number of third faults is the number of faults recorded during the operation of the PCI bridge link at the second adjusted transmission rate.
[0036] In one exemplary embodiment, the apparatus further includes: a first generation module configured to, when the second adjusted transmission rate is equal to the minimum rate of the PCI bridge link and the recorded third number of failures of the PCI bridge link is greater than or equal to a failure threshold, trigger adjustment of the network bandwidth of the PCI bridge link; after obtaining the adjusted network bandwidth, when the adjusted network bandwidth is equal to the minimum network bandwidth of the PCI bridge link, generate a second failure event, wherein the second failure event includes at least one of the following: a second failure event that can be repaired, a second failure event that cannot be repaired, a second failure event that causes the PCI bridge link to be unable to operate, and a second failure event that allows the PCI bridge link to continue operating; and a first transmission module configured to send the second failure event to a baseboard management controller in a server, wherein the baseboard management controller is configured to record the second failure event and generate fault handling information, wherein the fault handling information is used to instruct the replacement of the faulty PCI switch device, and / or, the replacement of the EP device connected to the faulty PCI switch device.
[0037] In one exemplary embodiment, the apparatus further includes: a second generation module configured to generate target information after adjusting the transmission rate of the PCI bridge link based on a fault threshold, and send the target information to a baseboard management controller deployed in a server, wherein the target information includes a target transmission rate, the target transmission rate being the rate after adjusting the transmission rate of the PCI bridge link, and the target transmission rate being less than the transmission rate of the PCI bridge link before adjustment; and a second recording module configured to record the target transmission rate through the baseboard management controller and generate a prompt message to indicate that the first number of faults in the PCI bridge link exceeds a fault threshold.
[0038] In an exemplary embodiment, the first detection module includes: a first detection unit configured to start a fault handling system when the server is started, and to detect N levels on the PCI bridge link, wherein N is a natural number greater than or equal to 1; and a first determination unit configured to determine the type of PCI device in each level.
[0039] According to another embodiment of this application, a fault handling system is also provided. The fault handling system is disposed on the motherboard of a server. The fault handling system includes a target processor, which is connected to a PCI bridge deployed in the server via a link. The target processor is configured to implement the steps of the above-described method.
[0040] According to yet another embodiment of this application, a server is also provided, including: a PCI bridge link, and a fault handling system as described above, wherein a PCI device is connected to the PCI bridge link.
[0041] According to yet another embodiment of this application, a computer program product is also provided, including a computer program that, when executed by a processor, implements the steps in any of the above method embodiments.
[0042] According to yet another embodiment of this application, a non-volatile computer-readable storage medium is also provided, wherein a computer program is stored in the non-volatile computer-readable storage medium, and the computer program is configured to perform the steps in any of the above method embodiments when it is run.
[0043] According to yet another embodiment of this application, an electronic device is also provided, including a memory and a processor, wherein a computer program is stored in the memory and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.
[0044] This application addresses the problem in related technologies where the PCI devices connected to the PCI bridge link are first detected. Then, different fault handling methods are determined based on the detected PCI devices, allowing for the resolution of PCI device faults according to these methods. Specifically, when the PCI devices connected to the PCI bridge link are only EP devices, the EP device's fault handling function is activated. When the PCI devices connected to the PCI bridge link include PCI switch devices, a fault threshold for the PCI switch devices is set, and the transmission rate of the PCI bridge link is adjusted based on this threshold. Therefore, this approach solves the problem of inaccurately identifying faulty PCI devices in related technologies, which leads to the inability to determine the accurate fault handling method, thereby improving the accuracy of PCI device fault handling. Attached Figure Description
[0045] Figure 1 is a hardware structure block diagram of a mobile terminal for a PCI device fault handling method according to an embodiment of this application.
[0046] Figure 2 is a flowchart of a fault handling method for a PCI device according to an embodiment of this application;
[0047] Figure 3 is a schematic diagram of the connection of a PCI device according to an embodiment of this application;
[0048] Figure 4 is a second schematic diagram of the connection of a PCI device according to an embodiment of this application;
[0049] Figure 5 is a schematic diagram of the connection of a PCI device according to an embodiment of this application;
[0050] Figure 6 is an interaction diagram of various devices performing error information processing according to an optional embodiment of this application;
[0051] Figure 7 is a flowchart of BIOS error information processing according to an optional embodiment of this application;
[0052] Figure 8 is a structural block diagram of the fault handling system according to an embodiment of this application;
[0053] Figure 9 is a structural block diagram of the server according to an embodiment of this application;
[0054] Figure 10 is a structural block diagram of a fault handling apparatus for a PCI device according to an embodiment of this application;
[0055] Figure 11 is a schematic diagram of a non-volatile computer-readable storage medium for a fault handling method of a PCI device according to an embodiment of this application;
[0056] Figure 12 is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the fault handling method for a PCI device according to an embodiment of this application. Detailed Implementation
[0057] The embodiments of this application will be described in detail below with reference to the accompanying drawings and examples.
[0058] It should be noted that the terms "first," "second," etc., in the specification, claims, and drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0059] The methods and embodiments provided in this application can be executed in a mobile terminal, computer terminal, or similar computing device. Taking a mobile terminal as an example, FIG1 is a hardware structure block diagram of a mobile terminal for a PCI device fault handling method according to an embodiment of this application. As shown in FIG1, the mobile terminal may include one or more (only one is shown in FIG1) processors 102 (processors 102 may include, but are not limited to, microprocessors MCU (Microcontroller Unit) or programmable logic devices FPGA (Field-Programmable Gate Array), etc.) and a memory 104 configured to store data. The mobile terminal may also include a transmission device 106 configured for communication functions and an input / output device 108. Those skilled in the art will understand that the structure shown in FIG1 is only illustrative and does not limit the structure of the mobile terminal. For example, the mobile terminal may also include more or fewer components than shown in FIG1, or have a different configuration than shown in FIG1.
[0060] The memory 104 may be configured to store computer programs, such as application software programs and modules, like the computer program corresponding to the PCI device fault handling method in this embodiment. The processor 102 executes various functional applications and data processing by running the computer programs stored in the memory 104, thereby implementing the aforementioned method. The memory 104 may include high-speed random access memory and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the mobile terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0061] The transmission device 106 is configured to receive or transmit data via a network. Optional examples of the network may include a wireless network provided by the mobile terminal's communication provider. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module, configured to communicate wirelessly with the Internet.
[0062] This embodiment provides a PCI device fault handling method, applied to a fault handling system. The fault handling system is installed on the motherboard of a server and is connected to a PCI bridge deployed in the server via a link. Figure 2 is a flowchart of the PCI device fault handling method according to an embodiment of this application. As shown in Figure 2, the process includes the following steps:
[0063] Step S202: Detect the PCI devices connected on the PCI bridge link, wherein the PCI devices include EP devices and / or PCI switch devices, and multiple EP devices are allowed to be connected on the PCI switch device;
[0064] Step S204: Based on the detected PCI device, determine the fault handling method for the PCI device, and handle the fault of the PCI device according to the fault handling method.
[0065] In cases where the PCI device connected to the PCI bridge link is only an EP (Endpoint) device, the fault handling function of the EP device is activated. The fault handling function is a function set in the EP device to handle faults of the EP device. In cases where the PCI device connected to the PCI bridge link includes a PCI Switch device, a fault threshold for the PCI Switch device is set, and the transmission rate of the PCI bridge link is adjusted based on the fault threshold. The fault threshold is used to represent the threshold for recording the number of faults of the PCI Switch device. The fault handling methods include the fault handling function and the method of setting the fault threshold of the PCI Switch device.
[0066] Through the above steps, the PCI devices connected to the PCI bridge link are first detected. Then, different fault handling methods are determined based on the detected PCI devices, and faults in the PCI devices are handled according to these methods. Specifically, when the PCI devices connected to the PCI bridge link are only EP devices, the fault handling function of the EP devices is activated. When the PCI devices connected to the PCI bridge link include PCI switch devices, a fault threshold for the PCI switch devices is set, and the transmission rate of the PCI bridge link is adjusted based on the fault threshold. Therefore, this solves the problem in related technologies where it is difficult to accurately identify the faulty PCI device, thus making it impossible to determine the accurate fault handling method, thereby improving the accuracy of PCI device fault handling.
[0067] In some embodiments, the fault handling system may be a Basic Input Output System (BIOS) deployed in a server, or it may be a device including a processor. The BIOS is firmware located on the computer's motherboard. It is responsible for hardware initialization, self-test, and booting the operating system when the computer starts. The BIOS provides an interface that allows the operating system and applications to communicate with the computer's hardware.
[0068] In some embodiments, a PCI bridge link refers to a communication link connecting two or more PCI bridges. A PCI bridge is a device configured to connect different PCI links, enabling the transfer of data from one PCI link to another. The PCI bridge link is responsible for transmitting data and control information between different PCI links to facilitate communication and data exchange between different PCI devices.
[0069] In some embodiments, a PCI device refers to a device with a PCI interface. A PCI interface is an external expansion bus, which is a bus designed to connect external devices to the computer. A PCI device can be a network card, storage medium card, accelerator card, GPU card, or other components with a PCI interface.
[0070] In some embodiments, a PCI switch device is a device configured to connect a computer system to external devices. A PCI switch device can expand the PCI bus interface of a computer system, enabling it to connect more external devices, such as network cards, graphics cards, and storage devices. By connecting multiple PCI bus interfaces together, it achieves the management and control of multiple external devices, thereby improving the system's scalability and flexibility. Therefore, a PCI switch device implements time-division multiplexing of the PCI bus, allowing multiple PCI devices to communicate simultaneously through the PCI bus, and can provide data flow control and routing functions, enabling efficient data exchange and communication between PCI devices. PCI switch chip devices are typically used in servers, workstations, and other systems requiring a large number of PCI device connections.
[0071] PCI switch devices have a fault threshold register, which is set to display a fault threshold. When the number of faults reaches or exceeds this threshold, the corresponding fault handling method is triggered. The fault threshold can include various types of faults, such as transmission faults, reception faults, clock faults, etc. By setting these fault thresholds, fault detection and handling can be performed on the device, ensuring its reliability and stability.
[0072] In some embodiments, an EP (Endpoint) device refers to a device connected to a PCI switch device or a device connected to a PCI link, such as a network card, graphics card, sound card, or storage controller. EP devices communicate and transmit data via the PCI bus, the PCI switch device motherboard, or other devices.
[0073] In some embodiments, fault handling functions in EP devices include, but are not limited to, Advanced Error Reporting (AER). AER is a function in PCI devices used to report and handle device errors. The AER function allows PCI devices to generate error reports when errors occur and send them to the host system's operating system or other management entities. These error reports can include the error type, error location, and other relevant information to help system administrators diagnose and handle device errors. The AER function can also notify the host system of link errors through the error reporting mechanism on the PCI link, so that the system can take appropriate measures to handle the error, such as resetting or reconfiguring the device, rerouting data flows, etc. This helps improve system reliability and stability and reduces the impact of errors on system performance and data integrity.
[0074] In some embodiments, PCI device failures include, but are not limited to:
[0075] 1) Correctable or repairable fault events: These refer to errors that can be corrected using error correction codes or other techniques, such as checksum errors in data transmission.
[0076] 2) Uncorrectable errors or fault events that cannot be repaired: These refer to errors that cannot be corrected by error correction codes or other technologies, such as serious errors in data transmission.
[0077] 3) Fatal error, or failure event that causes the PCI device to malfunction: refers to an error that seriously affects the function of the device, causing the device to malfunction.
[0078] 4) Non-fatal errors, or failure events that allow the PCI device to continue operating: These are errors that affect the device's functionality but allow the device to continue working.
[0079] In an exemplary embodiment, when the PCI device connected to the PCI bridge link is only an EP device, activating the fault handling function of the EP device includes: setting the AER function of the EP device to an enabled state or an activated state to activate the fault handling function of the EP device, wherein the AER function includes the fault handling function.
[0080] In some embodiments, the EP device can be a device in a server connected to multiple processors via PCI links. For example, as shown in Figure 3, when the fault handling system includes a BIOS, the devices connected to CPU0 and CPU1 only include the EP device. When the BIOS scans all PCI devices on the current PCI link and finds only the EP device, the BIOS enables the AER function of all EP devices on the PCI link.
[0081] In some embodiments, setting the AER function of the EP device to an enabled or activated state can be done by following these steps:
[0082] 1) After the BIOS system boots up, enter the BIOS system setup interface;
[0083] 2) In the BIOS setup interface, find the "PCI device configuration" or similar option;
[0084] 3) In the PCI device configuration menu, find the "Advanced Error Reporting (AER)" or similar option;
[0085] 4) Set the "Advanced Error Reporting (AER)" option to "Enabled" or "Enabled";
[0086] 5) Save the settings and exit the BIOS setup interface;
[0087] 6. Ensure that the AER function of all PCI devices has been successfully enabled. You can check this using the corresponding tools or commands in the operating system.
[0088] In this embodiment, when the PCI device connected to the PCI bridge link is only an EP device, the faults of the EP device can be accurately detected and handled by setting the AER function of the EP device to the enabled or started state.
[0089] In an exemplary embodiment, when the PCI device connected on the PCI bridge link is only an EP device, after activating the fault handling function of the EP device, the method further includes: monitoring a first fault event of the EP device, wherein the first fault event includes at least one of the following: a first fault event that can be repaired, a first fault event that cannot be repaired, a first fault event that causes the EP device to be unable to operate, and a first fault event that allows the EP device to continue operating; upon detecting the first fault event, reporting the first fault event to the target operating system and the baseboard management controller in the server; wherein the target operating system is configured to perform one of the following operations based on the first fault event: restarting the server, performing a self-test operation, and generating a fault prompt message; the baseboard management controller is configured to perform at least one of the following operations based on the first fault event: recording the first fault event, performing a fault troubleshooting operation, and repairing the first fault event.
[0090] In some embodiments, the first reparable fault event includes a correctable error in the EP device, such as a checksum error during data transmission; the first non-reparable fault event includes an uncorrectable error in the EP device, such as a critical error during data transmission; the first fault event causing the EP device to malfunction includes a fatal error in the EP device; and the first fault event allowing the EP device to continue operating includes a non-fatal error in the EP device. For example, when any one of a correctable, uncorrectable, fatal, or non-fatal error occurs in the EP device connected to the PCI link in the server, the BIOS immediately reports this fault information to the target operating system and the Board Management Controller (BMC), notifying the target operating system and the BMC to perform fault handling operations, or prompting the replacement of the EP device.
[0091] This embodiment monitors fault events of the EP device and reports these events to the target operating system and baseboard management controller in the server, thereby enabling timely fault handling.
[0092] In an exemplary embodiment, when the PCI device connected on the PCI bridge link includes a PCI switch device, setting a fault threshold for the PCI switch device includes:
[0093] When the PCI devices connected on the PCI bridge link include PCI switch devices, determine the target register set in the PCI switch device, wherein the target register is set to record the number of failures of the PCI switch device; and set the failure threshold of the target register according to preset conditions.
[0094] In some embodiments, the target register is a fault threshold register set in the PCI switch device. It is configured to set a fault threshold, and when the number of faults reaches or exceeds the fault threshold, a corresponding fault handling method is triggered. The target register includes a counter, which is configured to accumulate the number of faults of the PCI switch device.
[0095] In some embodiments, the PCI device connected on the PCI bridge link includes a PCI switch device, including one of the following:
[0096] The PCI devices connected to the PCI bridge link are only PCI switch devices; for example, as shown in Figure 4, CPU0 and CPU1 are both connected to PCI switch devices.
[0097] The PCI devices connected on the PCI bridge link are only PCI switch devices, and one or more EP devices are connected to the PCI switch device; for example, as shown in Figure 4, three EP devices are connected to the PCI switch device in CPU0.
[0098] The PCI devices connected on the PCI bridge link include PCI switch devices and EP devices. The PCI switch device is connected to the first processor in the server, and the EP device is connected to the second processor in the server. For example, as shown in Figure 5, CPU0 is connected to a PCI switch device, and CPU1 is connected to an EP device.
[0099] The PCI devices connected to the PCI bridge link include PCI switch devices and EP devices. The PCI switch device is connected to the first processor in the server, and one or more other EP devices are connected to the PCI switch device. The EP devices are connected to the second processor in the server. For example, as shown in Figure 5, CPU0 is connected to a PCI switch device, which has 3 EP devices connected to it, and CPU1 is connected to 3 EP devices.
[0100] It should be noted that there can be one or more PCI switch devices, and each PCI switch device can also connect to multiple EP devices. This embodiment can accurately determine the fault handling method for each type of PCI device by detecting the type of the connected device.
[0101] In some embodiments, setting a fault threshold for the target register according to preset conditions includes: setting the fault threshold for the target register according to preset conditions and through at least one of the following operations:
[0102] Determine the operating information of the PCI Switch device and set fault thresholds based on the operating information, wherein the operating information includes at least one of the following: operating environment, operating performance, operating security performance, and operating stability performance;
[0103] For example, the operating environment includes the temperature range, humidity range, and other environmental conditions of the PCI switch device. Based on the device's stability performance under different environmental conditions, corresponding fault thresholds can be set to ensure the device's reliability in different environments.
[0104] Operational performance: This includes performance indicators such as data transfer speed and bandwidth utilization of the PCI switch device. Based on the performance of the PCI switch device under different performance requirements, corresponding fault thresholds can be set to ensure the stability of the device under different loads.
[0105] Operational security features include security authentication, data encryption, and access control for the PCI switch device. Based on the device's performance under different security requirements, appropriate fault thresholds can be set to ensure the reliability of the PCI switch device under various security threats.
[0106] Operational stability performance: This includes the stability performance of the PCI switch device under long-term operation, such as whether disconnection or crash occurs. Based on the performance of the PCI switch device under different stability requirements, corresponding fault thresholds can be set to ensure the reliability of the device under long-term operation.
[0107] Based on the above operational information, appropriate fault thresholds can be set according to actual conditions to ensure the stable operation of PCI switch devices under different conditions.
[0108] A first number of PCI switch devices and a second number of EP devices connected to the PCI switch devices are determined, and a fault threshold is set based on the first and second numbers. In this embodiment, the fault threshold can be set according to the first number of PCI switch devices and the second number of connected EP devices. For example, if there are 5 PCI switch devices and 20 connected EP devices, an appropriate fault threshold (e.g., 100) can be set according to the actual situation to ensure the stability and reliability of the network. By monitoring and analyzing fault conditions in the network, the optimal fault threshold can be determined so that problems can be identified and resolved in a timely manner.
[0109] In one exemplary embodiment, after setting the fault threshold of the target register according to preset conditions, the method further includes: detecting the PCI switch device and one or more EP devices connected to the PCI switch device; when the PCI switch device or any EP device connected to the PCI switch device fails, triggering a counter in the target register to perform a fault accumulation operation, wherein the fault accumulation operation is used to accumulate the number of times the PCI switch device has failed. In this embodiment, when the PCI switch device or any EP device connected to it fails, the counter in the target register will be triggered to perform a fault accumulation operation. This means that the fault counter will record each fault event and accumulate it into the corresponding counter for subsequent analysis and troubleshooting. This helps to quickly identify the frequency and pattern of fault occurrence so that appropriate measures can be taken to repair the fault and improve the reliability and stability of the system.
[0110] In one exemplary embodiment, when the PCI device connected on the PCI bridge link includes a PCI switch device, before setting the fault threshold of the PCI switch device, the method further includes one of the following:
[0111] When one or more EP devices are connected to a PCI switch device, the fault handling function of the EP devices is enabled, and the PCI switch device receives the results of the EP devices' fault handling according to the fault handling function. For example, as shown in Figure 4, when three EP devices are connected to the PCI switch device connected to CPU0, the fault handling function (AER function) for all three EP devices is enabled. The fault handling function of the EP devices can also be enabled through the PCI switch device. The fault handling function of the EP devices can include automatic restart, fault alarm, error log recording, etc. When an EP device handles a fault according to the fault handling function, the PCI switch device can receive the processing results sent by the EP device, such as successful restart, alarm sending, etc. By receiving the processing results of the EP devices, the PCI switch device can promptly understand the fault handling status of the EP devices, helping administrators to take appropriate measures in a timely manner to ensure the normal operation of the system.
[0112] When a PCI switch device is connected to the first processor in the server, and an EP device is connected to the second processor in the server, the fault handling function of the EP device is disabled. For example, as shown in Figure 5, if CPU0 is connected to a PCI switch device and CPU1 is connected to three EP devices, then the AER function of the three EP devices connected to CPU1 is disabled. Thus, the error handling mechanism of the PCI link can be implemented by setting the fault threshold of the PCI switch device.
[0113] When a PCI switch is connected to the first processor in a server, and one or more other EP devices are connected to the PCI switch, and the EP devices are connected to the second processor in the server, the fault handling function of the EP devices is disabled, while the fault handling functions of the other EP devices are enabled. The PCI switch receives the results of the other EP devices' fault handling according to their respective fault handling functions. For example, as shown in Figure 5, if CPU0 is connected to a PCI switch with three EP devices connected to it, and CPU1 is also connected to three EP devices, then the AER function of the three EP devices connected to CPU1 is disabled, while the AER function of the three EP devices connected to the PCI switch is enabled. By receiving the processing results from the EP devices, the PCI switch can promptly understand the fault handling status of the EP devices, helping administrators to take timely measures to ensure the normal operation of the system.
[0114] In one exemplary embodiment, adjusting the transmission rate of a PCI bridge link based on a fault threshold includes: triggering an adjustment of the transmission rate of the PCI bridge link when the accumulated number of first faults of the PCI bridge link is greater than or equal to the fault threshold, thereby obtaining a first adjusted transmission rate.
[0115] In some embodiments, adjusting the transmission rate can be done by reducing the data transmission speed to decrease link load, for example, from 32GT / s to 16GT / s. This enables timely response to link failures, reduces data transmission interruptions and latency, and improves the stability and reliability of the PCI bridge link.
[0116] In some embodiments, if the cumulative number of failures of the PCI bridge link is greater than or equal to a failure threshold, the transmission rate of the PCI bridge link is adjusted to obtain a first adjusted transmission rate, including at least one of the following:
[0117] If the number of first failures is greater than or equal to the failure threshold, the transmission rate of the PCI Switch device is adjusted to obtain the first adjusted transmission rate.
[0118] If the number of first failures is greater than or equal to the failure threshold, the transmission rate of one or more EP devices connected to the PCI Switch device is adjusted to obtain the first adjusted transmission rate.
[0119] In some embodiments, the transfer rate of an EP device connected to a PCI switch device, or the transfer rate of the EP device itself, can be adjusted in the following ways: Software control: Using the management software or driver of the PCI switch device or EP device, the transfer rate of the connected EP device can be adjusted by setting parameters or configuration files. Hardware control: Some PCI switch devices or EP devices may have physical switches or buttons, allowing for manual operation directly on the device to adjust the transfer rate of the connected EP device. Remote control: Through a network or remote connection, the PCI switch device or EP device can be remotely controlled using remote management tools or protocols to adjust the transfer rate of the connected EP device. Regardless of the method used, it is necessary to ensure that the operation to adjust the transfer rate correctly identifies and affects the PCI switch device or EP device that needs adjustment to ensure the accuracy and effectiveness of the transfer rate adjustment. The first adjusted transfer rate can then be obtained using the above methods. Note that the transfer rate adjustment may require restarting the PCI switch device to take effect.
[0120] In some embodiments, if the accumulated number of PCI bridge link failures is greater than or equal to a failure threshold, the method triggers an adjustment to the PCI bridge link's transmission rate. After obtaining a first adjusted transmission rate, the method further includes: clearing the first failure count; re-recording the number of PCI bridge link failures to obtain a second failure count. If the second failure count is less than the failure threshold, the current transmission rate remains unchanged; if the second failure count is greater than or equal to the failure threshold, the method triggers another adjustment to the PCI bridge link's transmission rate. This process continues until the transmission rate stabilizes at a suitable value.
[0121] In some embodiments, after re-recording the number of PCI bridge link failures to obtain a second failure count, the method further includes: if the second failure count is greater than or equal to a failure threshold, continuing to trigger adjustment of the PCI bridge link's transmission rate to obtain a second adjusted transmission rate; and controlling the operation of the PCI bridge link based on the second adjusted transmission rate. If the second failure count does not reach the failure threshold, the operating status of the PCI bridge link will continue to be monitored, and the operation of adjusting the transmission rate will be triggered when necessary. This ensures that the PCI bridge link can make timely adjustments when a failure occurs, thereby guaranteeing the stable operation of the system.
[0122] In some embodiments, after controlling the operation of the PCI bridge link based on the second adjusted transmission rate, the method further includes: stopping the adjustment of the PCI bridge link's transmission rate when the second adjusted transmission rate is equal to the minimum rate of the PCI bridge link; and triggering an adjustment of the PCI bridge link's network bandwidth to obtain the adjusted network bandwidth when the second adjusted transmission rate is equal to the minimum rate of the PCI bridge link and the recorded third number of PCI bridge link failures is greater than or equal to a fault threshold, wherein the third number of failures is the number of failures recorded during the PCI bridge link's operation at the second adjusted transmission rate. When the second adjusted transmission rate is equal to the minimum rate of the PCI bridge link and the recorded number of PCI bridge link failures is greater than or equal to a set fault threshold, the network bandwidth of the PCI bridge link will be adjusted. This means that the system automatically reallocates the bandwidth of the PCI bridge link to ensure the stability and reliability of network transmission. This automatic adjustment helps the system react quickly in the event of a failure, thereby minimizing the impact on network performance.
[0123] In some embodiments, when the second adjusted transmission rate is equal to the minimum rate of the PCI bridge link and the number of recorded third failures of the PCI bridge link is greater than or equal to a failure threshold, the method triggers the adjustment of the network bandwidth of the PCI bridge link. After obtaining the adjusted network bandwidth, the method further includes: generating a second failure event when the adjusted network bandwidth is equal to the minimum network bandwidth of the PCI bridge link, wherein the second failure event includes at least one of the following: a second failure event that can be repaired, a second failure event that cannot be repaired, a second failure event that causes the PCI bridge link to be unable to operate, and a second failure event that allows the PCI bridge link to continue operating; sending the second failure event to a baseboard management controller in a server, wherein the baseboard management controller is configured to record the second failure event and generate fault handling information, wherein the fault handling information is used to instruct the replacement of the faulty PCI switch device, and / or, the replacement of the EP device connected to the faulty PCI switch device.
[0124] In some embodiments, a repairable second failure event refers to an error that can be corrected using error correction codes or other techniques, such as a checksum error in data transmission. An unrepairable second failure event refers to an error that cannot be corrected using error correction codes or other techniques, such as a critical error in data transmission. A second failure event that causes a PCI bridge link to malfunction refers to an error that severely affects the functionality of the PCI bridge link, causing the device to malfunction. A second failure event that allows the PCI bridge link to continue operating refers to an error that has some impact on the device's functionality, but the device can still continue to operate. A second failure event may be an interruption or loss of network connectivity, resulting in data transmission failure or severe delay.
[0125] In some embodiments, the baseboard management controller records detailed information about the second fault event, including the time, location, cause, and impact of the fault. The baseboard management controller then generates fault handling information, including the handling measures, repair plan, responsible person, and estimated completion time. This information helps managers and technicians better understand the fault situation, take timely countermeasures, and ensure timely fault repair.
[0126] In an exemplary embodiment, after adjusting the transmission rate of the PCI bridge link based on a fault threshold, the method further includes: generating target information and sending the target information to a baseboard management controller deployed in a server, wherein the target information includes a target transmission rate, the target transmission rate being the rate after adjusting the transmission rate of the PCI bridge link, and the target transmission rate being less than the transmission rate of the PCI bridge link before adjustment; recording the target transmission rate through the baseboard management controller and generating a prompt message to indicate that the first number of faults in the PCI bridge link exceeds the fault threshold.
[0127] In some embodiments, the process of generating target information includes determining a target transmission rate value, encapsulating that value into a data packet or message, and then sending it over the network to the server where the baseboard management controller resides. The baseboard management controller can adjust the transmission rate of the PCI bridge link based on the received target information to achieve the desired target transmission rate. This ensures that data transmission in the system can occur at the set rate, thereby meeting specific performance requirements.
[0128] In some embodiments, the baseboard management controller can record the target transmission rate and generate corresponding prompts to help the user monitor and manage the transmission rate. The prompts may include, but are not limited to:
[0129] 1) If the current transmission rate reaches or exceeds the preset threshold, the user is reminded that there may be network congestion or bandwidth limitations, and it is recommended to take corresponding measures to optimize the network environment.
[0130] 2) The current transmission rate is lower than expected, which may indicate a device malfunction or network problem. It is recommended to check the relevant devices and network connections to improve the transmission rate.
[0131] 3) For specific tasks or applications, based on the set target transmission rate, remind the user whether the current transmission rate meets expectations, so that settings can be adjusted or other measures can be taken in a timely manner.
[0132] By recording and generating alerts, the baseboard management controller can help users promptly identify and resolve transmission rate issues, thereby improving transmission efficiency and user experience.
[0133] In one exemplary embodiment, detecting PCI devices connected on a PCI bridge link includes: starting a fault handling system when the server is started, detecting N levels on the PCI bridge link, where N is a natural number greater than or equal to 1; and determining the type of PCI device in each level.
[0134] In some embodiments, when the server is started, a fault handling system is initiated to detect N layers on the PCI bridge link to ensure the stability and reliability of the connection.
[0135] In some embodiments, the types of PCI devices in each tier include: Hardware tier: PCI device types may include graphics cards, network cards, sound cards, RAID cards, etc. Driver tier: PCI device types may include graphics drivers, network drivers, sound card drivers, etc. Application tier: PCI device types may include graphics processors, network interface cards, audio input / output devices, etc.
[0136] The present application will now be described with reference to optional embodiments:
[0137] This optional embodiment uses the method of BIOS controlling PCI link error handling as an example for illustration. This optional embodiment is mainly for the method of error information handling when the CPU-side PCI bridge link is connected to different devices.
[0138] Because the PCI link bus can support both PCI terminal devices (i.e., EP devices) and PCI switch chips (i.e., PCI switch devices), but PCI terminal devices only support AER (Advanced Error Recovery) to recover from errors, unrecoverable errors, fatal errors, and non-fatal errors, while PCI switch chips can still support PCI terminal settings or PCI switch chips themselves, and since PCI switch chips are bridge-type devices, error repair can be performed on the terminal devices under the PCI switch by setting error threshold parameters. When the number of bridge errors on the PCI switch reaches the threshold setting, the PCI switch chip will implement a speed reduction mechanism to reduce the number of errors. However, because there are three possible scenarios for the PCI link: the first is that there is only a PCI terminal device and no PCI switch chip; the second is that there is only a PCI switch chip and no PCI terminal device; and the third is that there are both PCI terminal devices and PCI switch chips, it is easy to implement error handling mechanisms for the first and second scenarios. However, for the third scenario, it is necessary to disable the error handling mechanism for PCI terminal devices and enable the PCI switch chip. The error handling mechanism of the switch has been implemented, and the error handling mechanism of the PCI link has been implemented. Since the number of PCI bridges supported by different CPU architectures is different, this optional embodiment only describes the error information handling mechanism with one PCI bridge link as the main body.
[0139] Figure 6 shows the interaction diagram of the various devices performing error information processing in this optional embodiment. The interaction process is as follows:
[0140] S601, the server is powered on and the BIOS starts;
[0141] In S602, the BIOS scans the PCI devices under the PCI bridge link to confirm whether a PCI switch chip exists in the current link. If no PCI switch chip is found, the AER function of the PCI terminal devices under the current link is enabled, and errors are reported. If a PCI switch chip is detected, the AER function of the PCI terminal EP devices under the non-switch chip bridge is disabled, and only the AER function of the PCI terminal devices under the current switch chip bridge is enabled, and an error threshold is set. When the error counter reaches the threshold, the link speed and bandwidth are reduced, and the SMI (System Management Interrupt) interrupt signal is used to notify the BMC to record the error.
[0142] The S603 BMC receives error messages reported by the BIOS and records information about reduced link speed and bandwidth.
[0143] This embodiment adopts different measures for different connection situations, which effectively ensures that errors of PCI devices are reported in a timely manner. The measures taken through the Switch chip bridge reduce the time loss caused by frequent replacement of PCI terminal devices on the server, improve the security, reliability and stability of the server, and effectively ensure the cost-effectiveness of the server's RAS function.
[0144] Figure 7 shows a flowchart of the BIOS error information processing in this optional embodiment, including the following steps:
[0145] S701, BIOS confirms server power-on and startup;
[0146] S702, the BIOS scans and detects all PCI devices on the PCI bridge link, including all PCI devices such as PCI terminal devices and PCI switch chip devices;
[0147] S703, the BIOS determines whether the scanned PCI devices are only EP devices;
[0148] S704: When the BIOS scans all PCI devices in the current PCI bridge link and finds only EP devices, the BIOS enables the AER function for all EP devices in the PCI bridge link.
[0149] S705: When any of the correctable, uncorrectable, fatal, or non-fatal errors occur in the EP device, the BIOS immediately reports this fault information to the operating system (OS) and BMC, notifying the OS and BMC to replace or replace the device.
[0150] S706 When the BIOS scans and finds that the first level of all PCI devices in the current PCI bridge link is only the PCI Switch chip, even if the PCI Switch chip's downlink port is connected to the EP device, the AER function of the EP device will be disabled. At the same time, the error threshold register of the PCI Switch chip will be set to the fault threshold, for example, the fault threshold will be set to 1000.
[0151] S707: When any AER error occurs in the EP device at the lower end of the PCI Switch chip, the fault threshold of the PCI Switch chip is incremented by 1.
[0152] S708 determines whether the fault count in the threshold register of the PCI switch chip has reached the fault threshold.
[0153] S709 When the count in the threshold register reaches or exceeds the fault threshold, the process switches to S710. The PCI bridge link will automatically trigger the speed reduction function (e.g., from 32GT / s to 16GT / s) and restart the count in the threshold register.
[0154] S711, if S704 still exceeds the set fault threshold, proceed to S715, continue to reduce the speed until it reaches the current minimum transmission rate of the PCI link; if the minimum transmission rate still exceeds the set threshold, reduce the PCI link bandwidth (e.g., from x16 to x8) until the minimum speed and minimum bandwidth (e.g., 2.5GT / s, x1) are reduced, and continue to execute steps S704-S705;
[0155] S712 determines whether the PCI link bandwidth has been reduced to the minimum bandwidth;
[0156] In S713, the BIOS will notify the BMC to record the current PCI link speed reduction information via the SMI interrupt signal. This is set to inform the user that the error data of the current PCI link has exceeded the threshold setting.
[0157] S714, restarts the threshold register count.
[0158] It should be noted that when the BIOS scans the first level of all PCI devices in the current PCI link and there are both PCI switch chips and EP devices, even if the downstream port of the PCI switch chip is connected to the EP device, the AER function of the EP device still needs to be turned off. Only the error threshold of the PCI switch should be retained and the steps S706-S705 should be executed.
[0159] In summary, this optional embodiment involves the BIOS scanning PCI devices under the PCI bridge link during server startup to confirm the presence of a PCI switch chip bridge in the current link. If no PCI switch chip bridge is found, the AER (Advanced Error Reporting) function of the PCI terminal devices under the current link is enabled for error reporting. If a PCI switch chip bridge is detected, the AER function of PCI terminal EP (Extended Premises Equipment) devices not under the switch chip bridge is disabled, and only the AER function of PCI terminal devices under the current switch chip bridge is enabled, with an error threshold set. When the error counter reaches the threshold, link speed reduction and bandwidth reduction measures are implemented, and the BIOS uses an SMI interrupt to notify the BMC (Browser Control Center) for recording. This optional embodiment adopts different measures for different link situations, effectively ensuring timely reporting of PCI device errors and reducing the time loss caused by frequent PCI terminal device replacements through measures implemented via the switch chip bridge. It not only reduces the problem of frequent terminal device replacements by proactively maintaining errors but also ensures server reliability, providing highly beneficial measures for server stability and security, data center operation and maintenance, and the overall lifespan of the server.
[0160] This embodiment also provides a fault handling system. As shown in Figure 8, which is a structural block diagram of the fault handling system in this embodiment, the fault handling system is installed on the motherboard of the server. The fault handling system includes a target processor, which is connected to a PCI bridge deployed in the server. The target processor is configured to implement the steps of the above-described method.
[0161] This embodiment also provides a server, as shown in Figure 9, which is a structural block diagram of the server in this embodiment, including: a PCI bridge link, and a fault handling system as described above, wherein PCI devices are connected to the PCI bridge link.
[0162] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the related technology, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM (Read-Only Memory) / RAM (Random Access Memory), magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.
[0163] This embodiment also provides a fault handling device for a PCI device, which is configured to implement the above embodiments and optional implementations; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0164] Figure 10 is a structural block diagram of a fault handling device for a PCI device according to an embodiment of this application. The device is applied to a fault handling system, which is mounted on the motherboard of a server and connected to a PCI bridge deployed in the server. As shown in Figure 10, the device includes:
[0165] The first detection module 1002 is configured to detect PCI devices connected to the PCI bridge link, wherein the PCI devices include EP devices and / or PCI switch devices, and multiple EP devices are allowed to be connected to the PCI switch device; the first processing module 1004 is configured to determine the fault handling method of the PCI device based on the detected PCI devices, and handle the fault of the PCI device according to the fault handling method; wherein the first processing module includes: a first processing unit, configured to activate the fault handling function of the EP device when the PCI device connected to the PCI bridge link is only an EP device, the fault handling function is a function set in the EP device, and the fault handling function is used to handle the fault of the EP device; a second processing unit, configured to set the fault threshold of the PCI switch device when the PCI device connected to the PCI bridge link includes a PCI switch device, and adjust the transmission rate of the PCI bridge link based on the fault threshold, the fault threshold is used to represent the threshold for recording the number of faults of the PCI switch device, and the fault handling method includes the fault handling function and the method of setting the fault threshold of the PCI switch device.
[0166] In one exemplary embodiment, the first processing unit includes: a first setting subunit, configured to enable or start the AER function of the EP device when the PCI device connected on the PCI bridge link is only an EP device, so as to start the fault handling function of the EP device, wherein the AER function includes the fault handling function.
[0167] In one exemplary embodiment, the apparatus further includes: a first monitoring module, configured to monitor a first fault event of the EP device after activating the fault handling function of the EP device when the PCI device connected on the PCI bridge link is only an EP device, wherein the first fault event includes at least one of the following: a first fault event that can be repaired, a first fault event that cannot be repaired, a first fault event that causes the EP device to be unable to operate, and a first fault event that allows the EP device to continue operating; and a first reporting module, configured to report the first fault event to the target operating system and the baseboard management controller in the server when the first fault event is detected; wherein the target operating system is configured to perform one of the following operations based on the first fault event: restart the server, perform a self-test operation, and generate a fault prompt message; and the baseboard management controller is configured to perform at least one of the following operations based on the first fault event: record the first fault event, perform a fault troubleshooting operation, and repair the first fault event.
[0168] In one exemplary embodiment, the second processing unit includes: a first determining subunit configured to determine a target register set in a PCI switch device when the PCI device connected on the PCI bridge link includes a PCI switch device, wherein the target register is configured to record the number of failures of the PCI switch device; and a first setting subunit configured to set a failure threshold of the target register according to preset conditions.
[0169] In one exemplary embodiment, the PCI devices connected to the PCI bridge link include PCI switch devices, including one of the following: the PCI devices connected to the PCI bridge link are only PCI switch devices; the PCI devices connected to the PCI bridge link are only PCI switch devices, and one or more EP devices are connected to the PCI switch device; the PCI devices connected to the PCI bridge link include PCI switch devices and EP devices, and the PCI switch device is connected to a first processor in the server, and the EP device is connected to a second processor in the server; the PCI devices connected to the PCI bridge link include PCI switch devices and EP devices, and the PCI switch device is connected to a first processor in the server, and one or more other EP devices are connected to the PCI switch device, and the EP devices are connected to a second processor in the server.
[0170] In an exemplary embodiment, the first setting subunit includes: a first setting submodule, configured to set a fault threshold of a target register according to preset conditions and by at least one of the following operations: determining operating information of a PCI Switch device and setting a fault threshold based on the operating information, wherein the operating information includes at least one of the following: operating environment, operating performance, operating security performance, and operating stability performance; determining a first number of PCI Switch devices and a second number of EP devices connected in the PCI Switch devices, and setting a fault threshold based on the first number and the second number.
[0171] In one exemplary embodiment, the apparatus further includes: a first detection module configured to detect a PCI switch device and one or more EP devices connected to the PCI switch device after setting a fault threshold in the target register according to preset conditions; and a first trigger module configured to trigger a counter in the target register to perform a fault accumulation operation when the PCI switch device or any EP device connected to the PCI switch device fails, wherein the fault accumulation operation is used to accumulate the number of times the PCI switch device has failed.
[0172] In one exemplary embodiment, the above-described apparatus further includes one of the following: a first enabling module, configured to enable the fault handling function of the EP device when, in the case that the PCI device connected on the PCI bridge link includes a PCI Switch device, before setting a fault threshold for the PCI Switch device, and in the case that one or more EP devices are connected to the PCI Switch device, and receive the result of the EP device processing the fault according to the fault handling function through the PCI Switch device; a first disabling module, configured to disable the fault handling function of the EP device when, in the case that the PCI device connected on the PCI bridge link includes a PCI Switch device, before setting a fault threshold for the PCI Switch device, and in the case that the PCI Switch device is connected to a first processor in the server, and the EP device is connected to a second processor in the server; and a second processing module, configured to disable the fault handling function of the EP device, enable other fault handling functions of other EP devices, and receive the result of the other EP devices processing the fault according to other fault handling functions through the PCI Switch device when, in the case that the PCI device connected on the PCI bridge link includes a PCI Switch device, before setting a fault threshold for the PCI Switch device, and in the case that the PCI Switch device is connected to a first processor in the server, one or more other EP devices are connected to the PCI Switch device, and the EP devices are connected to a second processor in the server.
[0173] In one exemplary embodiment, the second processing unit includes: a first triggering subunit, configured to trigger an adjustment of the transmission rate of the PCI bridge link to obtain a first adjusted transmission rate when the number of accumulated failures of the PCI bridge link is greater than or equal to a failure threshold.
[0174] In one exemplary embodiment, the first trigger subunit includes at least one of the following: a first trigger submodule configured to trigger adjustment of the transmission rate of the PCI Switch device to obtain a first adjusted transmission rate when the number of first faults is greater than or equal to a fault threshold; and a second trigger submodule configured to trigger adjustment of the transmission rate of one or more EP devices connected in the PCI Switch device to obtain the first adjusted transmission rate when the number of first faults is greater than or equal to a fault threshold.
[0175] In an exemplary embodiment, the apparatus further includes: a first clearing module, configured to trigger an adjustment of the transmission rate of the PCI bridge link when the accumulated number of failures of the PCI bridge link is greater than or equal to a failure threshold, and after obtaining a first adjusted transmission rate, clear the first failure count; and a first recording module, configured to re-record the number of failures of the PCI bridge link to obtain a second failure count.
[0176] In one exemplary embodiment, the apparatus further includes: a first adjustment module, configured to re-record the number of times the PCI bridge link fails, obtain a second number of failures, and then, if the second number of failures is greater than or equal to a failure threshold, continue to trigger the adjustment of the transmission rate of the PCI bridge link to obtain a second adjusted transmission rate; and a first control module, configured to control the operation of the PCI bridge link based on the second adjusted transmission rate.
[0177] In one exemplary embodiment, the apparatus further includes: a first stop module configured to stop adjusting the transmission rate of the PCI bridge link after controlling the operation of the PCI bridge link based on the second adjusted transmission rate, provided that the second adjusted transmission rate is equal to the minimum rate of the PCI bridge link; and a second trigger module configured to trigger adjusting the network bandwidth of the PCI bridge link to obtain the adjusted network bandwidth when the second adjusted transmission rate is equal to the minimum rate of the PCI bridge link and the number of third faults recorded for the PCI bridge link is greater than or equal to a fault threshold, wherein the number of third faults is the number of faults recorded during the operation of the PCI bridge link at the second adjusted transmission rate.
[0178] In one exemplary embodiment, the apparatus further includes: a first generation module configured to, when the second adjusted transmission rate is equal to the minimum rate of the PCI bridge link and the recorded third number of failures of the PCI bridge link is greater than or equal to a failure threshold, trigger adjustment of the network bandwidth of the PCI bridge link; after obtaining the adjusted network bandwidth, when the adjusted network bandwidth is equal to the minimum network bandwidth of the PCI bridge link, generate a second failure event, wherein the second failure event includes at least one of the following: a second failure event that can be repaired, a second failure event that cannot be repaired, a second failure event that causes the PCI bridge link to be unable to operate, and a second failure event that allows the PCI bridge link to continue operating; and a first transmission module configured to send the second failure event to a baseboard management controller in a server, wherein the baseboard management controller is configured to record the second failure event and generate fault handling information, wherein the fault handling information is used to instruct the replacement of the faulty PCI switch device, and / or, the replacement of the EP device connected to the faulty PCI switch device.
[0179] In one exemplary embodiment, the apparatus further includes: a second generation module configured to generate target information after adjusting the transmission rate of the PCI bridge link based on a fault threshold, and send the target information to a baseboard management controller deployed in a server, wherein the target information includes a target transmission rate, the target transmission rate being the rate after adjusting the transmission rate of the PCI bridge link, and the target transmission rate being less than the transmission rate of the PCI bridge link before adjustment; and a second recording module configured to record the target transmission rate through the baseboard management controller and generate a prompt message to indicate that the first number of faults in the PCI bridge link exceeds a fault threshold.
[0180] In one exemplary embodiment, the first detection module includes: a first detection unit configured to start a fault handling system when the server is started, and to detect N layers on the PCI bridge link, where N is a natural number greater than or equal to 1; and a first determination unit configured to determine the type of PCI device in each layer.
[0181] It should be noted that the above modules can be implemented by software or hardware. For the latter, they can be implemented in the following ways, but are not limited to: all the above modules are located in the same processor; or, the above modules are located in different processors in any combination.
[0182] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above method embodiments.
[0183] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps in any of the above method embodiments.
[0184] The embodiments described herein also provide a computer program that includes computer instructions stored in a non-volatile computer-readable storage medium; a processor of a computer device reads the computer instructions from the non-volatile computer-readable storage medium and executes the computer instructions, causing the computer device to perform the steps in any of the above method embodiments.
[0185] This application also provides a non-volatile computer-readable storage medium storing a computer program that, when executed by a processor, performs a fault handling method for a PCI device.
[0186] Figure 11 shows a schematic diagram of an embodiment of the non-volatile computer-readable storage medium for the fault handling method of the PCI device provided in this application. Taking the non-volatile computer-readable storage medium shown in Figure 11 as an example, the non-volatile computer-readable storage medium 1101 stores a computer program 1102 that executes the above method when executed by a processor.
[0187] Figure 12 shows a schematic diagram of the hardware structure of an electronic device according to an embodiment of the fault handling method for the PCI device provided in this application.
[0188] Taking the device shown in Figure 12 as an example, the device includes a processor 1201 and a memory 1202.
[0189] The processor 1201 and the memory 1202 can be connected via a bus or other means. Figure 12 shows an example of connection via a bus.
[0190] The memory 1202, as a non-volatile computer-readable storage medium, can be configured to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the PCI device fault handling method in this embodiment. The processor 1201 executes various server functions and data processing by running the non-volatile software programs, instructions, and modules stored in the memory 1202, thereby implementing the PCI device fault handling method.
[0191] The memory 1202 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the PCI device's fault handling method, etc. Furthermore, the memory 1202 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 1202 may optionally include memory remotely located relative to the processor 1201, and these remote memories can be connected to the local module via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0192] The computer instructions 1203 corresponding to the fault handling methods of one or more PCI devices are stored in the memory 1202. When executed by the processor 1201, the fault handling methods of the PCI devices in any of the above method embodiments are executed.
[0193] Any embodiment of the computer device that performs the above-described fault handling method for the PCI device can achieve the same or similar effects as any of the aforementioned method embodiments.
[0194] Finally, it should be noted that those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program for the PCI device fault handling method can be stored in a non-volatile computer-readable storage medium. When executed, the program can include the processes of the embodiments of the above methods. The non-volatile readable storage medium for the program can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc. The above computer program embodiments can achieve the same or similar effects as any of the corresponding foregoing method embodiments.
[0195] The above are exemplary embodiments disclosed in this application. However, it should be noted that various changes and modifications can be made without departing from the scope of the embodiments disclosed in this application as defined by the claims. The functions, steps, and / or actions of the methods according to the disclosed embodiments described herein do not need to be performed in any particular order. Furthermore, although the elements disclosed in the embodiments of this application may be described or claimed individually, they may be understood as multiple unless explicitly limited to a singular number.
[0196] It should be understood that, as used herein, the singular form “a” is intended to include the plural form as well, unless the context clearly supports an exception. It should also be understood that, as used herein, “and / or” refers to any and all possible combinations of one or more of the associated listed items.
[0197] The example numbers disclosed in the above application are for descriptive purposes only and do not represent the superiority or inferiority of the examples.
[0198] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a non-volatile computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0199] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of the disclosure of the embodiments of this application (including the claims) is limited to these examples; under the concept of the embodiments of this application, the technical features of the above embodiments or different embodiments can also be combined, and there are many other variations of different aspects of the embodiments of this application as described above, which are not provided in detail for the sake of brevity. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the embodiments of this application should be included within the protection scope of the embodiments of this application.
Claims
1. A fault handling method for a peripheral component interconnect (PCI) device, characterized in that, The method, applied to a fault handling system mounted on a server motherboard and connected to a PCI bridge deployed within the server, includes: The PCI devices connected on the PCI bridge link are detected, wherein the PCI devices include terminal equipment (EP devices) and / or peripheral component interconnection switches (PCI Switch devices), and multiple EP devices are allowed to be connected on the PCI Switch device; Based on the detected PCI device, determine the fault handling method for the PCI device, and handle the fault of the PCI device according to the fault handling method; Where the PCI device connected to the PCI bridge link is only the EP device, the fault handling function of the EP device is activated. The fault handling function is a function set in the EP device and is used to handle the faults of the EP device. Where the PCI device connected to the PCI bridge link includes the PCI Switch device, a fault threshold of the PCI Switch device is set, and the transmission rate of the PCI bridge link is adjusted based on the fault threshold. The fault threshold is used to represent a threshold for recording the number of faults of the PCI Switch device. The fault handling method includes the fault handling function and the method of setting the fault threshold of the PCI Switch device.
2. The method according to claim 1, characterized in that, If the PCI device connected to the PCI bridge link is only the EP device, the fault handling function of the EP device is activated, including: When the PCI device connected to the PCI bridge link is only the EP device, the Advanced Error Reporting (AER) function of the EP device is set to enabled or started to activate the fault handling function of the EP device, wherein the fault handling function is included in the AER function.
3. The method according to claim 1, characterized in that, When the PCI device connected to the PCI bridge link is only the EP device, after activating the fault handling function of the EP device, the method further includes: The first fault event of the EP device is monitored, wherein the first fault event includes at least one of the following: a first fault event that can be repaired, a first fault event that cannot be repaired, a first fault event that causes the EP device to be unable to operate, and a first fault event that allows the EP device to continue to operate. Upon detecting the first fault event, the first fault event is reported to the target operating system and baseboard management controller in the server; The target operating system is configured to perform one of the following operations based on the first fault event: restart the server, perform a self-test, and generate a fault message; the baseboard management controller is configured to perform at least one of the following operations based on the first fault event: record the first fault event, perform a fault troubleshooting operation, and repair the first fault event.
4. The method according to claim 1, characterized in that, When the PCI device connected to the PCI bridge link includes the PCI switch device, setting a fault threshold for the PCI switch device includes: In the case where the PCI device connected on the PCI bridge link includes the PCI switch device, a target register set in the PCI switch device is determined, wherein the target register is set to record the number of failures of the PCI switch device; The fault threshold of the target register is set according to preset conditions.
5. The method according to claim 4, characterized in that, The PCI devices connected on the PCI bridge link include the PCI switch devices, including one of the following: The PCI device connected to the PCI bridge link is only the PCI switch device; The PCI device connected to the PCI bridge link is only the PCI switch device, and one or more of the EP devices are connected to the PCI switch device; The PCI devices connected on the PCI bridge link include the PCI switch device and the EP device, wherein the PCI switch device is connected to the first processor in the server, and the EP device is connected to the second processor in the server; The PCI devices connected on the PCI bridge link include the PCI switch device and the EP device. The PCI switch device is connected to a first processor in the server, and one or more other EP devices are connected to the PCI switch device. The EP devices are connected to a second processor in the server.
6. The method according to claim 4, characterized in that, Setting the fault threshold of the target register according to preset conditions includes: The fault threshold of the target register is set according to the preset conditions and by at least one of the following operations: Determine the operating information of the PCI Switch device and set the fault threshold based on the operating information, wherein the operating information includes at least one of the following: operating environment, operating performance, operating security performance, and operating stability performance; A first number of PCI switch devices and a second number of EP devices connected to the PCI switch devices are determined, and the fault threshold is set based on the first number and the second number.
7. The method according to claim 4, characterized in that, After setting the fault threshold of the target register according to preset conditions, the method further includes: Detect the PCI switch device and one or more EP devices connected to the PCI switch device; When the PCI Switch device fails or any of the EP devices connected to the PCI Switch device fails, the counter in the target register is triggered to perform a fault accumulation operation, wherein the fault accumulation operation is used to accumulate the number of times the PCI Switch device has failed.
8. The method according to claim 1, characterized in that, When the PCI device connected to the PCI bridge link includes the PCI switch device, before setting the fault threshold of the PCI switch device, the method further includes one of the following: When one or more EP devices are connected to the PCI Switch device, the fault handling function of the EP device is enabled, and the PCI Switch device receives the result of the EP device handling the fault according to the fault handling function. When the PCI Switch device is connected to the first processor in the server and the EP device is connected to the second processor in the server, the fault handling function of the EP device is disabled; When the PCI Switch device is connected to the first processor in the server, and one or more other EP devices are connected to the PCI Switch device, and the EP devices are connected to the second processor in the server, the fault handling function of the EP devices is turned off, other fault handling functions of the other EP devices are turned on, and the results of the other EP devices handling faults according to the other fault handling functions are received through the PCI Switch device.
9. The method according to claim 1, characterized in that, Adjusting the transmission rate of the PCI bridge link based on the fault threshold includes: If the cumulative number of failures of the PCI bridge link is greater than or equal to the failure threshold, the transmission rate of the PCI bridge link is adjusted to obtain a first adjusted transmission rate.
10. The method according to claim 9, characterized in that, If the cumulative number of failures of the PCI bridge link is greater than or equal to the failure threshold, the transmission rate of the PCI bridge link is adjusted to obtain a first adjusted transmission rate, including at least one of the following: If the first number of faults is greater than or equal to the fault threshold, the transmission rate of the PCI Switch device is adjusted to obtain the first adjusted transmission rate. If the first number of faults is greater than or equal to the fault threshold, the transmission rate of one or more EP devices connected to the PCI Switch device is adjusted to obtain the first adjusted transmission rate.
11. The method according to claim 9, characterized in that, If the cumulative number of failures of the PCI bridge link is greater than or equal to the failure threshold, the method further includes triggering an adjustment to the transmission rate of the PCI bridge link. After obtaining the first adjusted transmission rate, the method further includes: Clear the first fault count; The number of times the PCI bridge link failed was re-recorded to obtain the second failure count.
12. The method according to claim 11, characterized in that, After re-recording the number of failures of the PCI bridge link to obtain the second number of failures, the method further includes: If the second number of failures is greater than or equal to the failure threshold, the transmission rate of the PCI bridge link is adjusted to obtain the second adjusted transmission rate. The operation of the PCI bridge link is controlled based on the second adjusted transmission rate.
13. The method according to claim 12, characterized in that, After controlling the operation of the PCI bridge link based on the second adjusted transmission rate, the method further includes: If the second adjusted transmission rate is equal to the minimum rate of the PCI bridge link, the adjustment of the transmission rate of the PCI bridge link shall be stopped. If the second adjusted transmission rate is equal to the minimum rate of the PCI bridge link, and the third number of failures recorded for the PCI bridge link is greater than or equal to the failure threshold, the network bandwidth of the PCI bridge link is adjusted to obtain the adjusted network bandwidth. The third number of failures is the number of failures recorded during the operation of the PCI bridge link at the second adjusted transmission rate.
14. The method according to claim 12, characterized in that, If the second adjusted transmission rate is equal to the minimum rate of the PCI bridge link, and the recorded third number of failures of the PCI bridge link is greater than or equal to the failure threshold, the method further includes: triggering an adjustment of the network bandwidth of the PCI bridge link; after obtaining the adjusted network bandwidth, the method further includes: When the adjusted network bandwidth is equal to the minimum network bandwidth of the PCI bridge link, a second fault event is generated, wherein the second fault event includes at least one of the following: a second fault event that can be repaired, a second fault event that cannot be repaired, a second fault event that causes the PCI bridge link to be unable to operate, and a second fault event that allows the PCI bridge link to continue to operate; The second fault event is sent to the baseboard management controller in the server, wherein the baseboard management controller is configured to record the second fault event and generate the fault handling information, wherein the fault handling information is used to instruct the replacement of the faulty PCI switch device, and / or, the replacement of the EP device connected to the faulty PCI switch device.
15. The method according to claim 1, characterized in that, After adjusting the transmission rate of the PCI bridge link based on the fault threshold, the method further includes: Target information is generated and sent to the baseboard management controller deployed in the server. The target information includes a target transmission rate, which is the rate after adjusting the transmission rate of the PCI bridge link. The target transmission rate is less than the transmission rate of the PCI bridge link before adjustment. The target transmission rate is recorded by the baseboard management controller, and a prompt message is generated to indicate that the first number of failures of the PCI bridge link exceeds the fault threshold.
16. The method according to claim 1, characterized in that, Detecting PCI devices connected to the PCI bridge link includes: When the server is started, the fault handling system is started to detect N layers on the PCI bridge link, where N is a natural number greater than or equal to 1; Determine the type of the PCI device in each of the tiers.
17. A fault handling system, characterized in that, The fault handling system is mounted on the motherboard of the server. The fault handling system includes a target processor, which is connected to a PCI bridge deployed in the server. The target processor is configured to implement the steps of the method described in any one of claims 1-16.
18. A server, characterized in that, include: PCI bridge link, and fault handling system as described in claim 17 above, wherein PCI devices are connected to the PCI bridge link.
19. A fault handling device for a PCI device, characterized in that, An apparatus for use in a fault handling system, wherein the fault handling system is mounted on the motherboard of a server and is connected to a PCI bridge deployed in the server, the apparatus comprising: The first detection module is configured to detect PCI devices connected on the PCI bridge link, wherein the PCI devices include EP devices and / or PCI switch devices, and multiple EP devices are allowed to be connected on the PCI switch device; The first processing module is configured to determine the fault handling method of the PCI device based on the detected PCI device, and handle the fault of the PCI device according to the fault handling method. The first processing module includes: a first processing unit configured to, when the PCI device connected to the PCI bridge link is only the EP device, activate the fault handling function of the EP device, the fault handling function being a function set in the EP device for handling faults of the EP device; and a second processing unit configured to, when the PCI device connected to the PCI bridge link includes the PCI Switch device, set a fault threshold for the PCI Switch device and adjust the transmission rate of the PCI bridge link based on the fault threshold, the fault threshold representing a threshold for recording the number of faults of the PCI Switch device, and the fault handling method including the fault handling function and the method of setting the fault threshold of the PCI Switch device.
20. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method described in any one of claims 1 to 16.
21. A non-volatile computer-readable storage medium, characterized in that, The non-volatile computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the method described in any one of claims 1 to 16.
22. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method described in any one of claims 1 to 16.
Citation Information
Patent Citations
Fault processing method and device and server
CN111414268A
PCIe fault detection device and method, equipment and storage medium
CN113868051A
Fault equipment determination method and device, storage medium and electronic equipment
CN117499214A
Fault processing method and computing device
CN118051366A
Fault processing method, device and system for PCI (Peripheral Component Interconnect) equipment
CN118245269A
Cited By
Method for handling link failure, monitor, and server
CN122450728A