PCIe equipment loss retry method, device and system

By acquiring the power failure status information of PCIe devices and using the PCIe bus and I2C bus to scan the slot status, physical connection failures and protocol layer anomalies are distinguished. A warm start/cold start mechanism is adopted for retrying, which solves the problem of loss during PCIe device startup and improves the robustness and troubleshooting efficiency of the system.

CN122019241APending Publication Date: 2026-05-12SHANGHAI LINGHUA INTELLIGENT TECHNOLOGY CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-13
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

In existing technologies, PCIe devices have a small probability of being lost during startup, making it impossible to distinguish between physical connection failures and PCIe protocol layer anomalies. This can lead to false or missed retries, which in turn affects system availability and troubleshooting efficiency.

Method used

By acquiring the power failure status information of the PCIe device, it is determined whether it is in a warm start state. The PCIe bus and I2C bus are used to scan the slot status in parallel to distinguish between physical connection failures and protocol layer anomalies. A warm start/cold start coordination mechanism is adopted to retry the start-up and limit the number of retries to avoid invalid operations.

Benefits of technology

It improved troubleshooting efficiency, optimized equipment initialization success rate, enhanced system robustness and availability, and reduced system startup delay.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122019241A_ABST
    Figure CN122019241A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a PCIe (Peripheral Component Interconnect Express) equipment loss retry method. The method comprises the following steps: acquiring power supply fault state information of PCIe equipment; judging whether the power supply fault state information is warm start or not; if yes, the slot position of the PCIe equipment is scanned, and slot position state information is obtained; judging whether the slot position state information is first preset state information or not; and if yes, restarting the PCIe equipment. According to the method, the initialization success rate of the equipment can be optimized, and the robustness of the system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of equipment control, and more particularly to a PCIe device loss retry method, retry device, and retry system. Background Technology

[0002] PCIe (peripheral component interconnect express, high-speed serial computer expansion bus standard) devices in servers or embedded devices have a small probability of "disappearing" (i.e., device loss) during startup, which can affect system availability.

[0003] In existing technologies, if no PCIe presence signal is received from the target device or the initialization response times out, it is determined as a "device loss" retry mechanism. However, this retry mechanism cannot distinguish between "physical connection failure" and "PCIe protocol layer abnormality," which easily leads to false or missed retries, resulting in low troubleshooting efficiency and system startup delays. Summary of the Invention

[0004] Therefore, in order to overcome at least some of the defects and deficiencies of the prior art, this invention proposes a PCIe device loss retry method to improve system robustness.

[0005] On one hand, the PCIe device loss retry method proposed in this embodiment of the invention includes: obtaining power failure status information of the PCIe device; determining whether the power failure status information is warm start status information; if so, scanning the slots of the PCIe device based on the PCIe bus and I2C bus respectively to obtain slot status information; determining whether the slot status information is a first preset status information; if so, retrying to start the PCIe device.

[0006] In one embodiment of the present invention, the step of scanning the slots of the PCIe device based on the PCIe bus and the I2C bus respectively to obtain slot status information specifically includes: scanning the slots through the PCIe protocol to obtain PCIe status information; scanning the slots through the I2C protocol to obtain I2C status information; and summarizing the PCIe status information and the I2C status information to obtain the slot status information.

[0007] In one embodiment of the present invention, the first preset state information is that the PCIe state is not in place and the I2C state is in place.

[0008] In one embodiment of the present invention, the retry startup of the PCIe device specifically includes: determining whether the warm start counter is less than a preset number; wherein, the warm start counter is used to record the number of startups when the PCIe device is in a warm start state; if so, the warm start counter is incremented by 1, and the retry counter of the PCIe device is incremented by 1; wherein, the retry counter is the number of retry startups of the PCIe device; triggering the warm start state information of the PCIe device, and performing initial startup on the PCIe device.

[0009] In one embodiment of the present invention, the method further includes: if it is determined that the warm-start counter is greater than or equal to a preset number of times, then terminating the retry of the PCIe device and marking the PCIe device as abnormal.

[0010] In one embodiment of the present invention, the method further includes: when the power failure status information is determined to be cold start status information, clearing the power failure status information of the PCIe device, resetting the warm start counter, and performing initialization startup on the PCIe device.

[0011] On the other hand, an embodiment of the present invention proposes a PCIe device loss retry device, comprising: a first acquisition module, used to acquire power failure status information of the PCIe device; a first judgment module, used to determine whether the power failure status information is warm-start status information; a obtaining module, used to scan the slots of the PCIe device based on the PCIe bus and I2C bus respectively to obtain slot status information when the first judgment module determines that the power failure status information is warm-start status information; a second judgment module, used to determine whether the slot status information is a first preset status information; and a retry module, used to retry the start of the PCIe device when the second judgment module determines that the slot status information is the first preset status information.

[0012] In one embodiment of the present invention, the obtaining module includes: a PCIe status information obtaining unit, used to scan the slot through the PCIe protocol to obtain PCIe status information; an I2C status information obtaining unit, used to scan the slot through the I2C protocol to obtain I2C status information; and a slot status information obtaining unit, used to summarize the PCIe status information and the I2C status information to obtain the slot status information.

[0013] In one embodiment of the present invention, the device further includes: an initialization startup module, configured to clear the power failure status information of the PCIe device, reset the warm start counter, and perform initialization startup on the PCIe device when the first judgment module determines that the power failure status information is cold start status information.

[0014] In another aspect, embodiments of the present invention propose a PCIe device loss retry system, comprising: a memory and a processor connected to the memory, wherein the memory stores a computer program, and the processor executes a PCIe device loss retry method as described above when running the computer program.

[0015] This invention provides a PCIe device loss retry device, comprising: The first acquisition module detects power failure status information of PCIe devices; The processor, connected to the first acquisition module, determines whether the power failure status information is a warm-start status information. When the power failure status information is determined to be a warm-start status information, the processor scans the slots of the PCIe device based on the PCIe bus and I2C bus respectively to obtain the slot status information. Specifically, when the slot status information is determined to be the first preset status information, the processor retryes the PCIe device to start up. The first preset status information is that the PCIe status is not in place and the I2C status is in place.

[0016] In one embodiment of the present invention, the processor is a baseboard management controller for a server.

[0017] In one embodiment of the present invention, the processor monitors the power status via an I2C bus.

[0018] In one embodiment of the present invention, the processor monitors the power status via a power monitoring CPLD on the server motherboard.

[0019] In one embodiment of the present invention, the first acquisition module may be a voltage monitoring chip, used to detect power failure status information of PCIe devices.

[0020] In one embodiment of the present invention, the PCIe device is a GPU card or a storage controller.

[0021] In one embodiment of the present invention, the PCIe device loss retry device further includes a PCIe bus, a PCIe device and a PCIe slot connected to the PCIe bus, an I2C bus, and a management interface on the PCIe slot.

[0022] In one embodiment of the present invention, the processor reads the value of the power failure status bit by accessing the configuration register in the motherboard firmware in order to obtain the power failure status information of the PCIe device.

[0023] In one embodiment of the present invention, the processor reads the register via the I2C bus to obtain the "power failure flag" of the PCIe slot as "warm start status information".

[0024] In one embodiment of the present invention, the processor is a baseboard management controller of a server. When the slot of the PCIe device is scanned based on the PCIe bus and the I2C bus respectively, the baseboard management controller notifies the CPU to perform a PCIe scan on the slot. At the same time, the baseboard management controller sends a read request to the predefined I2C management address of the PCIe device through the I2C bus.

[0025] In one embodiment of the present invention, when scanning based on the I2C bus, the processor can perform the scanning through an I2C multiplexer or a processor I2C controller.

[0026] In one embodiment of the present invention, when the processor retryes the startup of the PCIe device, the processor performs a warm start on the PCIe device. The processor includes a warm start counter to record the number of startups when the PCIe device is in a warm start state.

[0027] As can be seen from the above, the technical features of the present invention can have the following beneficial effects: This application obtains the power failure status information of the PCIe device; determines whether the power failure status information is a warm start status information. If so, the slots of the PCIe device are scanned based on the PCIe bus and I2C bus respectively to obtain slot status information; determines whether the slot status information is a first preset status information; if so, the PCIe device is retried to start. This application first determines whether the power-on status in the fault status information of the PCIe device is a warm start, and uses a warm start / cold start coordination mechanism to first determine whether the device has a problem. When the device is in a warm start state, the slots of the device are scanned in parallel through the PCIe bus and I2C bus. By determining whether the PCIe and I2C are in place in the scanned structure, the retry mechanism is triggered, thereby improving the efficiency of fault diagnosis, optimizing the success rate of device initialization, and improving the robustness of the system. Attached Figure Description

[0028] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0029] Figure 1 This is a flowchart illustrating a PCIe device loss retry method provided in the first embodiment of the present invention.

[0030] Figure 2 for Figure 1 A flowchart illustrating the specific steps of step S103.

[0031] Figure 3 for Figure 1 A flowchart illustrating the specific steps of step S105.

[0032] Figure 4 for Figure 3 A flowchart illustrating the additional steps following step S1051.

[0033] Figure 5 for Figure 1 A flowchart illustrating the steps following step S102.

[0034] Figure 6 This is a schematic diagram of a PCIe device loss retry device provided in an embodiment of the present invention.

[0035] Figure 7 for Figure 6 The diagram shows the specific unit of the obtained module.

[0036] Figure 8 for Figure 6 The schematic diagram of the retry module shown is a detailed unit diagram. Figure 9 This is a schematic diagram of a PCIe device loss retry system provided in an embodiment of the present invention.

[0037] Figure 10 This is a schematic diagram of a storage medium module provided in an embodiment of the present invention. Detailed Implementation

[0038] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0039] In this embodiment, a PCIe device loss retry method can be applied to PCIe devices in servers and embedded devices. These PCIe devices can be, for example, various devices supporting both PCIe and I2C interfaces (such as GPU cards, storage controllers, etc.). The loss phenomenon refers to the server or embedded device system suddenly failing to detect the PCIe device during operation. This loss phenomenon is usually caused by one or more of the following reasons, including: 1. Physical connection problems: poor slot contact, oxidation / dirt on the gold fingers leading to signal transmission interruption, damage to the slot or graphics card itself. 2. Power supply problems: insufficient power supply can cause the device to restart or disconnect; poor power quality can cause system instability and loss. 3. Signal integrity problems: unstable transmission links can cause the system to attempt reconnection or disconnect directly. Cable / adapter card problems: using extension cables or adapter cards can introduce signal attenuation; poor-quality cables are prone to causing loss. 4. Heat dissipation problems: overheating of the device may trigger protection mechanisms, causing the system to fail to detect the device. 5. Driver / BIOS problems: driver conflicts or crashes can cause the device to enter an abnormal state and become unresponsive, among other reasons. When a loss occurs, a retry is necessary to determine the device malfunction. In one embodiment, the PCIe bus can output the PCIe protocol to the PCIe device slot, allowing feedback from the slot to determine if the PCIe device is present. However, this relies on software-level detection of the PCIe device's status and does not involve hardware-level judgment. In another embodiment, after system startup, if the device driver does not detect the PCIe device node (e.g., no corresponding device directory in / sys / bus / pci / devices / in a Linux system), a software-level retry can be triggered. During the retry, the software calls the PCIe configuration space interface to rescan the bus or restarts the device driver. The software-level judgment is based solely on whether the software recognizes the device driver loading, without verifying the physical connection status of the device (e.g., whether the hardware presence pin signal is read). If the physical connection of the device is normal but the driver adaptation is abnormal, the software retry will repeatedly execute invalid operations, failing to locate the true fault. Furthermore, this method cannot distinguish between physical connection failures and PCIe protocol layer anomalies, which can easily lead to false or missed retries, resulting in low troubleshooting efficiency and system startup delays.

[0040] like Figure 1 As shown, an embodiment of the present invention provides a PCIe device loss retry method, which includes, for example, the following steps: S101. Obtain the power failure status information of the PCIe device; S102. Determine whether the power fault status information is a warm start status information; S103. If so, the slots of the PCIe device are scanned based on the PCIe bus and the I2C bus respectively to obtain the slot status information. S104. Determine whether the slot status information is the first preset status information; S105. If so, the PCIe device will be restarted again.

[0041] In step S101, in one embodiment, after the system powers on, the power failure status bit can be read from the configuration register in the motherboard firmware (BIOS / UEFI) to obtain the power failure status information of the PCIe device. The power failure status bit is, for example, a binary flag bit, with values ​​including "1" and "0," used to record whether a power abnormality occurred during the previous system startup. When the power failure status bit is "1," it indicates that there was an abnormal power outage or power failure during the previous system startup, and the current startup is a startup after recovery from the abnormality. When the power failure status bit is "0," it indicates that the previous system shutdown was normal or the first power-on, and the current startup is a normal startup. The configuration register is a dedicated status storage unit preset by the system. Its address and access permissions are preset through firmware configuration to ensure the accuracy and security of reading the power failure status bit information and avoid misjudgment of the status due to external interference.

[0042] In step S102, based on the power fault status information obtained in step S101, the startup type is classified and corresponding operations are performed. If the power fault status bit is "1", the current startup type is determined to be a warm start, and the power fault status information is the warm start status information. In the warm start state, the system is not completely powered off and only a partial reset is required. Therefore, the system global initialization steps (such as firmware reinstallation or full bus reset) can be skipped directly, and the slot scanning stage can be entered to shorten the startup time.

[0043] When determining the cold start status, if the power fault status bit is "0", the current startup type is determined to be a cold start, and a full initialization process must be executed. This step distinguishes the startup type by the power fault status, avoiding redundant global initialization for temporarily abnormal devices and laying the foundation for efficient retrying later.

[0044] In step S103, the system determines that the received power fault status bit information is "1", indicating a warm start state. Then, the PCIe bus detection unit and I2C bus detection unit are activated to perform a parallel scan of all PCIe device slots in the system, synchronously collecting the device presence status of each slot and ultimately generating slot status information. In one embodiment, the PCIe device slots support 8 slots by default. In other embodiments, the number of PCIe device slots can be expanded to more than 8 slots. Parallel scanning avoids the queuing delay of single-bus scanning, ensuring that all slot detection is completed within a short time (e.g., within 500ms), adapting to multi-device server deployment scenarios.

[0045] In step S104, a retry is only performed when the physical connection of the device is normal but the PCIe communication is abnormal. The first preset status information may be, for example, that the PCIe status is not present and the I2C status is present. This step filters out the target devices that need to be retried by comparing the slot status information with the first preset status information, thus excluding physically faulty devices and normal devices.

[0046] In step S105, when the slot status information matches the first preset status, the hierarchical retry mechanism is triggered.

[0047] This application first determines whether the power-on status in the PCIe device's fault status information is a warm start. A coordinated warm start / cold start mechanism is used to pre-determine if the device has a problem. When the device is in a warm start state, the partial reset time can be reduced to a very short time (e.g., within 1 second), significantly shorter than a cold start, thus improving retry efficiency. Then, the device's slots are scanned in parallel using the PCIe and I2C buses. By determining whether the PCIe and I2C ports are present in the scanned structure, a retry mechanism is triggered, accurately distinguishing between physical faults and protocol layer anomalies. This improves fault diagnosis efficiency, optimizes device initialization success rate, and enhances system robustness.

[0048] Furthermore, such as Figure 2 As shown, step S103 may include: S1031. Scan the slot using the PCIe protocol to obtain PCIe status information; S1032. Scan the slot using the I2C protocol to obtain I2C status information; S1033. Summarize the PCIe status information and the I2C status information to obtain the slot status information.

[0049] Specifically, in step S1031, the slots can be scanned via the PCIe protocol. The process can be, for example, activating the PCIe protocol detection unit and sending a "device presence detection command" (e.g., PRSNT# signal interaction logic based on the PCIe Base Specification 5.0 protocol) to each PCIe device slot. During the detection process, a read request is sent to the PCIe configuration space corresponding to the slot (e.g., the device ID register at address 0x00). If a valid response (device ID matching a preset value) is received within a preset response timeout period (e.g., 100ms), the slot is determined to be "PCIe present". During anomaly detection, if no response is received within the preset response timeout period, the response times out, or the device ID is invalid, the slot is determined to be "PCIe absent". The PCIe status information, i.e., the "PCIe present / absent" identifier for each slot, reflects whether PCIe protocol layer communication is normal.

[0050] In step S1032, the I2C bus detection unit is activated. It can access the physical presence pin (preset I2C slave address, such as 0x50) of the device corresponding to each slot, for example, based on the I2C protocol (such as a subset of SMBus). During the detection process, a "physical presence query command" is sent to the target slot. If the device returns an ACK (acknowledgment) signal, it indicates that the device's physical pin contact is normal, and it is determined to be "I2C present". When an anomaly is detected, if no ACK signal is returned, the signal times out, or a NACK (invalid response) occurs, it is determined to be "I2C absent", indicating that the device may have a physical fault such as improper insertion or slot damage. The I2C status information, namely the "I2C present / absent" identifier for each slot, is used to verify the physical connection status of the device.

[0051] In step S1033, the PCIe status information and I2C status information of the same slot are combined to generate, for example, multiple possible slot status information. For example, PCIe in place + I2C in place (device normal, no processing required), PCIe in place + I2C not in place (physical fault, no retry required), PCIe not in place + I2C in place (protocol layer abnormality, retry required), PCIe not in place + I2C not in place (physical fault, no retry required).

[0052] This step can accurately distinguish fault types by summarizing the status of both buses, providing a basis for subsequent retry decisions and avoiding the misjudgment problems of traditional single-bus detection. However, it is not limited to this; other buses can also be used for detection, such as SMBus, JTAG, MDIO, SPI / eSPI, etc.

[0053] Furthermore, when the first preset status information is "the PCIe status is not in place, and the I2C status is in place," a collaborative verification of physical connection and protocol communication is performed. Here, "I2C status in place" indicates that the device's physical connection is normal (e.g., good pin contact, no slot damage), ruling out physical faults and providing a basis for retry recovery. "PCIe status not in place" indicates abnormal communication between the device and the system's PCIe protocol layer (e.g., link training failure, signal interference, protocol handshake timeout, etc.), which is a temporary recoverable fault and needs to be resolved through retry. If the slot status information matches the first preset status, it indicates that the loss is due to a temporary abnormality in the PCIe protocol layer, rather than physical damage, and performing a retry operation can effectively recover the device; if they do not match (e.g., a physical fault status), the retry is terminated to avoid wasting resources.

[0054] Furthermore, such as Figure 3 As shown, step S105 specifically includes: S1051. Determine whether the warm-start counter is less than a preset number. The warm-start counter is the number of times the PCIe device has been started when it is in a warm-start state. The preset number can be between 2 and 5 times, for example, 3 times. S1052. If so, the warm-start counter is incremented by 1, and the retry counter of the PCIe device is incremented by 1; wherein, the retry counter is the number of retry starts of the PCIe device; S1053. Trigger the warm start status information of the PCIe device and initialize the PCIe device.

[0055] Specifically, in step S1051, the current value of the system-preset Warm Boot Counter is read, and it is determined whether it is less than a preset number. The core function of the Warm Boot Counter is to "limit the number of retries and avoid infinite loops." Its definition and characteristics are as follows: Counting object: Only counts the number of retries when the PCIe device is in a warm boot state; cold boot does not affect its value; Initial value: Only resets to zero during cold boot, and maintains numerical continuity during warm boot. In one embodiment, the upper limit is set to, for example, 3 times, which is the optimal value based on extensive experimental verification. If it is less than, for example, 3 times, it may miss the opportunity to recover from temporary anomalies; if it is more than 3 times, it will significantly prolong the boot time.

[0056] In step S1052, if the counter has not reached its upper limit, the following operation may be performed: Counter Update: Increment the Warm Boot Counter value by 1 (e.g., from 0 to 1), and simultaneously update the value of the global configuration register Setup.PCIEretry (increment the global retry count by 1, e.g., from 0 to +). Setup.PCIEretry records the cumulative number of retries for this PCIe device and is stored in a non-volatile register for easy troubleshooting. The value of Setup.PCIEretry can be obtained, for example, by querying the retry history through the system log. This triggers step S1053. In step S1053, the warm start process of the PCIe device can be actively triggered, which specifically includes: performing a partial reset only on the currently lost PCIe device and the corresponding PCIe link (such as restarting the link training process of the PCIe root complex), and skipping the global system initialization (such as firmware reinstallation, other device reset, etc.). Re-execute PCIe device enumeration (read device ID, configuration space) and basic driver loading to ensure that the device and system establish normal communication. Since the warm start time is controlled within 1 second, the efficiency is improved by more than 80% compared to the cold start (which takes 5-10 seconds).

[0057] In steps S1051-S1053 of this embodiment, a local warm start can be used to replace a global cold start, which can significantly shorten the startup delay while ensuring the effectiveness of retry.

[0058] Furthermore, such as Figure 4 As shown, after step S1051, the method further includes: S1054. If it is determined that the warm-start counter has been greater than or equal to 3 times, then the retry of the PCIe device is terminated and the PCIe device is marked as abnormal.

[0059] This implementation describes the handling logic after the maximum number of retries has been reached, to prevent wireless retries from causing the system to freeze.

[0060] Specifically, when it is determined that the warm-up counter is greater than or equal to a preset number of times, the following operations are performed: Terminate the retry process and immediately stop subsequent retry operations on the PCIe device to avoid occupying bus resources and prolonging system startup time. Mark the device as abnormal and record its abnormal information in the system log (such as / var / log / messages in Linux systems or Event Viewer in Windows), including: Basic information: the slot number of the abnormal device (e.g., Slot3), device model (e.g., NVIDIA A100 GPU); Retry history: the final value of the warm start counter (3 times), the number of global retries (Setup.PCIEretry value), and the slot status information for each retry; Abnormal identifier: mark "PCIe protocol layer unrecoverable abnormality" to prompt the user to check the hardware links (e.g., PCIe slots, cables).

[0061] This step, through forced termination and detailed logging, balances system startup efficiency with ease of subsequent troubleshooting.

[0062] Furthermore, such as Figure 5 As shown, after step S102, the method further includes: S106. When the power failure status information is determined to be cold start status information, the power failure status information of the PCIe device is cleared, the warm start counter is reset, and the PCIe device is initialized and started.

[0063] This step determines the current startup type as a cold start if the power failure information, i.e., the power failure status bit, is "0". This ensures the integrity of initialization and the accuracy of the counter in cold start scenarios. The specific steps are as follows: When the power failure status information is determined to be cold start status information (i.e., the power failure status bit is "0"), perform the following operations: Clear the power failure status information and reset the power failure status bit to "0" (to ensure accurate status judgment during subsequent startup and avoid residual abnormal indicators). Reset the warm boot counter to zero, because a cold boot is a "full initialization", and the number of possible warm boot retries needs to be counted again. Performing a cold start initialization process, which involves a full initialization procedure for the PCIe device, may include the following steps: Restart the PCIe bus and clear the configuration space cache of all PCIe devices; Perform a complete link rate negotiation (such as upgrading from Gen1 to Gen4), lane binding, and other processes for all PCIe slots; Read the complete configuration information of each PCIe device (such as device ID, vendor ID, BAR space) and load the corresponding underlying driver; The system checks the stability of the communication link for each device. If an anomaly is detected, it is marked as "cold start initialization failed," prompting the user to check the hardware.

[0064] This step can perform a full initialization via cold start to improve device configuration issues during initial startup or recovery from severe anomalies, providing a stable foundation for subsequent warm start retries.

[0065] This embodiment can eliminate physical connection faults by utilizing the I2C presence signal, accurately pinpoint PCIe protocol layer anomalies, avoid false retries, and accurately distinguish between physical faults and protocol layer anomalies through dual-bus collaborative detection. Secondly, this embodiment performs a cold start, for example, initializing only once, and then uses a warm start to achieve a rapid partial reset after an anomaly, reducing system latency. Furthermore, this embodiment uses a counter to forcibly limit the number of retries to less than a preset number, balancing fault tolerance and startup reliability.

[0066] like Figure 6 As shown, a second embodiment of the present invention provides a PCIe device loss retry device 300. As... Figure 6 As shown, the PCIe device loss retry device 300 includes, for example, a first acquisition module 301, a first judgment module 302, an acquisition module 303, a second judgment module 304, and a retry module 305.

[0067] The first acquisition module 301 is used to acquire power failure status information of the PCIe device. The first judgment module 302 is used to determine whether the power failure status information is a warm-start status information. The obtaining module 303 is used to scan the slots of the PCIe device based on the PCIe bus and I2C bus respectively to obtain slot status information when the first judgment module 302 determines that the power failure status information is a warm-start status information. The second judgment module 304 is used to determine whether the slot status information is a first preset status information. The retry module 305 is used to retry the startup of the PCIe device when the second judgment module 304 determines that the slot status information is the first preset status information.

[0068] Furthermore, such as Figure 6 As shown, the PCIe device loss retry device 300 further includes, for example, an initialization startup module 306, which is used to clear the power failure status information of the PCIe device, reset the warm start counter, and perform initialization startup on the PCIe device when the first judgment module determines that the power failure status information is cold start status information.

[0069] like Figure 7 As shown, the obtaining module 303 in this embodiment specifically includes: PCIe status information obtaining unit 3031, I2C status information obtaining unit 3032 and slot status information obtaining unit 3033.

[0070] Specifically, the PCIe status information obtaining unit 3031 is used to scan the slot via the PCIe protocol to obtain PCIe status information. The I2C status information obtaining unit 3032 is used to scan the slot via the I2C protocol to obtain I2C status information. The slot status information obtaining unit 3033 is used to summarize the PCIe status information and the I2C status information to obtain the slot status information.

[0071] Furthermore, such as Figure 8 As shown, the retry module 305 specifically includes: a first judgment unit 3051, a counter increment unit 3052, a warm-start trigger unit 3053, and a retry termination unit 3054.

[0072] Specifically, the first judgment unit 3051 is used to determine whether the warm start counter is less than a preset number, wherein the warm start counter is the number of times the PCIe device is started when it is in the warm start state information.

[0073] The counter increment unit 3052 is used to increment the warm-start counter by 1 and the retry counter of the PCIe device by 1 when the first judgment unit 3051 determines that the warm-start counter is less than a preset number (e.g., 3 times); wherein, the retry counter is the number of retry startups of the PCIe device. The warm-start triggering unit 3053 is used to trigger the warm-start status information of the PCIe device and perform initialization startup on the PCIe device.

[0074] The retry termination unit 3054 is used to terminate the retry of the PCIe device and mark the PCIe device as abnormal when the first judgment unit 3051 determines that the warm-start counter is greater than or equal to a preset number (e.g., 3 times). It should be noted that the PCIe device loss retry method implemented by the PCIe device loss retry device 300 provided in this embodiment is as described in the previous embodiment, and therefore will not be described in detail here. Optionally, the various modules and other operations or functions in another embodiment are respectively for implementing the methods of the foregoing embodiments of the present invention, and the beneficial effects are the same as those in the first embodiment. For the sake of brevity, they will not be described in detail here.

[0075] like Figure 9As shown, an embodiment of the present invention provides a PCIe device loss retry device 400. The retry device or system 400 includes, for example, a memory 410 and a processor 420 connected to the memory 410. The memory 410 may be, for example, a non-volatile memory, on which a computer program 411 is stored. The processor 420 may be, for example, an embedded processor, or a BMC (Baseboard Management Controller), CPU (Central Processing Unit), CPLD (Complex Programmable Logic Device), or other processing unit with logic operation and program execution capabilities in a server. When the processor 420 runs the computer program 411, it executes a PCIe device loss retry method according to the aforementioned embodiment. In one embodiment, the BMC can monitor the power status via the I2C bus and, in cooperation with the CPU (i.e., the PCIe device), perform a scan of the PCIe bus. The BMC and the CPU can communicate via the LPC bus or shared memory to collaboratively complete the entire retry process. In some embodiments, the PCIe device loss retry device (or system) 400 may further include a PCIe bus, PCIe devices and PCIe slots connected to the PCIe bus, an I2C bus, a BMC connector, and a management interface on the PCIe slot.

[0076] In one embodiment, this application also provides a server that includes the aforementioned PCIe device loss retry device 400.

[0077] In one embodiment, specifically illustrated, a server employing the device 400 of this application is equipped with a PCIe device (e.g., a high-performance GPU card). When a PCIe device is lost, for example, a sudden, minute voltage drop occurs in the server's power module. Although this voltage drop is very short (milliseconds), it may cause the GPU card's main power supply (12V) to momentarily drop below its operating threshold, resulting in a PCIe link disconnection. However, the GPU card's auxiliary power supply and I2C management interface still function normally. At this time, the voltage drop can be detected by the power monitoring CPLD on the server motherboard, and a "power failure flag" is set in the internal register for the corresponding PCIe slot, with the status "warm start status information". Then, the server's processor 420 (e.g., BMC) can perform periodic checks (e.g., every 10 seconds). The processor 420 (e.g., BMC) can read the CPLD register via the I2C bus to obtain the "power failure flag" of the PCIe slot, which is set to "warm start status information".

[0078] Next, a dual-bus detection process can be initiated via processor 420 (e.g., BMC). For example, the BMC instructs the CPU (via system management interrupt or shared memory) to perform a PCIe scan on the slot. After the CPU performs the scan, it returns the result: no device was found (PCIe status is not present). Simultaneously, the BMC sends a read request to the GPU card's predefined I2C management address via the I2C bus. The GPU card responds normally and returns its device ID and temperature information (I2C status is present). Then, the processor 420 (e.g., BMC) can combine the scan results into slot status information: (PCIe = not present, I2C = present). This information perfectly matches the predefined "first preset status information". The BMC determines that the current value of the "warm-start counter" associated with the GPU card is 1 (assuming no such failure has occurred before), and then the BMC performs the following operations: increments the warm-start counter to 2 and records a retry event.

[0079] When a warm-start operation is triggered on the PCIe slot, specific actions may include: sending a command to the power control unit via the I2C bus to perform a "power-down-power-on" cycle reset of the GPU card's main power supply, or directly resetting the link via the PCIe hot reset signal. After the reset operation is complete, the CPU re-enumerates the PCIe bus. After the PCIe bus is re-enumerated, the CPU successfully recognizes the GPU card and reloads the driver. The GPU card resumes operation, and the computing task can continue. If the BMC finds the device status to be normal through subsequent inspections, it clears the power failure flag but retains the warm-start counter value (e.g., 2 times) in case of a subsequent failure.

[0080] Through the above operation process, detection and retry can be performed automatically without manual intervention, and server services can continue uninterrupted. Therefore, by using both the PCIe and I2C buses to perform dual detection of device status, the specific loss problem of device logic loss caused by power failure can be accurately identified, effectively avoiding misjudgments caused by relying solely on single bus detection (such as misjudging a physical disconnection as a power failure, or misjudging a brief signal jitter as a serious fault). After accurately determining the fault type, the system can automatically execute a warm-start retry process without manual plugging / unplugging of devices or restarting the server. For scenarios such as data centers that require uninterrupted operation, this can greatly improve operational efficiency and system availability. Furthermore, this application limits the number of retries (e.g., a maximum of 3 times), effectively avoiding infinite restart loops caused by permanent hardware damage, preventing system resources from being exhausted by invalid operations, and providing administrators with clear fault location information. Moreover, different processing strategies can be adopted according to the different power failure status information (cold / warm). During a cold start, all states are cleared and a completely new initialization is performed; during a warm start, the state information is retained for targeted recovery. This design ensures the correctness of the system state machine and the rationality of the initialization process. Furthermore, by recording detailed counter values ​​(such as the number of retries) and device status (such as anomaly flags), it provides maintenance personnel with fault diagnosis information, helping to quickly pinpoint whether the problem is an intermittent power supply issue or a hardware failure.

[0081] like Figure 10 As shown, one embodiment of the present invention provides a storage medium, such as a computer-readable storage medium 500. The computer-readable storage medium 500 is, for example, a non-volatile memory, such as a magnetic medium (e.g., hard disk, floppy disk, and magnetic tape), an optical medium (e.g., CD-ROM and DVD), a magneto-optical medium (e.g., optical disc), and a hardware device specifically configured for storing and executing computer-executable instructions (e.g., read-only memory (ROM), random access memory (RAM), flash memory, etc.). Computer-executable instructions 510 are stored on the computer-readable storage medium 500. The computer-readable storage medium 500 can be executed by one or more processors or processing devices to implement a PCIe device loss retry method as described in the first embodiment above.

[0082] The methods, systems, and / or servers provided in this application can be applied to, for example, network platform servers (e.g., 5G network servers), data centers (e.g., large-scale cloud data centers), databases (e.g., financial database servers), AI computing servers (e.g., edge computing platform servers), servers for monitoring and data acquisition systems, traffic control system servers, smart grid platform servers, distributed control system (DCS) operator stations, artificial intelligence and high-performance computing (HPC) servers, big data analytics servers, network security platform servers, communication and collaboration platform servers, industrial control computers, medical systems / computers, industrial tablet computers, automated testing systems, drones, robots, general-purpose humanoid robots, AOI platform servers, embedded computing platforms, industrial IoT, vehicle networking, fanless minicomputers, or other suitable types of servers.

[0083] Furthermore, it is understood that the foregoing embodiments are merely illustrative examples of the present invention. Provided that the technical features do not conflict, the structure is not contradictory, and the purpose of the invention is not violated, the technical solutions of the various embodiments can be arbitrarily combined and used.

[0084] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for retrying a lost PCIe device, characterized in that, include: Obtain power failure status information for PCIe devices; Determine whether the power failure status information is a warm-start status information; If so, the slots of the PCIe device are scanned based on the PCIe bus and I2C bus respectively to obtain the slot status information; Determine whether the slot status information is the first preset status information; If so, the PCIe device will be restarted.

2. The method according to claim 1, characterized in that, The step of scanning the slots of the PCIe device using both the PCIe bus and the I2C bus to obtain slot status information specifically includes: The slot is scanned using the PCIe protocol to obtain PCIe status information; The slot is scanned using the I2C protocol to obtain I2C status information; The slot status information is obtained by summarizing the PCIe status information and the I2C status information.

3. The method according to claim 2, characterized in that, The first preset status information is that the PCIe status is not in place and the I2C status is in place.

4. The method according to claim 3, characterized in that, The retry startup of the PCIe device specifically includes: Determine whether the warm-start counter is less than a preset number. The warm-start counter is used to record the number of times the PCIe device is in a warm-start state. If so, the warm-start counter is incremented by 1, and the retry counter of the PCIe device is incremented by 1; wherein, the retry counter is the number of retry startups of the PCIe device; The warm start status information of the PCIe device is triggered, and the PCIe device is initialized and started.

5. The method according to claim 4, characterized in that, The method further includes: If the warm-start counter is determined to be greater than or equal to the preset number of times, then the retry for the PCIe device is terminated and the PCIe device is marked as abnormal.

6. The method according to claim 4, characterized in that, The method further includes: When the power failure status information is determined to be cold start status information, the power failure status information of the PCIe device is cleared, the warm start counter is reset, and the PCIe device is initialized and started.

7. A PCIe device loss retry device, characterized in that, include: The first acquisition module is used to acquire power failure status information of PCIe devices; The first judgment module is used to determine whether the power failure status information is a warm start status information; The module is used to scan the slots of the PCIe device based on the PCIe bus and the I2C bus respectively when the first judgment module determines that the power failure status information is a warm start status information, and obtain the slot status information. The second judgment module is used to determine whether the slot status information is the first preset status information; The retry module is used to retry the startup of the PCIe device when the second judgment module determines that the slot status information is the first preset status information.

8. The PCIe device loss retry device according to claim 7, characterized in that, The obtained module includes: The PCIe status information acquisition unit is used to scan the slot through the PCIe protocol to obtain PCIe status information. The I2C status information obtaining unit is used to scan the slot through the I2C protocol to obtain I2C status information; The slot status information obtaining unit is used to summarize the PCIe status information and the I2C status information to obtain the slot status information.

9. The PCIe device loss retry device according to claim 7, characterized in that, The device further includes: The initialization startup module is used to clear the power failure status information of the PCIe device, reset the warm start counter, and perform initialization startup on the PCIe device when the first judgment module determines that the power failure status information is cold start status information.

10. A PCIe device loss retry system, characterized in that, include: A memory and a processor connected to the memory, the memory storing a computer program, the processor executing the computer program to perform a PCIe device loss retry method as described in any one of claims 1 to 6.