Link recovery methods and computing devices

By introducing a tiered decision-making mechanism into the PCIe switching chip, the recovery operation level is selected based on the link status and recovery record, which solves the problem of service interruption caused by PCIe link instability and achieves the maintenance of service continuity and link reliability while ensuring system stability.

CN122496543APending Publication Date: 2026-07-31XFUSION DIGITAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XFUSION DIGITAL TECH CO LTD
Filing Date
2026-04-28
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively restore the stability of PCIe links while avoiding severe business interruptions, especially when link instability is caused by factors such as changes in environmental temperature and humidity, device aging, or poor contact. In such cases, the technologies fail to maintain business continuity while ensuring system stability.

Method used

By implementing a tiered decision-making mechanism based on link status and current recovery status in the PCIe switching chip, link instability events can be identified, and appropriate recovery operation levels can be selected according to the event type and historical recovery records. These operations include gradual speed reduction, clearing equalization parameters, and adjusting transmission parameters, thereby gradually restoring link stability.

Benefits of technology

While ensuring system stability, the impact of link maintenance operations on services is minimized, and system downtime caused by repeated link reconstruction or error accumulation is avoided, thereby improving the reliability and service continuity of PCIe switching chips.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122496543A_ABST
    Figure CN122496543A_ABST
Patent Text Reader

Abstract

This application relates to a link recovery method and computing device. The method includes: acquiring the link status of a high-speed PCIe port interconnecting peripheral components; when a link instability event is determined based on the link status, determining the level of the link recovery operation to be performed according to the link instability event and the current recovery status of the PCIe port, wherein the current recovery status is determined based on the record of link recovery operations already performed on the PCIe port, and different levels represent different degrees of intervention in the link; and performing a link recovery operation corresponding to the determined level. This can minimize the impact of link maintenance operations on services while ensuring system stability, enabling the link to maintain connection in the lowest available state, avoiding system downtime caused by repeated link reconstruction or error accumulation, and improving the reliability and service continuity of the PCIe switching chip.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communication technology, and in particular to link recovery methods and computing devices. Background Technology

[0002] PCIe (Peripheral Component Interconnect Express) switches, as a type of switching chip used to expand the connectivity of the PCIe bus, can virtualize multiple downlink ports on a small number of host-side PCIe lanes to achieve data routing and forwarding. They are widely used in systems that require connecting multiple PCIe devices. With the PCIe protocol speed increasing to Gen3 and above, the signal integrity requirements of the physical link are becoming increasingly stringent, making link equalization technology crucial for maintaining a low bit error rate. However, in practical applications, factors such as changes in environmental temperature and humidity, device aging, and poor contact can all affect the stability of high-speed SerDes (Serializer / Deserializer), leading to correctable errors (CE) or uncorrectable errors (UCE) in the link. This can result in link deceleration, repeated rebuilds, or even system crashes, posing challenges to service continuity.

[0003] To address the aforementioned link instability issues, relevant technologies primarily employ two approaches: First, passively recording link errors and issuing alarms, allowing the system to handle or ignore them automatically. This approach lacks effective intervention in link status, potentially allowing services to continue operating with errors until a serious failure occurs. Second, utilizing the DPC (downstream port containment) feature in the PCIe specification to directly disable or isolate the faulty port upon detecting a serious error. While this prevents error propagation, it immediately interrupts all services on that port, significantly impacting customers. Therefore, these technologies struggle to effectively restore link stability while avoiding severe service interruptions. Summary of the Invention

[0004] This application provides a link recovery method and computing device that can improve the reliability and service continuity of PCIe switching chips.

[0005] According to a first aspect of the embodiments of this application, a link recovery method is provided, the method comprising:

[0006] Obtain the link status of high-speed PCIe ports interconnecting peripheral components; When a link instability event is determined to exist in a link based on the link status, the level of the link recovery operation to be performed is determined according to the link instability event and the current recovery status of the PCIe port; wherein, the current recovery status is determined based on the record of link recovery operations already performed on the PCIe port, and different levels represent different degrees of intervention in the link; Perform the link recovery operation corresponding to the determined level.

[0007] This application implements a tiered decision-making mechanism based on link status and current recovery status in the PCIe switching chip. This mechanism executes corresponding recovery operations according to the link's instability level to achieve link recovery. This minimizes the impact of link maintenance operations on services while ensuring system stability, allowing the link to maintain connection in the lowest available state. It avoids system downtime caused by repeated link reconstruction or error accumulation, thus improving the reliability and service continuity of the PCIe switching chip.

[0008] In one possible implementation, the aforementioned link instability events include at least one of the following: Correctable error CE generated by PCIe port; Uncorrectable errors (UCE) generated by PCIe ports; The frequency with which the PCIe port's Link Training and State Machine (LTSSM) enters the recovery state exceeds a first preset threshold. The frequency of LTSSM-triggered link retraining events on the PCIe port exceeds the second preset threshold.

[0009] This embodiment uses abnormal CE, UCE, and LTSSM states (frequent entry into recovery state or triggering retraining events) as the criteria for determining link instability events. This allows for comprehensive and accurate identification of different stages of link degradation, from slight signal quality degradation (increased CE) to the occurrence of serious errors (UCE), and finally to the link falling into a dead loop of repeated negotiations (LTSSM anomaly). Based on these precise triggering conditions, the PCIe switching chip can promptly initiate the corresponding step-by-step recovery process, intervening before the link completely fails. This avoids system downtime due to error accumulation and prevents overreaction and service interruption while the link is still maintainable. Thus, while ensuring system stability, it maximizes service continuity.

[0010] In one possible implementation, the method further includes: When the link instability event is detected, the recovery record of the PCIe port is obtained. The recovery record includes the level of the link recovery operation performed on the PCIe port, the execution time, and the link response result after the execution. Based on the recovery record, determine the current recovery status of the PCIe port.

[0011] This embodiment introduces a method for acquiring and analyzing recovery records. By obtaining historical recovery behavior of PCIe ports, link recovery decisions no longer rely on a single instantaneous state but are based on complete, time-series historical information. This avoids repeatedly executing the same level of invalid operations within a short period and prevents the link from oscillating repeatedly between deceleration and acceleration. Furthermore, by recording the link response results after each operation, the current degree of link degradation can be accurately assessed, thereby selecting the most appropriate next-level operation and significantly improving the link recovery success rate of the PCIe switching chip.

[0012] In one possible implementation, when the aforementioned level is Level 1, a link recovery operation corresponding to the determined level is performed, including: The maximum link negotiation rate supported by the PCIe port is gradually reduced; Each reduction triggers a link renegotiation until no link instability event occurs, or the maximum link negotiation rate has been reduced to the PCIe Gen1 rate.

[0013] This embodiment achieves minimal intervention-based recovery of unstable links. This level of operation attempts to circumvent signal degradation caused by environmental changes by gradually reducing the link negotiation rate and triggering renegotiation, without affecting core physical layer parameters. Therefore, it minimizes link disturbance and achieves the fastest recovery speed. When the link can operate stably at a lower rate, services can continue to run without interruption, sacrificing only some bandwidth performance and avoiding service outages caused by complete link failure.

[0014] In one possible implementation, when the aforementioned level is level two, a link recovery operation corresponding to the determined level is performed, including: Clear the link balancing parameters of the PCIe port physical layer and trigger link renegotiation based on the PCIe Gen1 rate. After successfully establishing the link, gradually increase the maximum negotiation rate supported by the PCIe port; After each lift, determine whether a link instability event has occurred. If no link instability event occurs, continue to increase the maximum negotiation rate until the preset maximum attempt rate is reached; When a link instability event occurs, stop increasing the maximum negotiation rate and revert the maximum negotiation rate back to the previous stable rate.

[0015] This embodiment implements a deep recovery mechanism for unstable links. Building upon the ineffectiveness of the first-level speed reduction, this level of operation further removes potentially faulty equalization parameters from the physical layer, enabling the link to break free from its old processing mode and relearn and adapt to the current physical environment starting at Gen1 speed. This solves the problem of existing equalization parameters becoming faulty due to environmental changes, which cannot be recovered by simple speed reduction. By rebuilding the link from Gen1, a minimum level of connectivity is ensured even under the worst conditions, guaranteeing service continuity.

[0016] In one possible implementation, when the aforementioned level is level three, a link recovery operation corresponding to the determined level is performed, including: Adjust the physical layer transmit parameters of the PCIe port; Link renegotiation is triggered based on PCIe Gen3 rate; This embodiment, after the second-level parameter clearing proved ineffective, proactively adjusts the physical layer transmission parameters to match the severely degraded physical link. This addresses the issue of auto-negotiation and auto-learning mechanisms failing in extreme environments, providing another opportunity for link recovery through proactive intervention.

[0017] In one possible implementation, determining the level of the link recovery operation to be performed based on the link instability event and the current recovery status of the PCIe port includes: If the current recovery status of the PCIe port indicates that no link recovery operation of any level has been performed, the level of the link recovery operation to be performed is determined to be Level 1. When the current recovery status of the PCIe port indicates that a first-level link recovery operation has been performed but a second-level link recovery operation has not been performed, the level of the link recovery operation to be performed is determined to be a second-level link. If the current recovery status of the PCIe port indicates that a Level 2 link recovery operation has been performed but a Level 3 link recovery operation has not been performed, then the level of the link recovery operation to be performed is determined to be Level 3.

[0018] This embodiment can progressively select more in-depth intervention methods based on the persistence of link instability events and the levels of recovery operations already performed. This ensures that each intervention is a natural escalation after the previous milder methods have failed, avoiding excessive impact on services caused by escalating operations. Simultaneously, by maintaining the current link state when no link instability events are detected, unnecessary link oscillations and performance losses are avoided, significantly improving the PCIe switching chip's adaptability and operational efficiency in complex environments.

[0019] In one possible implementation, the method further includes: Maintain the current link state of the PCIe port when no link instability event is detected.

[0020] In this embodiment, if no link instability events are detected during the monitoring period, it indicates that the current link is operating stably, and no intervention is required.

[0021] In one possible implementation, the method further includes: After performing the above link recovery operation, if no link instability event occurs within the preset time window, the current recovery status of the PCIe port will be reset to the state where no link recovery operation of any level has been performed.

[0022] This embodiment introduces a recovery state reset mechanism. When the link remains stable for a long time, the recovery state can be rolled back to the initial state. This allows the system to start from the mildest first-level operation (slowing down) and retry when an unstable event occurs again in the future, avoiding unnecessary advanced interventions caused by historical records.

[0023] According to a second aspect of the embodiments of this application, a link restoration apparatus is provided, the apparatus comprising: The link status acquisition module is used to acquire the link status of the PCIe port; The level determination module is used to determine the level of the link recovery operation to be performed when the link instability event is determined to exist in the link based on the link status, according to the link instability event and the current recovery status of the PCIe port; wherein, the current recovery status is determined based on the record of the link recovery operations already performed on the PCIe port, and different levels represent different degrees of intervention in the link; The recovery operation execution module is used to perform link recovery operations corresponding to the determined level.

[0024] According to a third aspect of the embodiments of this application, a computing device is provided. The computing device includes: a memory and a PCIe switching chip, the memory storing a computer program, and the PCIe switching chip executing the program to implement the method as described above.

[0025] According to a fourth aspect of the embodiments of this application, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a PCIe switching chip, implements the methods described in the embodiments of this application.

[0026] According to a fifth aspect of the embodiments of this application, a computer program product is provided, including a computer program that, when executed by a PCIe switching chip, implements the methods described above in the embodiments of this application. Attached Figure Description

[0027] More details, features, and advantages of embodiments of the present application are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which: Figure 1 A schematic diagram of a system architecture for PCIe link recovery provided as an exemplary embodiment of this application; Figure 2 A flowchart of a link recovery method provided as an exemplary embodiment of this application; Figure 3 A flowchart of a link recovery method provided as yet another exemplary embodiment of this application; Figure 4 for Figure 2 Flowchart of step S230; Figure 5 for Figure 2 Another flowchart for step S230; Figure 6 A flowchart of a link recovery method provided as yet another exemplary embodiment of this application; Figure 7 for Figure 2 Flowchart of step S220; Figure 8 A schematic block diagram of the functional modules of a link recovery device provided in an exemplary embodiment of this application; Figure 9 A structural block diagram of a computing device provided for an exemplary embodiment of this application. Detailed Implementation

[0028] Embodiments of this application will now be described in more detail with reference to the accompanying drawings. While some embodiments of this application are shown in the drawings, it should be understood that embodiments of this application can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the embodiments of this application. It should be understood that the accompanying drawings and embodiments of this application are for illustrative purposes only and are not intended to limit the scope of protection of this application.

[0029] It should be understood that the various steps described in the method implementation of this application may be performed in different orders and / or in parallel. Furthermore, the method implementation may include additional steps and / or omit the steps shown. The scope of this application is not limited in this respect.

[0030] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the following description. It should be noted that the concepts of "first", "second", etc., mentioned in the embodiments of this application are only used to distinguish different devices, modules, or units, and are not used to limit the order of functions performed by these devices, modules, or units or their interdependencies.

[0031] It should be noted that the terms "one" and "more" mentioned in the embodiments of this application are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0032] The names of the messages or information exchanged between multiple devices in the embodiments of this application are for illustrative purposes only and are not intended to limit the scope of these messages or information.

[0033] Figure 1 This is a schematic diagram of a system architecture for PCIe link recovery provided in an embodiment of this application. Figure 1 As shown, the system architecture includes a server 10, which internally contains a CPU (central processing unit) 11, a PCIe switch 12, and multiple devices connected to the PCIe switch 12. Figure 1 The arrows in the text only indicate the direction of the downlink. Figure 1 The "Link" in this context refers to, for example, a PCIe link. A PCIeSwitch is a hardware switching chip used to expand and manage PCIe bus connections, and can play a role in data routing, bandwidth allocation, and interconnection of multiple devices within a system.

[0034] For example: CPU11, as the main control unit of server 10, has a built-in PCIe controller and is connected to the uplink port of PCIe switching chip 12 via a PCIe link. CPU11 is responsible for initiating PCIe transactions, sending data access requests to downstream devices, and processing the returned responses.

[0035] The PCIe Switch 12 contains multiple PCIe ports, one of which is configured as an uplink port connected to the CPU 11, while the remaining ports are configured as downlink ports connected to different devices. The PCIe switch chip 12 integrates firmware that executes link recovery strategies, monitors the link status of each port in real time, and performs corresponding link recovery operations based on the current recovery status when a link instability event is detected.

[0036] In this embodiment, the multiple devices connected to the PCIe Switch 12 may include, but are not limited to: NVMe (non-volatile memory express) 13 and NVMe 14 can be connected to the downstream port of PCIe Switch 12 via a PCIe link. NVMe 13 and NVMe 14 can be solid-state drives, serving as high-speed storage devices for persistent data storage and retrieval.

[0037] The GPU (graphics processing unit) 15 is connected to the downlink port of the PCIeSwitch 12 via a PCIe link to accelerate graphics processing or general computing tasks.

[0038] NetCard 16: Connects to the downlink port of PCIe Switch 12 via a PCIe link to enable communication between server 10 and external networks.

[0039] CPU11 establishes a communication connection with the uplink port of PCIe Switch 12 via a PCIe link; PCIe Switch 12 forwards the PCIe transaction packets sent by CPU11 to the downlink port of the corresponding target device through its internal routing and forwarding mechanism; each device (such as NVMe13, NVMe14, GPU15, NetCard 16) connects to the downlink port of PCIe Switch 12 via its own PCIe link to receive and respond to transaction requests from CPU11.

[0040] In the above system architecture, each port of the PCIe Switch 12 independently maintains its own PCIe link. When any link experiences a correctable error (CE), an uncorrectable error (UCE), or repeated link reconstruction due to environmental changes (such as temperature and humidity changes, device aging) or physical contact problems, the PCIe Switch 12 will initiate a link recovery process. Through operations such as step-by-step speed reduction, clearing equalization parameters for renegotiation, and adjusting transmission parameters, it will restore and maintain a stable link connection as much as possible, thereby ensuring the reliable operation of the entire server 10.

[0041] Based on the above embodiments, this application also provides a link recovery method, which can be implemented by a PCIe switching chip (such as...). Figure 1 The PCIe Switch 12 shown is executed. Figure 2 As shown, the method may include the following steps: In step S210, the link status of the PCIe port is obtained.

[0042] Specifically, the PCIe switching chip monitors the link status of each PCIe port in real time through periodic polling or interrupt triggering. This link status can include: the current operating speed of the link, the negotiation capability of the link (such as the maximum supported speed), the state of the LTSSM (Link Training and Status State Machine), and various error messages generated by the link.

[0043] In step S220, when it is determined that a link instability event exists in the link based on the link status, the level of the link recovery operation to be performed is determined according to the link instability event and the current recovery status of the PCIe port. The current recovery status is determined based on the record of link recovery operations already performed on the PCIe port, and different levels represent different degrees of intervention in the link.

[0044] In this embodiment, the link status obtained in step S210 is analyzed to determine whether there is a link instability event.

[0045] For example, when determining the existence of link instability events based on various error messages of link parameters, the registers of the PCIe port can be read to count the occurrence of CE and UCE. If multiple CEs (e.g., more than 10) are detected consecutively within a preset time window (e.g., 1 second), or any UCE occurs, a link instability event can be confirmed. The occurrence of CE usually indicates a deterioration in link signal quality, but it is still recoverable; the occurrence of UCE indicates that the link can no longer guarantee correct data transmission, which is a serious link instability event.

[0046] When determining whether a link is unstable based on the LTSSM state, the LTSSM state transitions of each port can be monitored in real time. For example, if the LTSSM frequently enters the recovery state or frequently triggers link retraining events, it indicates that there is a link instability event even without obvious CE / UCE.

[0047] When determining whether a link instability event has occurred by checking the current operating rate of the link, the current operating rate can be compared with the maximum negotiated rate supported by the port. If the current operating rate has been reduced to the minimum (Gen1), but the link still frequently experiences the aforementioned CE, UCE, or LTSSM anomalies, it can be further confirmed that a link instability event exists. Otherwise, if the current operating rate is lower than the port's capacity and the link is stable, it indicates normal speed reduction and does not constitute a link instability event.

[0048] Therefore, it should be noted that the determination of link instability events in the embodiments can be based on a single link state trigger or a comprehensive determination based on a combination of multiple link states. Specifically, a link instability event can be determined when any of the following situations are detected: the PCIe port generates a CE or UCE; or, the frequency of LTSSM entering recovery state or triggering link retraining events exceeds a preset threshold. In addition, to further improve the accuracy of the judgment, information such as the current operating rate and negotiation capability can also be combined for auxiliary confirmation. For example, when the link has been reduced to Gen1 rate but the above-mentioned CE / UCE or LTSSM anomalies still occur frequently, the existence of a link instability event can be further determined.

[0049] In this embodiment, the link instability event may include at least one of the following: CE generated by PCIe port; UCE generated by PCIe port; The frequency of PCIe port link training and LTSSM entering recovery state exceeds the first preset threshold. The frequency of LTSSM-triggered link retraining events on the PCIe port exceeds the second preset threshold.

[0050] In the event of a link instability event, it indicates that the link can no longer maintain stable operation and the corresponding recovery intervention process needs to be initiated.

[0051] Specifically: (1) For CE or UCE.

[0052] In this embodiment, the PCIe switching chip obtains error information occurring on the link in real time by monitoring the AER (PCI Express Advanced Error Reporting) register or device-specific error status register of each PCIe port. Wherein: Error Response (CE) refers to a type of error that occurs at the link layer or physical layer and can be automatically recovered through retransmission or error correction mechanisms, such as ECRC (end-to-end cyclic redundancy check error) and bad transaction layer packets. Although CE does not immediately lead to data transmission failure, a continuous large number of CEs is often a precursor to a decline in link signal quality.

[0053] UCE refers to a critical error that cannot be automatically recovered from by hardware, such as a TLP infection or a flow control protocol error. The occurrence of UCE usually means that the link can no longer guarantee the correct transmission of data, and if not intervened in time, it may lead to system crashes or data corruption.

[0054] The PCIe switching chip can be set with a time window (e.g., within 1 second) and a corresponding threshold (e.g., 10 consecutive CEs or 1 UCE). When the number of detected errors exceeds the preset threshold, it is determined that a link instability event has occurred.

[0055] (2) The frequency of LTSSM entering recovery state or triggering link retraining events exceeds the corresponding preset threshold.

[0056] Specifically, the PCIe switching chip identifies the following two abnormal states by monitoring the state transitions of the LTSSM on each PCIe port: Entering Recovery State: Recovery state is a sub-state in LTSSM used for equalization parameter renegotiation or bit-locking adjustment when the link has been established. If the link frequently enters recovery state, i.e., the frequency of entering recovery state exceeds a first preset threshold (e.g., more than 5 times per minute), it indicates that although the link is not broken, the signal quality fluctuates greatly, and frequent adjustments are required to maintain the connection.

[0057] Link retraining event trigger: Link retraining refers to the process of completely re-initializing the link, usually initiated by the devices at both ends of the link or triggered by accumulated errors. If the link frequently triggers retraining events, that is, the frequency of triggering link retraining events exceeds the second preset threshold (e.g., more than 3 times per minute), it indicates that the link may be trapped in a dead loop of "link establishment-failure-link reconstruction" and cannot operate stably at any speed.

[0058] The embodiment can set a frequency threshold for the above state transitions. When the number of times the state enters the recovery state or triggers the retraining event exceeds the preset threshold within a unit of time, it is determined that a link instability event has occurred.

[0059] This embodiment uses correctable errors (CE), uncorrectable errors (UCE), and LTSSM state anomalies (frequent entry into recovery state or triggering retraining events) as the criteria for determining link instability events. This allows for comprehensive and accurate identification of different stages of link degradation, from slight signal quality degradation (increased CE) to the occurrence of serious errors (UCE), and finally to the link entering a dead loop of repeated negotiations (LTSSM anomalies). Based on these precise triggering conditions, the PCIe switching chip can promptly initiate the corresponding step-by-step recovery process, intervening before the link completely fails. This avoids system downtime due to error accumulation and prevents overreaction and service interruption while the link is still maintainable. Thus, while ensuring system stability, it maximizes service continuity.

[0060] When the aforementioned link instability event is detected, the current recovery status of the PCIe port is further queried. This current recovery status can be determined based on recovery records, which store the level, execution time, and subsequent link response results of the link recovery operations performed on the port. The embodiment can divide recovery operations into multiple levels, with different levels addressing different instability situations in the corresponding link. For example, the recovery status can be determined in reverse chronological order based on the operation type, execution time, and response results in the recovery record: if the record is empty or the most recent operation occurred after a preset time interval and the link is stable, the current recovery status of the link is that no level has been executed; if the most recent operation was at level one and the link remains unstable after execution, the current recovery status of the link is that level one has been executed; if the most recent operation was at level two and the link remains unstable after execution, the current recovery status of the link is that level two has been executed.

[0061] In one implementation, the embodiment can determine the level of the link recovery operation to be performed based on the level of the link instability event and the current recovery status. For example: If the current recovery status indicates that no link recovery operation of any level has been performed, then the first-level operation will be performed. If the current recovery status indicates that Level 1 has been performed but Level 2 has not, then Level 2 operations will be performed.

[0062] If the current recovery status indicates that Level 2 has been performed but Level 3 has not, then Level 3 operation will be performed.

[0063] In step S230, a link recovery operation corresponding to the determined level is performed.

[0064] Specifically, the PCIe switching chip can call the corresponding link recovery operation module to perform the corresponding link recovery operation according to the level determined in step S220.

[0065] If the first-level operation is determined to be performed, the PCIe switching chip will gradually reduce the maximum link negotiation rate supported by the PCIe port (e.g., from Gen5 to Gen4, then to Gen3, etc.), and trigger link renegotiation after each reduction, until the link no longer experiences link instability events, or the maximum negotiation rate has been reduced to the PCIe Gen1 rate. This level of operation does not modify the physical layer balancing parameters of the PCIe port and has minimal impact on the link.

[0066] If the second-level operation is determined, the PCIe switch chip first clears the link balancing parameters stored in the PCIe port's physical layer, and then triggers link renegotiation based on the PCIe Gen1 rate. After successfully establishing a link, the PCIe switch chip gradually increases the maximum negotiation rate supported by the port (e.g., from Gen1 to Gen2, Gen3, etc.), and monitors for link instability events after each increase. If no link instability event occurs, the increase continues until the preset maximum attempt rate is reached; if a link instability event occurs, the increase stops and the maximum negotiation rate is rolled back to the previous stable rate.

[0067] If a Level 3 operation is determined, the PCIe switching chip adjusts the physical layer transmit parameters of the PCIe port (e.g., Preset parameters, which can be adjusted in the order of P6, P7...P10) and triggers link renegotiation based on the PCIe Gen3 rate. If link instability events still occur after adjustment, the PCIe port is ultimately disabled or isolated to prevent data transmission and reception on that port, thus preventing the spread of errors and affecting the overall stability of the system.

[0068] This application implements a tiered decision-making mechanism based on link status and current recovery status in the PCIe switching chip. This mechanism executes corresponding recovery operations according to the link's instability level to achieve link recovery. This minimizes the impact of link maintenance operations on services while ensuring system stability, allowing the link to maintain connection in the lowest available state. It avoids system downtime caused by repeated link reconstruction or error accumulation, thus improving the reliability and service continuity of PCIe switch-based systems.

[0069] Based on the above embodiments, in another embodiment provided in this application, such as Figure 3 As shown, the method may further include the following steps: In step S240, when a link instability event is detected, the recovery record of the PCIe port is obtained. The recovery record includes the level of the link recovery operation performed on the PCIe port, the execution time, and the link response result after execution.

[0070] Specifically, the firmware of a PCIe switch chip can maintain a recovery log table for each PCIe port, which can be stored in the chip's internal registers or a dedicated memory area. By reading this log table, the following key information can be obtained: The levels of link recovery operations performed include Level 1, Level 2, and Level 3 operations. The record identifies the most recent operation level performed on this port, as well as a sequence of all operation levels performed historically.

[0071] Execution Time: Records the timestamp of each recovery operation to determine how close the operation was to the current time, preventing repeated execution of the same operation within a short period and thus avoiding link oscillations. For example, if a port has already performed a speed reduction operation within the last 5 seconds, even if another link instability event is detected, the next level of operation may be temporarily suspended, allowing the link some time to stabilize.

[0072] Post-operation link response results: Record the link's response after each recovery operation, such as whether the link was successfully established, the stable runtime after establishment, and whether link instability events reappeared within a certain period after the operation. These results can be used to assess the current port's link degradation level and determine whether the executed recovery operations were effective. For example, if a port only runs stably for 10 seconds after performing a Level 2 operation (clear parameters and renegotiation) before experiencing another UCE, then the port's degradation level can be considered severe, requiring an upgrade to a Level 3 operation.

[0073] In step S250, the current recovery status of the PCIe port is determined based on the recovery record.

[0074] Specifically, the current recovery stage of the PCIe port, i.e., the current recovery status, can be comprehensively determined based on the recovery record obtained in step S240. This determination can be made by considering three factors: the operation level, execution time, and link response result in the recovery record. For example, if the record is empty or the most recent operation has passed a preset time period and the link is stable, the current recovery status is "no level executed"; otherwise, based on the most recent operation level and its response result, the current recovery status is determined to be "first level executed," "second level executed," or "third level executed."

[0075] The current recovery status can be one of the following: State 0: No intervention state. There are no link recovery operation records in the recovery record, or the time of the most recent operation execution has exceeded the preset forgetting period (e.g., 30 minutes) from the current time. The link can be considered to be in the initial state.

[0076] Status 1: Level 1 operation has been performed. The recovery log shows that the most recent recovery operation was a Level 1 operation (gradual speed reduction), and this operation failed to bring the link to a stable operating state. This includes the following two situations: Scenario 1: Link instability events still occur after the operation. That is, after the first-level operation is completed, the link still detects instability such as CE, UCE, LTSSM frequently entering recovery state or triggering link retraining events within a preset time period (e.g., 5 seconds).

[0077] Scenario 2: The link stabilizes for a short period after the first-level operation before instability occurs again. That is, after the first-level operation is completed, the link briefly enters a stable state (e.g., CE / UCE or LTSSM anomalies no longer occur), but the stable duration does not exceed the preset stability threshold (e.g., 10 seconds), and then the link instability occurs again.

[0078] If any of the above conditions are met, the current recovery status is determined to be that the first-level operation has been performed but the second-level operation has not been performed.

[0079] Status 2: Level 2 operation has been performed. The recovery record shows that the most recent recovery operation was a Level 2 operation (clear parameters and renegotiation), and the link is still unstable after the operation was performed.

[0080] Status 3: Level 3 operation has been performed. The recovery record shows that the most recent recovery operation was a Level 3 operation (parameter tuning and renegotiation), and the link is still unstable after the operation was performed.

[0081] The embodiment can determine the current recovery status of the PCIe port by matching the recovery record obtained in step S240 with the preset status determination rules.

[0082] This embodiment introduces a method for acquiring and analyzing recovery records. By obtaining historical recovery behavior of PCIe ports, link recovery decisions no longer rely on a single instantaneous state but are based on complete, time-series historical information. This avoids repeatedly executing the same level of invalid operations within a short period and prevents the link from oscillating repeatedly between deceleration and acceleration. Furthermore, by recording the link response results after each operation, the current degree of link degradation can be accurately assessed, thereby selecting the most appropriate next-level operation and significantly improving the link recovery success rate of the PCIe switching chip.

[0083] Based on the above embodiments, in another embodiment provided in this application, when the link recovery operation level is the first level, such as Figure 4 As shown, step S230 above may specifically include the following steps: Step S231: Gradually reduce the maximum link negotiation rate supported by the PCIe port.

[0084] Specifically, when it is determined according to step S220 that a first-level link recovery operation needs to be performed, the first-level operation procedure is initiated. This level of operation can, without affecting the physical layer equalization parameters, force the link to operate at a rate with lower signal quality requirements by limiting the highest negotiable rate of the link, thereby avoiding signal degradation problems in the current environment.

[0085] The specific method for performing the first-level operation can include: reading the maximum link negotiation rate supported by the current PCIe port (e.g., Gen5) and downgrading it by one level (e.g., from Gen5 to Gen4). The PCIe protocol supports multiple rate levels, including Gen1 (2.5GT / s), Gen2 (5GT / s), Gen3 (8GT / s), Gen4 (16GT / s), Gen5 (32GT / s), etc., and there are clear multiple relationships between the levels. The maximum supported rate field in the port configuration register is adjusted level by level, from high to low rate.

[0086] Step S232: After each reduction, link renegotiation is triggered until no link instability event occurs, or the maximum link negotiation rate has been reduced to the PCIe Gen1 rate.

[0087] Specifically, after each reduction in the maximum negotiation rate configuration, the link renegotiation process of the PCIe port is actively triggered. This can be achieved by writing a specific renegotiation command bit to the port control register, causing the LTSSM state machine to enter Recovery state or retrain the link.

[0088] After the link renegotiation is completed, you can return to step S210 to continue monitoring the link status of the port and determine whether there are still link instability events (such as CE / UCE or frequent Recovery). If no more link instability events occur, it means that the current rate can meet the requirements for stable link operation, the first-level operation is successfully completed, and the link maintains the current reduced speed operation.

[0089] If link instability events still occur after reducing the level by one, repeat steps S231 and S232 to continue reducing the maximum negotiation rate to the next level (e.g., from Gen4 to Gen3), trigger renegotiation again, and monitor the link status.

[0090] This process repeats until one of the following two conditions is met: Scenario 1: The link no longer experiences link instability events at a certain rate level. In this case, the first-level operation is successful, and the link operates stably at that rate.

[0091] Scenario 2: The maximum link negotiation rate has been reduced to the lowest rate level defined by the PCIe protocol—Gen1 (2.5GT / s), but link instability events still occur. This indicates that simply reducing the speed cannot solve the link problem. This state will be recorded, and a higher-level recovery operation (Level 2) will be triggered subsequently.

[0092] It should be noted that during the entire Level 1 operation, the equalization parameters (such as Preset values, filter coefficients, etc.) stored in the PCIe port physical layer are not modified. These parameters are learned and solidified by the equalization algorithm during high-speed link operation, and retaining these parameters helps to quickly restore the link to a higher speed after it stabilizes.

[0093] By performing a first-level link recovery operation, this embodiment of the application achieves minimal intervention in the recovery of unstable links. This level of operation attempts to avoid signal degradation caused by environmental changes by gradually reducing the link negotiation rate and triggering renegotiation, without affecting core physical layer parameters. Therefore, it minimizes the disturbance to the link and achieves the fastest recovery speed. When the link can operate stably at a lower rate, services can continue to run without interruption, sacrificing only some bandwidth performance and avoiding service outages caused by a complete link failure.

[0094] Based on the above embodiments, in another embodiment provided in this application, when the link recovery operation is at the second level, such as Figure 5 As shown, step S230 above may specifically include the following steps: Step S233: Clear the link balancing parameters of the PCIe port physical layer and trigger link renegotiation based on the PCIe Gen1 rate.

[0095] In this embodiment, the second-level link recovery operation is used to clear old equalization parameters in the physical layer that may have been fixed and are no longer suitable for the current environment, allowing the link to relearn and adapt to the changed physical environment.

[0096] Specifically, when it is determined that a second-level link recovery operation is required, a clear command can first be written to the physical layer control register of the target PCIe port to reset various saved equalization parameters, including the de-emphasis and pre-emphasis coefficients at the transmitting end, the CTLE (continuous time linear equalization) and DFE (decision feedback equalization) coefficients at the receiving end, and the Preset value defined by the PCIe protocol. The clear operation restores the physical layer to a state similar to factory defaults. Subsequently, the maximum negotiation rate supported by the port can be temporarily locked to PCIe Gen1 (2.5GT / s), and link renegotiation can be actively triggered. Gen1 is chosen as the base rate because it has the lowest signal integrity requirements and is the easiest to successfully establish a link. During the renegotiation process, the PCIe port and the peer device re-execute physical layer initialization, including bit locking, symbol locking, and renegotiation of equalization parameters.

[0097] Step S234: After successfully establishing the link, gradually increase the maximum negotiation rate supported by the PCIe port.

[0098] Once the link is successfully established at PCIe Gen1 rate and passes stability monitoring, the rate escalation process can be initiated. Following the rate levels defined by the PCIe protocol, it attempts to escalate from Gen1 upwards, first modifying the maximum negotiation rate register to Gen2 to trigger link renegotiation. If the link operates stably at Gen2 (i.e., no link instability events occur within a preset time period), it continues to escalate to Gen3, and so on. After each escalation, a stability observation window (e.g., 30 seconds) is set to monitor for CE / UCE or LTSSM anomalies. If an instability event is detected at a certain rate, the escalation stops, and the maximum negotiation rate is rolled back to the previously confirmed stable rate. If the link successfully escalates to the preset maximum attempt rate (e.g., the port natively supports Gen5 and operates stably), the second-level operation is successfully completed, and the link operates at the highest possible stable rate.

[0099] Step S235: After each lift, determine whether a link instability event has occurred.

[0100] Specifically, after each rate increase and link renegotiation trigger, the system re-enters monitoring mode to continuously observe the link status of the target PCIe port. The judgment criteria are consistent with those in step S220, namely, monitoring the following two types of link instability events: Error event monitoring: Monitors whether the frequency of correctable errors (CEs) or uncorrectable errors (UCEs) exceeds a preset threshold. For example, if multiple CEs occur consecutively within 30 seconds after boosting to Gen2, it indicates poor signal quality at the current rate.

[0101] State anomaly monitoring: Monitor whether the LTSSM state machine frequently enters the recovery state or triggers link retraining events. For example, if the link enters the Recovery state more than 3 times per minute at Gen3 rate, it indicates that the equalization parameters may not be fully adapted to the current rate.

[0102] Set a stable observation window (e.g., 30 seconds to 1 minute) and continuously monitor the occurrence of the aforementioned link instability events within this window. If no link instability events occur or only a very small number (below the threshold) occur within the window, the link is considered to be operating stably at the current rate.

[0103] Step S236: If no link instability event occurs, continue to increase the maximum negotiation rate until the preset maximum attempt rate is reached.

[0104] Specifically, if the above judgment result is "no link instability event has occurred", that is, the link can operate stably at the current upgraded rate, then the current rate is recorded as "available stable rate", and the next round of upgrade attempts will continue.

[0105] For example, after the link operates stably under Gen2 for 30 seconds, the maximum negotiation rate is increased to Gen3, triggering renegotiation and entering the monitoring window again. If Gen3 also operates stably, the rate is increased to Gen4, and so on.

[0106] This process repeats until one of the following two conditions is met: Scenario 1: The link is successfully upgraded to the highest speed supported by the hardware capabilities of the PCIe port (e.g., the port natively supports Gen5 and can still run stably after being upgraded to Gen5). In this case, the second-level operation is successfully completed and the link is restored to full performance.

[0107] Scenario 2: The link is raised to the preset maximum attempt rate (which can be configured by the system, for example, set to Gen4 or Gen5), and further raising stops. The link continues to operate at the current stable rate.

[0108] Step S237: When a link instability event occurs, stop increasing the maximum negotiation rate and revert the maximum negotiation rate back to the previous stable rate.

[0109] Specifically, if the above judgment result is "a link instability event has occurred", that is, the link cannot operate stably at the current increased rate, then further increase attempts will be stopped and a rollback operation will be performed.

[0110] The rollback operation can be implemented by resetting the maximum negotiation rate supported by the PCIe port to the rate that has been confirmed to be stable at the previous level. For example, if the link experiences frequent UCE or Recovery events at Gen3, while Gen2 has been confirmed to be stable, the maximum negotiation rate can be rolled back to Gen2, and link renegotiation can be triggered to make the link run at the Gen2 rate again.

[0111] After the rollback is complete, re-enter monitoring mode to confirm that the link can operate stably at the rolled-back rate. If stability is confirmed, the second-level operation ends, and the link is locked at the current rolled-back rate. If the link is still unstable after the rollback (which is unlikely because the previous-level rate has been confirmed to be stable), the exception handling process is triggered or the third-level operation is initiated directly.

[0112] Understandably, by gradually increasing the speed and then rolling back when instability is detected, it is possible to find the optimal speed for each PCIe port that is both stable and as high as possible under the current environmental conditions.

[0113] Through the aforementioned second-level link recovery operation, this embodiment achieves a link recovery mechanism that, after clearing the equalization parameters and successfully establishing a link, gradually increases the rate from the lowest Gen1 rate and dynamically searches for the optimal stable rate. This avoids repeated link establishment failures caused by directly attempting a high rate all at once, making the recovery process smoother and more controllable. Furthermore, by setting a stable observation window and introducing a fallback mechanism after each rate increase, it ensures that the link operates at a verified stable rate at all times, maximizing service continuity.

[0114] For example, the specific execution process of the second-level operation can be as follows: The first step is to clear the link balancing parameters of the PCIe port physical layer.

[0115] By writing a specific clear command to the physical layer control register of the PCIe port, various equalization parameters stored in the SerDes module are reset. These parameters include, but are not limited to: Equalization parameters at the transmitting end, such as de-emphasis amplitude and pre-emphasis coefficient, are used to compensate for high-frequency losses in the signal during transmission.

[0116] Receiver equalization parameters, such as CTLE coefficients and DFE coefficients, are used to eliminate inter-symbol interference.

[0117] Preset value: A preset combination of equalization parameters. The PCIe protocol defines multiple presets, such as P0 to P10, for quickly configuring the equalizer status.

[0118] By clearing these parameters, the physical layer is restored to a state similar to the factory default, eliminating the misadaptation of the old parameters to the current link environment.

[0119] The second step is to trigger link renegotiation based on the PCIe Gen1 rate.

[0120] After the parameters are cleared, the maximum negotiation rate supported by the PCIe port is temporarily locked to PCIe Gen1, and link renegotiation is triggered. The reason for choosing Gen1 as the base rate is as follows: Gen1 (2.5GT / s) is the lowest speed and most lenient signal integrity level in the PCIe protocol. It has the least dependence on equalization parameters and is the easiest to establish a chain successfully.

[0121] Rebuild the link starting at the lowest rate so that there is a clear basis for judging stability at each step as the rate is gradually increased.

[0122] Renegotiation can be triggered by writing a link retraining command to the port control register or by simulating a link disconnection and reconnection via software. During renegotiation, the PCIe port and the peer device will re-initialize the physical layer, including renegotiating bit locking, symbol locking, and equalization parameters.

[0123] The third step is link stability monitoring.

[0124] After renegotiation, the link re-establishes a connection at Gen1 rate. The system then re-enters monitoring mode to determine if the link can operate stably at Gen1 rate. If no further instability events occur at Gen1 rate, the second-level operation is successful, and the link continues to operate at Gen1 rate. If instability events still occur at Gen1 rate, it indicates that clearing the physical layer equalization parameters cannot resolve the issue. This state will be recorded, and a higher-level recovery operation (third level) will be triggered subsequently.

[0125] It should be noted that the second-level operation has a greater impact on the link than the first level. Clearing the equalization parameters means that the link will discard the optimized configurations previously acquired through long-term learning. Thus, when the environment changes significantly (such as large temperature fluctuations or device aging), the old equalization parameters may have become an obstacle to link stability. Clearing them and relearning allows the link to better adapt to the new environment.

[0126] By performing a second-level link recovery operation, this embodiment of the application implements a deep recovery mechanism for unstable links. This level of operation, building upon the ineffectiveness of the first-level speed reduction, further removes potentially faulty equalization parameters in the physical layer, enabling the link to break free from its old environment processing mode and relearn and adapt to the current physical environment starting from the Gen1 rate. This solves the problem of existing equalization parameters becoming faulty due to environmental changes, which cannot be recovered by simple speed reduction. By rebuilding the link from Gen1, a minimum level of connectivity is ensured even under the worst conditions, guaranteeing service continuity.

[0127] Based on the above embodiments, in another embodiment provided in this application, when the level of the above link recovery operation is the third level, such as Figure 6 As shown, step S230 above may include the following steps: In step S238, the transmit parameters of the PCIe port physical layer are adjusted; In step S239, link renegotiation is triggered based on the PCIe Gen3 rate.

[0128] In this embodiment, if the link still experiences instability after the second-level link recovery operation (clearing equalization parameters, rebuilding from Gen1, and attempting to increase the rate), a third-level operation can be performed. First, the transmission parameters of the PCIe port's physical layer are adjusted. Specifically, these transmission parameters can be Preset values ​​defined by the PCIe protocol. A Preset is a set of preset transmission equalizer configurations used to control the de-emphasis amplitude and pre-emphasis intensity of the transmitted signal. Different Preset values ​​can be tried sequentially according to a preset adjustment sequence, for example, starting with P6 and trying P7, P8, P9, and P10 in sequence. After each adjustment, the corresponding configuration value is written to the physical layer control register. After completing the transmission parameter adjustment, the maximum negotiation rate supported by the PCIe port can be set to PCIe Gen3 (8GT / s), and link renegotiation can be actively triggered. Since Gen3 is the first rate class in the PCIe protocol to introduce equalization mechanisms, it is sensitive to transmission parameters and can effectively verify whether the adjusted Preset value improves signal quality. Therefore, Gen3 can be chosen as the base rate. Furthermore, compared to Gen4 / Gen5, Gen3 has relatively less stringent channel requirements and a higher success rate. During renegotiation, the port and the peer device retrain the physical layer and communicate using the specified transmission parameters.

[0129] In this embodiment, if a link instability event occurs, the LTSSM of the PCIe port can be disabled to prevent data transmission and reception on the PCIe port.

[0130] Specifically, after performing the aforementioned transmission parameter adjustments and Gen3 renegotiation, the link status monitoring phase begins. If the link no longer exhibits CE, UCE, or LTSSM anomalies within the preset observation window, it indicates that the third-level link recovery operation was successful, the link is running stably at the Gen3 rate, and the currently valid Preset configuration is recorded. If, after multiple rounds of Preset value adjustments (e.g., trying from P6 to P10), the link continues to experience unstable events, or even fails to establish a connection at all, it indicates that the current physical link has undergone extremely severe degradation, and software adjustments to the physical layer parameters cannot restore it. At this point, a final measure can be taken: disabling the LTSSM on the PCIe port. Specifically, a disable command is written to the port control register to stop the LTSSM from running. This port will be unable to participate in any link training or data transmission, thus completely preventing data transmission and reception on that port. Optionally, the PCIe DPC feature can also be used to mark the port as faulty and isolate it. This safety net measure completely isolates the faulty port, preventing errors (such as UCE and repeated link rebuilds) from spreading to other ports of the PCIe switching chip or even the entire system, thus avoiding system downtime due to a single point of failure.

[0131] For example, the specific execution process of the third-level link recovery operation is as follows: The first step is to adjust the transmit parameters of the PCIe port physical layer.

[0132] Specific configuration values ​​are written to the physical layer control register of the PCIe port to actively modify the equalization parameters of the transmitter. These parameters mainly include the Preset value, which is a set of preset transmit equalizer configuration combinations defined by the PCIe protocol, used to control characteristics such as de-emphasis amplitude and pre-emphasis intensity of the transmitted signal.

[0133] Preset values ​​typically include multiple levels, from P0 to P10, each corresponding to a different signal shaping effect. For example: P6: Medium deemphasis, suitable for channels with medium loss; P7: Stronger de-emphasis, suitable for high-loss channels; P8-P10: A more robust combination of de-emphasis and pre-emphasis, suitable for extreme wear scenarios.

[0134] The implementation can try different preset values ​​step by step according to a preset adjustment sequence (for example, starting from P6, trying P7, P8, P9, and P10 in sequence), and record the response result of the link after each adjustment. The purpose of adjusting the transmission parameters is to compensate for signal distortion caused by environmental deterioration (such as connector oxidation and PCB aging) by forcibly changing the waveform of the transmitted signal, so that the peer device can correctly receive and lock the signal.

[0135] It's important to note that adjusting transmission parameters can be a proactive intervention, unlike the second-level operation where parameters are cleared and allowed to learn automatically. Automatic learning relies on the negotiation algorithm of both ends, while proactive adjustment directly specifies the parameters, forcing the link to use those parameters for communication.

[0136] The second step is to trigger link renegotiation based on the PCIe Gen3 rate.

[0137] After parameter adjustments, the maximum negotiation rate supported by the PCIe port is set to PCIe Gen3, and link renegotiation is triggered. The reason for choosing Gen3 as the base rate is as follows: Gen3 (8GT / s) is the first rate class in the PCIe protocol to introduce a leveling mechanism. It is highly sensitive to transmission parameters and can fully verify whether the adjusted Preset value is effective. If a link can be successfully established under Gen3, it indicates that the method of actively adjusting the transmission parameters is effective, and the link still has the potential to recover to a medium-to-high speed. The method for triggering renegotiation is the same as in the previous embodiments, which can be writing a retraining command to the port control register or simulating a link disconnection and reconnection through software.

[0138] The third step is link stability monitoring and subsequent processing after performing the third-level link recovery operation.

[0139] After renegotiation is complete, enter monitoring mode to determine if the link is stable when running at Gen3 speed after parameter adjustments. The following three scenarios may occur: Scenario 1: Stable Link Operation. If no further link instability events occur at Gen3 speed, the Level 3 operation is successful, and the link maintains its Gen3 speed operation. Record the currently valid Preset configuration, and if necessary, attempt to further increase the speed later (e.g., to Gen4).

[0140] Scenario 2: The link is still unstable, but there are signs of improvement. If the link still has a small number of unstable events under Gen3, but it has improved compared to before the adjustment (e.g., the number of CEs is significantly reduced, and the recovery frequency is reduced), you can try to continue adjusting the Preset value (e.g., from P7 to P8) to trigger renegotiation again until the optimal configuration is found or all available Preset values ​​have been tried.

[0141] Scenario 3: The link cannot be established at all or remains severely unstable. If, after multiple rounds of preset adjustments, the link still cannot operate stably under Gen3, or even cannot be established at all, it indicates that the current physical link has experienced an extremely serious anomaly, and adjusting the physical layer parameters through software cannot restore it. In this case, the final fallback measure is to disable or isolate the PCIe port, preventing any data transmission or reception on that port, to prevent the error from spreading and affecting the overall stability of the system.

[0142] It should be noted that port isolation can be specifically implemented by disabling the LTSSM state machine of the port so that it no longer participates in any link training and data transmission, or by marking the port as faulty and isolating it through the PCIe DPC mechanism.

[0143] By executing a third-level link recovery operation, this embodiment implements a software intervention mechanism in harsh environments. This level of operation, after the second-level parameter clearing proved ineffective, proactively adjusts the physical layer transmission parameters (Preset value) to match the severely degraded physical link. This addresses the failure of auto-negotiation and auto-learning mechanisms in extreme environments, providing another opportunity for link recovery through proactive intervention. Furthermore, port isolation serves as a final safety net, ensuring timely mitigation even when the physical link is completely unrecoverable, preventing a single point of failure from escalating into a system-wide disaster. This allows the embodiment to maximize service continuity while ensuring the robustness of the entire system.

[0144] This embodiment completely blocks data transmission and reception on the faulty port by disabling LTSSM, effectively preventing a single point of failure from spreading into a bus storm or system-level deadlock, thus ensuring the stability and reliability of the entire system.

[0145] Based on the above embodiments, in another embodiment provided in this application, when the link instability event is detected, as follows: Figure 7 As shown, step S220 above may specifically include the following steps: Step S221: When the current recovery status of the PCIe port indicates that no link recovery operation of any level has been performed, determine that the level of the link recovery operation to be performed is the first level.

[0146] Specifically, when a link instability event is detected (such as continuous CE / UCE or frequent Recovery), and after querying the recovery record of the PCIe port, it is found that the port has not performed any level of link recovery operation (i.e., it is in the initial state), the first level of link recovery operation with the least impact can be determined to be performed according to the principle of from shallow to deep.

[0147] Step S222: When the current recovery status of the PCIe port indicates that a first-level link recovery operation has been performed but a second-level link recovery operation has not been performed, determine that the level of the link recovery operation to be performed is a second-level link.

[0148] Specifically, if a link instability event is detected, and the recovery log shows that the port has already undergone Level 1 operations (e.g., attempted to reduce speed to Gen1), but has not yet undergone Level 2 operations, it indicates that the Level 1 operations failed to resolve the issue and the intervention intensity needs to be escalated. In this case, it is determined to execute Level 2 link recovery operations.

[0149] Step S223: When the current recovery status of the PCIe port indicates that a second-level link recovery operation has been performed but a third-level link recovery operation has not been performed, determine that the level of the link recovery operation to be performed is the third level.

[0150] If a link instability event is detected, and the recovery log shows that the port has already undergone Level 2 operations (e.g., parameters have been cleared and an attempt to raise the port's elevation), but not Level 3 operations, it indicates that Level 2 operations have also failed to resolve the issue, and the link degradation is severe. In this case, it is determined to perform Level 3 link recovery operations.

[0151] If no link instability events are detected during the monitoring period, it indicates that the current link is operating stably and no intervention is required. In this case, the current link state of the PCIe port should be maintained, including the current operating speed, load balancing parameters, and other configurations, and normal data transmission should continue.

[0152] This embodiment can progressively select more in-depth intervention methods based on the persistence of link instability events and the levels of recovery operations already performed. This ensures that each intervention is a natural escalation after the previous milder methods have failed, avoiding excessive impact on services caused by escalating operations. Simultaneously, by maintaining the current link state when no link instability events are detected, unnecessary link oscillations and performance losses are avoided, significantly improving the PCIe switching chip's adaptability and operational efficiency in complex environments.

[0153] In this embodiment, after performing the above link recovery operation, if no link instability event occurs within the preset time window, the current recovery status of the PCIe port can be reset to a state where no link recovery operation of any level has been performed.

[0154] Specifically, after completing any level of link recovery operation (Level 1, Level 2, or Level 3) and the link has entered a stable operating state, a status reset timer can be started. By continuously monitoring the link status of this PCIe port, if no link instability events (including frequent entry into recovery state by CE, UCE, or LTSSM, or triggering of link retraining events) occur within a preset time window (e.g., 10 minutes or 30 minutes), it is considered that the current link quality has been well restored, and the environmental conditions have improved or the deteriorating factors have been eliminated.

[0155] In this situation, a state reset operation can be performed. This clears the historical record of the port stored in the recovery log table, or marks it as failed, and sets the current recovery state of the PCIe port to the initial state where no link recovery operation at any level has been performed. Subsequently, if the PCIe port experiences link instability again, the firmware will restart the stepwise recovery process from the first level of operation (gradually reducing speed) instead of jumping directly to a higher level previously used.

[0156] It should be noted that the preset time window should not be too short (e.g., greater than the preset first threshold) to avoid frequent resets due to short-term environmental fluctuations; nor should it be too long (e.g., less than the preset second threshold), otherwise the link may remain in a high-level recovery state after long-term stability, potentially missing the opportunity to handle newly emerging problems with gentler methods. This time window can be configured by the system according to the actual application scenario, for example, it can be set to 30 minutes.

[0157] This embodiment introduces a recovery state reset mechanism. When the link remains stable for a long time, the recovery state can be rolled back to the initial state. This allows the system to start from the mildest first-level operation (slowing down) and retry when an unstable event occurs again in the future, avoiding unnecessary advanced interventions caused by historical records.

[0158] By dividing each functional module according to its corresponding function, this application provides a link recovery device, which can be a server or a chip applied to a server. Figure 8 A schematic block diagram of the functional modules of a link recovery device provided for an exemplary embodiment of this application. Figure 8 As shown, the link recovery device includes: Link status acquisition module 81 is used to acquire the link status of the PCIe port; The level determination module 82 is used to determine the level of the link recovery operation to be performed when a link instability event is determined to exist in the link based on the link status, according to the link instability event and the current recovery status of the PCIe port; wherein, the current recovery status is determined based on the record of the link recovery operations already performed on the PCIe port, and different levels represent different degrees of intervention in the link; The recovery operation execution module 83 is used to perform link recovery operations corresponding to the determined level.

[0159] In yet another embodiment provided in this application, The aforementioned link instability events include at least one of the following: Correctable error CE generated by PCIe port; Uncorrectable errors (UCE) generated by PCIe ports; The frequency with which the PCIe port's Link Training and State Machine (LTSSM) enters the recovery state exceeds a first preset threshold. The frequency of LTSSM-triggered link retraining events on the PCIe port exceeds the second preset threshold.

[0160] This embodiment uses abnormal CE, UCE, and LTSSM states (frequent entry into recovery state or triggering retraining events) as the criteria for determining link instability events. This allows for comprehensive and accurate identification of different stages of link degradation, from slight signal quality degradation (increased CE) to the occurrence of serious errors (UCE), and finally to the link falling into a dead loop of repeated negotiations (LTSSM anomaly). Based on these precise triggering conditions, the PCIe switching chip can promptly initiate the corresponding step-by-step recovery process, intervening before the link completely fails. This avoids system downtime due to error accumulation and prevents overreaction and service interruption while the link is still maintainable. Thus, while ensuring system stability, it maximizes service continuity.

[0161] In another embodiment provided in this application, the device further includes a recovery status confirmation module, specifically used for: When the link instability event is detected, the recovery record of the PCIe port is obtained. The recovery record includes the level of the link recovery operation performed on the PCIe port, the execution time, and the link response result after the execution. Based on the recovery record, determine the current recovery status of the PCIe port.

[0162] This embodiment introduces a method for acquiring and analyzing recovery records. By obtaining historical recovery behavior of PCIe ports, link recovery decisions no longer rely on a single instantaneous state but are based on complete, time-series historical information. This avoids repeatedly executing the same level of invalid operations within a short period and prevents the link from oscillating repeatedly between deceleration and acceleration. Furthermore, by recording the link response results after each operation, the current degree of link degradation can be accurately assessed, thereby selecting the most appropriate next-level operation and significantly improving the link recovery success rate of the PCIe switching chip.

[0163] In another embodiment provided in this application, when the above-mentioned level is the first level, the recovery operation execution module 83 is specifically used for: The maximum link negotiation rate supported by the PCIe port is gradually reduced; Each reduction triggers a link renegotiation until no link instability event occurs, or the maximum link negotiation rate has been reduced to the PCIe Gen1 rate.

[0164] This embodiment achieves minimal intervention-based recovery of unstable links. This level of operation attempts to circumvent signal degradation caused by environmental changes by gradually reducing the link negotiation rate and triggering renegotiation, without affecting core physical layer parameters. Therefore, it minimizes link disturbance and achieves the fastest recovery speed. When the link can operate stably at a lower rate, services can continue to run without interruption, sacrificing only some bandwidth performance and avoiding service outages caused by complete link failure.

[0165] In another embodiment provided in this application, when the above-mentioned level is the second level, the recovery operation execution module 83 is specifically used for: Clear the link balancing parameters of the PCIe port physical layer and trigger link renegotiation based on the PCIe Gen1 rate. After successfully establishing the link, gradually increase the maximum negotiation rate supported by the PCIe port; After each lift, determine whether a link instability event has occurred. If no link instability event occurs, continue to increase the maximum negotiation rate until the preset maximum attempt rate is reached; When a link instability event occurs, stop increasing the maximum negotiation rate and revert the maximum negotiation rate back to the previous stable rate.

[0166] This embodiment implements a deep recovery mechanism for unstable links. Building upon the ineffectiveness of the first-level speed reduction, this level of operation further removes potentially faulty equalization parameters from the physical layer, enabling the link to break free from its old processing mode and relearn and adapt to the current physical environment starting at Gen1 speed. This solves the problem of existing equalization parameters becoming faulty due to environmental changes, which cannot be recovered by simple speed reduction. By rebuilding the link from Gen1, a minimum level of connectivity is ensured even under the worst conditions, guaranteeing service continuity.

[0167] In another embodiment provided in this application, when the above-mentioned level is the third level, the recovery operation execution module 83 is specifically used for: Adjust the physical layer transmit parameters of the PCIe port; Link renegotiation is triggered based on PCIe Gen3 rate; This embodiment, after the second-level parameter clearing proved ineffective, proactively adjusts the physical layer transmission parameters to match the severely degraded physical link. This addresses the issue of auto-negotiation and auto-learning mechanisms failing in extreme environments, providing another opportunity for link recovery through proactive intervention.

[0168] In another embodiment provided in this application, the level determination module 82 is specifically used for: If the current recovery status of the PCIe port indicates that no link recovery operation of any level has been performed, the level of the link recovery operation to be performed is determined to be Level 1. When the current recovery status of the PCIe port indicates that a first-level link recovery operation has been performed but a second-level link recovery operation has not been performed, the level of the link recovery operation to be performed is determined to be a second-level link. If the current recovery status of the PCIe port indicates that a Level 2 link recovery operation has been performed but a Level 3 link recovery operation has not been performed, then the level of the link recovery operation to be performed is determined to be Level 3.

[0169] This embodiment can progressively select more in-depth intervention methods based on the persistence of link instability events and the levels of recovery operations already performed. This ensures that each intervention is a natural escalation after the previous milder methods have failed, avoiding excessive impact on services caused by escalating operations. Simultaneously, by maintaining the current link state when no link instability events are detected, unnecessary link oscillations and performance losses are avoided, significantly improving the PCIe switching chip's adaptability and operational efficiency in complex environments.

[0170] In another embodiment provided in this application, the device further includes a link state maintenance module, specifically used for: Maintain the current link state of the PCIe port when no link instability event is detected.

[0171] In this embodiment, if no link instability events are detected during the monitoring period, it indicates that the current link is operating stably, and no intervention is required.

[0172] In another embodiment provided in this application, the device further includes a state reset module, specifically used for: After performing the above link recovery operation, if no link instability event occurs within the preset time window, the current recovery status of the PCIe port will be reset to the state where no link recovery operation of any level has been performed.

[0173] This embodiment introduces a recovery state reset mechanism. When the link remains stable for a long time, the recovery state can be rolled back to the initial state. This allows the system to start from the mildest first-level operation (slowing down) and retry when an unstable event occurs again in the future, avoiding unnecessary advanced interventions caused by historical records.

[0174] This application also provides a computing device, such as... Figure 9 As shown, Figure 9 This is a schematic diagram of a computing device provided in an embodiment of this application. Specifically, the computing device can be the aforementioned server, including a PCIe switch chip 1901, a communication interface 1902, and a memory 1903. The PCIe switch chip 1901, communication interface 1902, and memory 1903 communicate with each other via a communication bus. Furthermore, multiple PCIe devices can be connected to the PCIe switch chip 1901 via the communication bus, such as PCIe device 1904 and PCIe device 1905, etc., but the embodiment is not limited to this.

[0175] Memory 1903 is used to store computer programs; The PCIe switching chip 1901 is used to implement the above-described method provided in the embodiments of this application when executing the program stored in the memory 1903.

[0176] The communication bus mentioned in the above computing device may specifically be a PCIe bus, etc., but the embodiments are not limited to this. For ease of illustration, it is represented by thick lines in the figure, but this does not indicate the number or type of bus.

[0177] Communication interface 1902 is used for communication between the aforementioned computing device and other devices.

[0178] The memory 1903 may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device.

[0179] In another embodiment provided in this application, a computer-readable storage medium is also provided, which stores a computer program that, when executed by a PCIe switching chip, implements the methods described above in the embodiments of this application.

[0180] In another embodiment provided in this application, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute the methods described above in the embodiments of this application.

[0181] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).

[0182] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0183] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments of apparatus, computing devices, and computer-readable storage media are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0184] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application are included within the scope of protection of this application.

Claims

1. A link recovery method, characterized in that, The method includes: Obtain the link status of high-speed PCIe ports interconnecting peripheral components; When a link instability event is determined to exist in the link based on the link status, the level of the link recovery operation to be performed is determined according to the link instability event and the current recovery status of the PCIe port; wherein, the current recovery status is determined based on the record of link recovery operations already performed on the PCIe port, and different levels represent different degrees of intervention in the link; Perform the link recovery operation corresponding to the determined level.

2. The method according to claim 1, characterized in that, The link instability event includes at least one of the following: The correctable error CE generated by the PCIe port; The uncorrectable error (UCE) generated by the PCIe port; The frequency at which the link training and state machine (LTSSM) of the PCIe port enters the recovery state exceeds a first preset threshold. or The frequency of LTSSM-triggered link retraining events on the PCIe port exceeds a second preset threshold.

3. The method according to claim 1, characterized in that, The method further includes: When the link instability event is detected, the recovery record of the PCIe port is obtained. The recovery record includes the level, execution time, and link response result of the link recovery operation performed on the PCIe port. Based on the recovery record, the current recovery status of the PCIe port is determined.

4. The method according to claim 1, characterized in that, When the level is Level 1, performing the link recovery operation corresponding to the determined level includes: The maximum link negotiation rate supported by the PCIe port is gradually reduced. Each reduction triggers a link renegotiation until no link instability event occurs, or the maximum link negotiation rate has been reduced to the PCIe Gen1 rate.

5. The method according to claim 1, characterized in that, When the level is level two, performing the link recovery operation corresponding to the determined level includes: Clear the link balancing parameters of the physical layer of the PCIe port and trigger link renegotiation based on the PCIe Gen1 rate. After successfully establishing the link, gradually increase the maximum negotiation rate supported by the PCIe port; After each elevation, determine whether the link experiences an instability event; When no link instability event occurs on the link, the maximum negotiation rate continues to be increased until the preset maximum attempt rate is reached; When the link becomes unstable, stop increasing the maximum negotiation rate and revert the maximum negotiation rate to the previous stable rate.

6. The method according to claim 1, characterized in that, When the level is level three, performing the link recovery operation corresponding to the determined level includes: Adjust the physical layer transmission parameters of the PCIe port; Link renegotiation is triggered based on PCIe Gen3 rate.

7. The method according to claim 1, characterized in that, The step of determining the level of the link recovery operation to be performed based on the link instability event and the current recovery status of the PCIe port includes: When the current recovery status of the PCIe port indicates that no link recovery operation of any level has been performed, the level of the link recovery operation to be performed is determined to be Level 1. When the current recovery status of the PCIe port indicates that a first-level link recovery operation has been performed but a second-level link recovery operation has not been performed, the level of the link recovery operation to be performed is determined to be a second-level link. When the current recovery status of the PCIe port indicates that a second-level link recovery operation has been performed but a third-level link recovery operation has not been performed, the level of the link recovery operation to be performed is determined to be the third level.

8. The method according to claim 1, characterized in that, The method further includes: If no link instability event is detected, maintain the current link state of the PCIe port.

9. The method according to claim 1, characterized in that, The method further includes: After performing the link recovery operation, if the link does not experience any link instability events within a preset time window, the current recovery status of the PCIe port will be reset to a state where no link recovery operation at any level has been performed.

10. A computing device, characterized in that, include: PCIe switching chip; Memory used to store executable instructions of the PCIe switching chip; The PCIe switching chip is configured to execute the instructions to implement the method as described in any one of claims 1-9.