SerDes link initialization method and apparatus

By detecting and marking the health level of SerDes channels, and implementing parallel alignment and dynamic grouping strategies, the problems of low initialization efficiency and low resource utilization in existing technologies are solved, achieving faster and more reliable SerDes link initialization.

CN121070853BActive Publication Date: 2026-05-05SUZHOU YIGE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SUZHOU YIGE TECH CO LTD
Filing Date
2025-08-21
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing technologies suffer from low efficiency in SerDes link initialization due to the use of a serial verification process, and the fixed packet strategy lacks flexibility and results in low resource utilization.

Method used

By detecting multiple SerDes channels and marking their health levels, single-channel symbol alignment and multi-channel collaborative alignment are performed in parallel. Grouping and binding strategies are dynamically generated by combining historical initialization data, and clock or delay compensation logic is automatically inserted when necessary.

Benefits of technology

It significantly reduces initialization time, improves resource utilization and grouping flexibility, and ensures reliable communication in complex fault scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121070853B_ABST
    Figure CN121070853B_ABST
Patent Text Reader

Abstract

This application relates to the field of high-speed serial communication technology and discloses a SerDes link initialization method and apparatus. The method first evaluates the operating status of multiple SerDes channels and assigns a corresponding health level to each channel. Then, for channels with acceptable health levels, it simultaneously initiates both single-channel and multi-channel alignment attempts to obtain channels that have completed symbol alignment. Finally, the method determines a grouping and binding strategy for successfully aligned channels based on historical initialization data and executes the binding operation. This method improves the speed and adaptability of SerDes link initialization by evaluating channel health status, parallelizing the alignment process, and using historical data for intelligent grouping.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of high-speed serial communication technology, specifically to a SerDes link initialization method and apparatus. Background Technology

[0002] In the field of high-speed serial communication, multi-channel SerDes (serializer / deserializer) technology is widely used to meet the ever-increasing bandwidth demands. During the initialization of a SerDes link, multiple channels need to be grouped and aligned to establish stable and reliable communication. Existing technologies typically employ fixed grouping strategies and serial verification mechanisms. For example, they might first attempt to align all channels as a whole; if this fails, they are downgraded to predefined sub-groups for retry. This linear, serialized process requires waiting for a fixed timeout period before proceeding to the next attempt when some channels experience transient interference or faults, resulting in excessively long initialization times. Furthermore, its fixed grouping strategy lacks flexibility and cannot adaptively adjust to actual fault modes, leading to repeated invalid attempts in complex fault scenarios and reducing the utilization of hardware resources.

[0003] Therefore, there is an urgent need for a SerDes link initialization method to solve the problems of low initialization efficiency caused by the use of serial verification process in existing technologies, as well as poor flexibility and low resource utilization caused by the use of fixed grouping strategy. Summary of the Invention

[0004] In view of this, this application provides a SerDes link initialization method and apparatus to solve the problems of low initialization efficiency caused by the use of serial verification process in the prior art, as well as poor flexibility and low resource utilization caused by the use of fixed grouping strategy.

[0005] Firstly, this application provides a SerDes link initialization method, which includes:

[0006] Multiple SerDes channels are tested, and each SerDes channel is marked with its corresponding health level based on the test results;

[0007] For SerDes channels that meet the preset health level, single-channel sign alignment and multi-channel collaborative alignment are performed in parallel to obtain SerDes channels with achieved sign alignment.

[0008] Based on historical initialization data, determine the group binding strategy for the SerDes channel with implemented symbol alignment, and perform binding according to the group binding strategy.

[0009] In one alternative implementation, multiple SerDes channels are detected, including:

[0010] Hardware self-tests are performed simultaneously on multiple SerDes channels to obtain the signal integrity and power stability of each SerDes channel.

[0011] In one optional implementation, the health level includes three levels: healthy, minor fault, and major fault.

[0012] Each SerDes channel is labeled with its corresponding health level, including:

[0013] SerDes channels with repairable signal integrity or power stability issues are marked as minor fault levels.

[0014] SerDes channels with physical damage are marked as having a severe fault level.

[0015] In one alternative implementation, detecting multiple SerDes channels further includes:

[0016] Iterate through multiple SerDes channels to determine the physical resource group or die to which each SerDes channel belongs, and check for clock domain conflicts or input / output resource contention.

[0017] In one alternative implementation, single-channel symbol alignment includes establishing a separate symbol lock for each SerDes channel;

[0018] Multi-channel collaborative alignment includes initiating multi-channel binding preparation, which includes packet pre-synchronization or packet pre-alignment.

[0019] In one alternative implementation, performing single-channel symbol alignment and multi-channel collaborative alignment in parallel further includes:

[0020] Based on the SerDes channel that was aligned first in single-channel symbol alignment and multi-channel collaborative alignment, subsequent group binding steps are performed.

[0021] In one alternative implementation, the historical initialization data includes at least one of fault type, component power, and bandwidth performance.

[0022] In one alternative implementation, a grouping strategy for the symbol-aligned SerDes channel is determined based on historical initialization data, including:

[0023] Based on historical initialization data, a pre-trained adaptive learning optimization model is used to generate a grouping binding strategy for SerDes channels with already achieved symbol alignment.

[0024] In one alternative implementation, binding is performed, including:

[0025] When the grouping binding strategy involves SerDes channels spanning different physical resource groups or chip dies, clock or delay compensation logic is automatically inserted based on the detection results.

[0026] Secondly, this application discloses a SerDes link initialization device, the device comprising:

[0027] The grading module is used to detect multiple SerDes channels and mark each SerDes channel with a corresponding health level based on the detection results.

[0028] The alignment module is used to perform single-channel sign alignment and multi-channel collaborative alignment in parallel for SerDes channels that meet the preset health level, so as to obtain SerDes channels with achieved sign alignment.

[0029] The grouping module is used to determine the grouping binding strategy for SerDes channels with implemented symbol alignment based on historical initialization data, and to perform binding according to the grouping binding strategy.

[0030] This application provides a SerDes link initialization method. First, it lays a solid data foundation for the entire initialization process through a defined comprehensive detection mechanism. This mechanism not only quickly obtains physical layer health indicators such as signal integrity of each channel through parallel hardware self-testing, but also performs fine-grained classification of channel status through a three-level health level division. Furthermore, it traverses the channels to obtain their physical topology location and resource conflict information within the system. This combination of detection steps ensures that subsequent alignment and grouping steps are based on accurate data. Second, based on the above comprehensive detection, this application solves the problem of excessively long initialization time in existing technologies through a defined efficient parallel alignment mechanism. This mechanism initiates two complementary alignment attempts in parallel for detected healthy channels: one is a single-channel fast symbol locking for speed, and the other is a multi-channel collaborative binding preparation focused on establishing a wide bus. Crucially, this application also introduces an arbitration mechanism to ensure that the entire process can continue based on the first successful alignment result. This mode directly translates the advantages of parallelization into a reduction in initialization time, thereby efficiently obtaining a batch of aligned and usable channels, providing input for subsequent intelligent grouping. Finally, this application achieves optimal allocation of link resources through a defined intelligent grouping decision-making and execution mechanism. This mechanism utilizes a pre-trained adaptive learning optimization model to predict and generate the current optimal grouping strategy based on historical data including objective performance indicators such as fault type and success rate. This makes grouping decisions dynamic and adaptive, rather than fixed and rigid. More importantly, when executing this strategy, this application cleverly links the grouping step with the initial detection step: when complex binding across physical resource groups is required, the system automatically inserts corresponding clock or delay compensation logic using the detected topology information. This interconnected design ensures that intelligent decisions are accurately implemented, effectively solving the channel binding problem under complex layouts. In summary, this application forms a complete technical closed loop by combining comprehensive detection, efficient parallel alignment, and intelligent grouping decision-making in a tightly linked manner. Overall, it solves the core problems of low initialization efficiency, poor flexibility, and low resource utilization in existing technologies, achieving faster, more reliable, and more intelligent SerDes link initialization. Attached Figure Description

[0031] To more clearly illustrate the technical solutions in the specific embodiments of this application or the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0032] Figure 1 This is a schematic diagram of a SerDes link initialization method according to an exemplary embodiment;

[0033] Figure 2 This is a schematic diagram of the structure of a SerDes link initialization device provided in an embodiment of this application;

[0034] Figure 3 This is a schematic diagram of the structure of a computer device provided in an optional embodiment of this application. Detailed Implementation

[0035] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0036] It should be understood that the term "instruction" mentioned in the embodiments of this application can be a direct instruction, an indirect instruction, or an indication of a relationship. For example, A instructing B can mean that A directly instructs B, such as B being able to obtain information through A; it can also mean that A indirectly instructs B, such as A instructing C, so B can obtain information through C; or it can mean that there is a relationship between A and B.

[0037] In the description of the embodiments of this application, the term "correspondence" may indicate that there is a direct or indirect correspondence between two things, or that there is an association between two things, or that there is a relationship of instruction and being instructed, configuration and being configured, etc.

[0038] In the embodiments of this application, "predefined" can be achieved by pre-storing corresponding codes, tables or other means that can be used to indicate relevant information in the device (e.g., including terminal devices and network devices). This application does not limit the specific implementation method.

[0039] First, let me introduce the terminology used in this application.

[0040] SerDes: Serializer / Deserializer. It refers to a core high-speed communication circuit or technology whose function is to convert (serialize) parallel data within a chip into a high-speed serial data stream for transmission, and at the receiving end, convert the serial data stream back (deserialize) it into parallel data.

[0041] Lane: A channel or link. It refers to a single, complete SerDes serial communication path. When multi-lane parallel transmission is mentioned, it means that multiple such independent channels are working in parallel to form a communication link with a higher total bandwidth.

[0042] AI: Artificial Intelligence.

[0043] FPGA: Field-Programmable Gate Array.

[0044] Bank: Physical Resource Group. In FPGAs or large chips, this is a unit of physical area or resource grouping.

[0045] Die: A bare silicon chip or wafer. It refers to a single, independent silicon chip. In modern advanced packaging technologies, a single chip product may contain multiple dies.

[0046] PLL: Phase-Locked Loop. It refers to a key circuit used to generate and stabilize high-frequency clock signals.

[0047] IO: Input / Output. It refers to the physical pins or interface resources that enable a chip to communicate with the outside world.

[0048] Channel Ready: This is a final state in the initialization process. Once a lane or a group of lanes successfully passes all the detection, alignment, binding, and validation processes, it enters the Channel Ready state, signifying that the channel is ready to begin normal, reliable data communication.

[0049] In the field of high-speed serial communication technology, multi-channel SerDes technology is the core for realizing high-bandwidth data transmission between chips, and the efficiency and reliability of its link initialization process are crucial. Existing technologies typically employ a linear, serial verification process for handling multi-channel SerDes initialization, especially packet alignment. For example, the system first attempts to align all channels as a whole; if this attempt fails due to timeout, it then reverts to a preset sub-packet strategy (such as 4+4 or 2+2 packets) for another attempt. The fundamental flaw of this method lies in its serial dependency and the rigidity of the strategy. When some channels experience transient interference or physical performance degradation, the entire process must wait for a fixed, long timeout period to end before proceeding to the next step, which significantly prolongs the link establishment time and reduces initialization efficiency.

[0050] Furthermore, existing technologies have significant shortcomings in their initialization decision-making. First, they lack a comprehensive, proactive detection mechanism. They typically do not perform parallel hardware self-tests on all channels to obtain physical layer health indicators such as signal integrity and power stability before attempting alignment, nor do they systematically traverse channels to determine their respective physical resource groups or chip dies, and preemptively investigate topology risks such as clock domain conflicts. Second, their fault handling model is relatively coarse, usually only having two states: success or failure. It cannot finely classify faults, such as distinguishing between repairable minor faults and physically damaged major faults, resulting in a single processing approach: either wasting time trying to repair unrecoverable channels or directly discarding channels with a chance of repair. Finally, the decision-making logic of existing technologies is static. They cannot utilize historical initialization data (such as past fault types, grouping success rates, or bandwidth performance) to dynamically and intelligently adjust the current grouping strategy, lacking a mechanism for adaptive learning and optimization to cope with the impact of hardware aging or environmental changes.

[0051] Therefore, this application provides a SerDes link initialization method. Figure 1 This is a flowchart of the SerDes link initialization method according to an embodiment of this application, as follows: Figure 1 As shown, the process includes the following steps:

[0052] S101. Detect multiple SerDes channels and mark each SerDes channel with its corresponding health level based on the detection results.

[0053] Specifically, the health rating is a classification label used to indicate the current operational status or quality of a SerDes channel. It is a qualitative or quantitative rating derived from the test results. For example, one rating might indicate that the channel is working perfectly, while another rating might indicate that the channel has a problem.

[0054] Optionally, multiple SerDes channels can be detected, including:

[0055] Hardware self-tests are performed simultaneously on multiple SerDes channels to obtain the signal integrity and power stability of each SerDes channel.

[0056] Hardware self-test refers to a function integrated into the SerDes hardware itself, which can check the operating status of its internal circuits and signal links without the need for external testing equipment.

[0057] Signal integrity is a metric that measures the quality of high-speed serial signals. A high-quality signal has a clear waveform and low distortion. It is usually quantified by the bit error rate (BER), which is the probability of an erroneous bit occurring during transmission.

[0058] Power supply stability refers to the stability of the power supply voltage supplied to the SerDes circuit. Unstable power supply (such as voltage fluctuations exceeding a threshold) will directly affect circuit performance and lead to data transmission errors.

[0059] Compared to sequentially inspecting each channel one by one, parallel inspection can acquire a snapshot of the physical state of all channels in a very short time. This provides a timely and comprehensive data foundation for subsequent health level classification, improving the efficiency of overall initialization from the very beginning of the process.

[0060] Health levels are categorized into three categories: healthy, minor faults, and major faults. Minor faults refer to non-permanent faults that can be recovered from through software or firmware adjustments. This includes clearing historical error states or fine-tuning equalizer parameters. Major faults refer to permanent faults caused by physical hardware damage that cannot be recovered from through simple adjustments.

[0061] Each SerDes channel is labeled with its corresponding health level, including:

[0062] SerDes channels with repairable signal integrity or power stability issues are marked as minor fault levels.

[0063] SerDes channels with physical damage are marked as having a severe fault level.

[0064] In other words, if a channel's signal integrity or power stability problem is repairable, it is marked as a minor fault; if the test results indicate that the channel has physical damage, it is marked as a major fault.

[0065] This hierarchical mechanism enables the method to accurately distinguish and differentiate faults of different natures. Compared with the general fault / normal binary division in the prior art, this embodiment can identify minor fault channels that can be quickly repaired, thereby avoiding unnecessary channel discarding. At the same time, it isolates major fault channels to prevent them from affecting the initialization of the entire link, significantly improving the granularity of fault handling and resource utilization.

[0066] Optionally, detecting multiple SerDes channels also includes:

[0067] Iterate through multiple SerDes channels to determine the physical resource group or die to which each SerDes channel belongs, and check for clock domain conflicts or input / output resource contention.

[0068] In addition to physical health checks on each channel itself, a system-level topology check is also required. All channels are traversed to determine their physical location on the chip (which physical resource group or die they belong to) and to check for potential resource conflicts, such as shared clocks or I / O pins. The information obtained in this step is crucial for subsequent intelligent grouping, especially when it is necessary to combine physically dispersed channels together.

[0069] Physical resource groups or die blocks refer to the physical grouping or location of SerDes channels on a chip. In complex chips such as FPGAs, circuits are typically divided into different regions or packaged together from multiple independent silicon dies.

[0070] Clock domain conflict refers to data synchronization problems that may occur when circuit parts that need to work together use clocks from different sources or with different phases.

[0071] Input / output resource contention refers to a conflict that occurs when multiple functional modules attempt to use the same physical I / O pin simultaneously.

[0072] S102. For SerDes channels that meet the preset health level, perform single-channel sign alignment and multi-channel collaborative alignment in parallel to obtain SerDes channels with achieved sign alignment.

[0073] Specifically, based on the health level obtained in the first step, only channels in good condition that meet preset conditions are processed. Then, two different types of alignment attempts are initiated simultaneously: one for a single channel and the other for multiple channels working together. The goal is to improve efficiency. Regardless of which attempt succeeds, the ultimate goal is to obtain a batch of channels that have completed symbol alignment and are ready to proceed to the next step.

[0074] Parallel execution refers to starting and performing two or more operations simultaneously within the same time period, rather than executing them one after another in sequence. This is a processing method designed to reduce the overall processing time.

[0075] Single-channel symbol alignment is a fundamental synchronization process for a single SerDes channel. Because serial data is a long, continuous stream of bits, the receiver needs this process to accurately identify the boundaries of symbols (typically 8-bit or 10-bit data units). Only after this process is completed can a single channel correctly parse the data.

[0076] Multi-channel collaborative alignment is a more complex synchronization process that not only aligns the symbols within each channel but also ensures that the data across multiple channels is aligned in time. Due to physical differences, there will be slight deviations (offsets) in the transmission time of data on different channels. This process is designed to compensate for these deviations, enabling multiple channels to work as a wider, unified whole.

[0077] Optionally, single-channel symbol alignment includes establishing a separate symbol lock for each SerDes channel;

[0078] Multi-channel collaborative alignment includes initiating multi-channel binding preparation, which includes packet pre-synchronization or packet pre-alignment.

[0079] Symbol locking is a technique used at the SerDes receiver. In order to correctly parse data, the receiver must accurately identify the boundaries of symbols (typically 8-bit or 10-bit data blocks) within a continuous bit stream. Establishing symbol locking is the first step in achieving reliable communication.

[0080] Multichannel binding preparation is a series of preparatory operations performed to achieve channel binding (that is, logically merging multiple independent SerDes channels into a wider channel).

[0081] Pre-synchronization or pre-alignment of channels is a preliminary time alignment and state coordination of the channels within a potential channel group before formal binding, in order to improve the success rate and speed of the final binding.

[0082] This design allows the two parallel approaches to have clear and complementary goals. Single-channel alignment aims for the fastest possible availability of a single channel, while multi-channel collaborative alignment focuses on the preparation for establishing a wider bus. This dual-strategy parallel approach covers different optimization directions, enabling a usable intermediate state to be reached in a shorter time, regardless of whether the final system requires a fast-available single channel or a stable multi-channel group.

[0083] Optionally, performing single-channel sign alignment and multi-channel co-alignment in parallel also includes:

[0084] Based on the SerDes channel that was aligned first in single-channel symbol alignment and multi-channel collaborative alignment, subsequent group binding steps are performed.

[0085] The entire initialization process does not wait for both alignment attempts to complete. Instead, it adopts a first-come, first-served strategy. Whether it's single-channel alignment or multi-channel collaborative alignment, whichever produces a batch of successfully aligned channels first, that result is used as the basis for seamlessly moving to the next step (grouping and binding). This mechanism is key to ensuring that parallel design can truly shorten initialization time.

[0086] This mechanism ensures that the initialization process does not have to wait for both parallel attempts to complete, but can immediately proceed to the next step after either attempt achieves a partial success. This avoids unnecessary waiting time and is key to translating the advantages of parallelization into actual time savings, thereby further reducing the overall initialization time.

[0087] S103. Based on historical initialization data, determine the group binding strategy for the SerDes channel that has achieved symbol alignment, and perform binding according to the group binding strategy.

[0088] Specifically, the core of this step lies in the fact that its decisions are based on historical initialization data. This means that its decisions are not fixed but based on past experience. Based on this historical data, a grouping and binding strategy is formulated, which determines how to combine these available channels (e.g., combining four channels into a group, or two channels into a group, etc.). After the strategy is determined, binding is executed, putting the strategy into practice and ultimately forming one or more usable communication links.

[0089] Historical initialization data includes at least one of fault type, component power, and bandwidth performance.

[0090] Fault type is a classification record of various errors that occurred during the initialization process in the past.

[0091] Grouping success rate is a statistical measure of the historical success or failure frequency of different channel grouping schemes.

[0092] Bandwidth performance is a record of the data throughput achieved by different packet schemes in actual operation in the past.

[0093] Optionally, based on historical initialization data, determine the grouping binding strategy for SerDes channels with implemented symbol alignment, including:

[0094] Based on historical initialization data, a pre-trained adaptive learning optimization model is used to generate a grouping binding strategy for SerDes channels with already achieved symbol alignment.

[0095] A pre-trained adaptive learning optimization model is a computational model (e.g., based on machine learning or artificial intelligence algorithms) that first undergoes an initial training phase. Adaptive learning means that it can continuously adjust and optimize its decision-making capabilities based on newly acquired data in subsequent runs. Its ultimate goal is to find the optimal configuration for the system. Compared to simple lookup tables or fixed algorithms, this model can handle more complex scenarios and predict the current optimal strategy, fundamentally improving the intelligence level and long-term adaptability of the entire initialization method.

[0096] Optional, perform binding, including:

[0097] When the grouping binding strategy involves SerDes channels spanning different physical resource groups or chip dies, clock or delay compensation logic is automatically inserted based on the detection results.

[0098] This step addresses a specific and challenging scenario: when the model-generated strategy requires binding physically dispersed channels (i.e., across different physical resource groups or die dies) together, it automatically inserts clock or delay compensation logic. Crucially, this compensation action is based on the initial detection results (i.e., ensuring that intelligent decision-making can be physically reliably implemented, thus resolving signal timing deviations caused by physical distance).

[0099] Clock or delay compensation logic is a dedicated circuit or logic used to calibrate minute time differences caused by variations in signal transmission path length. Delay compensation typically adds a small delay to the first arriving signal, while clock compensation adjusts the phase of the clock for each channel to ensure that all channels sample data at the same time.

[0100] This step directly addresses the critical technical challenge of inter-channel clock or data skew caused by differences in physical wiring in multi-channel SerDes systems. By proactively compensating based on prior detection results, the method ensures that multiple physically dispersed channels can work synchronously and reliably as a whole, significantly improving the success rate and stability of wide-channel bonding in complex layouts.

[0101] In summary, the SerDes link initialization method provided in this application firstly detects multiple SerDes channels and marks each SerDes channel with a corresponding health level based on the detection results. This step enables the initialization process to identify channels in different states in advance. Compared with the prior art's approach of blindly aligning the entire system without distinguishing channel states, this provides a clear basis for subsequent alignment and grouping steps, avoiding wasting time on channels with known faults, thereby improving the targeting and efficiency of the initialization process. Secondly, for SerDes channels that meet the preset health level, single-channel symbol alignment and multi-channel collaborative alignment are performed in parallel. This parallel execution mechanism fundamentally changes the prior art's process of serially trying different alignment methods. By simultaneously initiating two alignment attempts, it is not necessary to wait for a timeout after one alignment method fails before trying another; instead, the faster method is used, directly shortening the total time required for link establishment and significantly improving the initialization speed. Finally, a grouping and binding strategy is determined based on historical initialization data. This introduces a dynamic, experience-based decision-making mechanism, replacing the fixed, preset grouping pattern in the prior art. This method enables the grouping strategy to adaptively adjust based on past success rates or failure modes, thereby selecting a grouping scheme with a higher success rate in the current scenario. This avoids repeated ineffective attempts caused by rigid strategies, improves the success rate of grouping and the utilization of hardware resources, and enhances the flexibility and adaptability of the method.

[0102] For example, a specific example will be used below to specifically illustrate the SerDes link initialization method of the above embodiments.

[0103] This example provides a SerDes link failure tolerance method, which specifically includes the following process:

[0104] First, upon detecting a power-on signal, reset command, or channel fault, the channel initialization process is immediately initiated, entering the global resource scanning phase to establish a basic triggering mechanism for subsequent link configuration. The triggering conditions for the entire method initialization are thus clearly defined.

[0105] Subsequently, a global resource scan and pre-detection process is performed in parallel. This step aims to understand the hardware resource status, mark the Lane health, and prepare for subsequent processes. Specifically, it includes the following steps:

[0106] Cross-Bank / Die Resource Check: Traverse all SerDesLanes within the FPGA, recording the Bank and Die to which each Lane belongs. Check for clock domain conflicts (e.g., whether different Banks share a PLL) and IO resource contention (e.g., pin multiplexing conflicts), and mark conflict-free Lane groups for priority use. If cross-Die Lanes exist, measure the physical delay of the wiring between dies and save the delay data (for subsequent compensation).

[0107] Parallel pre-testing across multiple lanes: Hardware self-tests are performed simultaneously on all lanes to check signal integrity (e.g., data bit error rate) and verify power stability (e.g., whether voltage fluctuations exceed thresholds). Lanes are categorized into three types: healthy, minor fault (repairable), and major fault (physical damage).

[0108] Next, lane health status is categorized. This step differentiates lanes in different states based on pre-detection results to avoid a complete restart for a single fault. Specifically, it includes the following steps:

[0109] Healthy lanes are directly added to the initialization candidate queue and given priority in subsequent symbol alignment and grouping.

[0110] For minor fault lanes, an automatic fast repair process is triggered, resetting the symbol alignment logic (clearing historical error states); fine-tuning the equalizer and pre-emphasis parameters (compensating for signal attenuation); after repair, the lane is marked as pending verification. If the verification passes, it is upgraded to a healthy lane; otherwise, it is downgraded to a major fault lane.

[0111] For severe fault lanes, mark them as standby lanes, temporarily skip the initialization phase, and only activate them when there are insufficient healthy lanes (for dynamic degradation).

[0112] Subsequently, multi-mode symbol alignment (parallel attempts) is performed. Two alignment modes are initiated simultaneously to reduce serial wait time and accelerate initialization. Specifically, the following steps are included:

[0113] Parallel startup of dual-mode alignment, performing the following two operations simultaneously on healthy lanes and lanes with minor faults after repair (without waiting for one to fail before switching):

[0114] Mode 1: Single Lane Fast Alignment. This mode employs a low-latency algorithm and prioritizes attempting to establish symbol alignment for each lane individually. If successful, the lane is marked as a single lane available.

[0115] Mode 2: Multi-Lane collaborative alignment, initiate multi-Lane binding preparation (such as 8-Lane pre-synchronization, 4+4 group pre-alignment); if successful, mark the Lane group as multi-Lane available.

[0116] Priority arbitration: If single-lane fast alignment succeeds first, the single-lane result is retained first; if multi-lane collaborative alignment succeeds first, switch directly to multi-lane mode; if one mode fails, seamlessly switch to another mode (without restarting the alignment process).

[0117] Next, the optimal group is dynamically selected based on historical data to adapt to different scenario requirements. This includes the following steps:

[0118] The self-learning decision model is invoked to read historical initialization data (such as fault type, grouping success rate, and bandwidth performance); the model predicts the optimal grouping strategy in the current scenario (such as direct 8-lane binding or splitting into 4+4 groups).

[0119] When performing group binding, if the strategy is 8-lane binding, global binding is performed on all healthy lanes, and cross-bank / die clock compensation logic (such as shared PLL synchronization, dynamic phase adjustment) is automatically inserted. If the strategy is 4+4 group binding, the lanes are split into two groups, binding is performed separately, and then the results are merged; for cross-bank / die lane groups, inter-die compensation (such as delay chain, phase fine-tuning) is enabled to offset routing delays.

[0120] Dynamic verification and timeout adjustment are implemented, flexibly adjusting the verification time based on lane quality to avoid misjudgments caused by a one-size-fits-all timeout. Specifically, this includes the following steps:

[0121] Dynamically set timeout thresholds: For high-quality lanes (low signal error rate, low clock jitter): set a short timeout (e.g., 100ms) to quickly complete verification; For low-quality lanes (e.g., long wiring, high jitter): set a long timeout (e.g., 500ms) to allow sufficient verification time.

[0122] Perform a channel integrity check to verify the content, symbol alignment status, data transmission and reception consistency, and whether the bit error rate meets the standard.

[0123] If the verification passes, the Lane group is marked as available; if it fails, it is reclassified.

[0124] Implement tiered retry and hot-insertion (fault self-healing) to differentiate fault types, prevent minor faults from slowing down the entire process, and support dynamic channel expansion. Specifically, this includes the following steps:

[0125] For minor faults, if the verification failure is due to symbol alignment drift (minor fault), directly restart the multi-mode symbol alignment and skip the full initialization; only reset the alignment logic and quickly retry.

[0126] For critical fault degradation and hot insertion, if the verification failure is due to physical disconnection (critical fault), the lane is marked as pending periodic detection, and communication is established using healthy subgroups first (e.g., if 4 lanes are successful, they are maintained; if 2 lanes are successful, they are added to the main channel). When the periodic detection (e.g., every 100ms) reaches "critical fault lane recovery", it is hot inserted into the current channel (without restarting initialization), and the bandwidth is dynamically expanded.

[0127] Finally, channel readiness and policy updates are performed to complete initialization and accumulate experience for the next process. This includes the following steps:

[0128] Record the channel status. All verified Lanes / groups enter the "ChannelReady" state and start normal communication.

[0129] Update the self-learning model and record the initialization data (such as fault type, grouping strategy, compensation parameters, and timeout settings);

[0130] The model automatically learns and optimizes, providing better strategies for the next initialization (such as adjusting grouping priority and timeout threshold).

[0131] In summary, the SerDes link fault tolerance method provided in this application achieves accurate identification and differentiated processing of faulty lanes through global resource scanning and lane health grading; it significantly shortens initialization time by employing parallel execution in both single-lane and multi-lane alignment modes and priority arbitration; it dynamically generates optimal grouping strategies using a self-learning decision model to adapt to different fault modes and hardware environments; it introduces hierarchical retry and hot-insertion mechanisms to support the maintenance of healthy subgroup communication and dynamic repair of faulty lanes; and it combines cross-bank / Die clock compensation and dynamic timeout adjustment to solve resource coordination and misjudgment problems, forming a closed-loop fault tolerance system of detection, decision-making, repair, and optimization, thereby improving the efficiency and reliability of high-speed communication links.

[0132] By employing dynamic grouping and intelligent fault-tolerant design, the efficiency and reliability of SerDes link initialization are significantly improved. Global resource scanning combined with Lane health grading enables accurate fault identification and differentiated handling, avoiding the inefficiency of full restarts in traditional solutions for a single fault. A dual-mode aligned parallel execution mechanism drastically shortens link establishment time, and a self-learning model dynamically adapts to the optimal grouping strategy in different scenarios, improving resource coordination efficiency. Layered retry and hot-insertion mechanisms ensure rapid repair of minor faults, uninterrupted communication during major faults, and dynamic capacity expansion after fault recovery. This solution effectively addresses the serial bottleneck and cross-resource coordination difficulties of traditional fixed grouping, achieving low-latency, high-reliability link transmission in high-speed communication scenarios, enhancing system stability and adaptability.

[0133] This application also provides a SerDes link initialization device for implementing the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0134] This application provides a SerDes link initialization device. Figure 2 This is a schematic diagram of a SerDes link initialization device provided in an embodiment of this application. The device includes:

[0135] The grading module 201 is used to detect multiple SerDes channels and mark each SerDes channel with a corresponding health level based on the detection results.

[0136] Alignment module 202 is used to perform single-channel sign alignment and multi-channel collaborative alignment in parallel for SerDes channels that meet the preset health level, so as to obtain SerDes channels with achieved sign alignment;

[0137] Grouping module 203 is used to determine the grouping binding strategy for the SerDes channel with implemented symbol alignment based on historical initialization data, and to perform binding according to the grouping binding strategy.

[0138] Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.

[0139] In this embodiment, the SerDes link initialization device is presented in the form of a functional unit. Here, a unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

[0140] This application also provides a computer device having the above-described features. Figure 2 The SerDes link initialization device shown.

[0141] Please see Figure 3 , Figure 3 This is a schematic diagram of the structure of a computer device provided in an optional embodiment of this application, such as... Figure 3As shown, the computer device includes one or more processors 10, memory 20, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components communicate with each other via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the computer device, including instructions stored in or on memory to display graphical information in a graphical user interface on an external input / output device (such as a display device coupled to the interface). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple computer devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 3 Take a processor 10 as an example.

[0142] Processor 10 may be a central processing unit, a network processor, or a combination thereof. Processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GDA), or any combination thereof.

[0143] The memory 20 stores instructions executable by at least one processor 10 to cause the at least one processor 10 to perform the method shown in the above embodiments.

[0144] The memory 20 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the computer device. Furthermore, the memory 20 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some alternative embodiments, the memory 20 may optionally include memory remotely located relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0145] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid-state drive; the memory 20 may also include a combination of the above types of memory.

[0146] The computer device also includes an input device 30 and an output device 40. The processor 10, memory 20, input device 30, and output device 40 can be connected via a bus or other means. Figure 3 Taking the example of a connection between China and Israel via a bus.

[0147] This application also provides a computer-readable storage medium. The methods described in this application can be implemented in hardware or firmware, or implemented as recordable on a storage medium, or implemented as computer code downloaded over a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and subsequently stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the methods shown in the above embodiments are implemented.

[0148] A portion of this application can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to this application through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.

[0149] Although embodiments of this application have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of this application, and all such modifications and variations fall within the scope defined by the appended claims.

Claims

1. A SerDes link initialization method, characterized in that, The method includes: Multiple SerDes channels are tested, and each SerDes channel is marked with its corresponding health level based on the test results; For SerDes channels that meet the preset health level, single-channel symbol alignment and multi-channel collaborative alignment are performed in parallel to obtain a SerDes channel with symbol alignment achieved; wherein, the single-channel symbol alignment includes establishing a symbol lock for each SerDes channel individually; the multi-channel collaborative alignment includes initiating multi-channel binding preparation, and the multi-channel binding preparation includes group pre-synchronization or group pre-alignment; Based on historical initialization data, a packet binding strategy for the SerDes channel with implemented symbol alignment is determined, and binding is performed according to the packet binding strategy; wherein, the historical initialization data includes at least one of fault type, packet power and bandwidth performance.

2. The method according to claim 1, characterized in that, The detection of multiple SerDes channels includes: Hardware self-tests are performed simultaneously on the multiple SerDes channels to obtain the signal integrity and power stability of each SerDes channel.

3. The method according to claim 2, characterized in that, The health level includes three levels: healthy, minor fault, and major fault. The step of marking each SerDes channel with its corresponding health level includes: SerDes channels with repairable signal integrity or power stability issues are marked as minor fault levels. SerDes channels with physical damage are marked as having a severe fault level.

4. The method according to any one of claims 1 to 3, characterized in that, The detection of multiple SerDes channels also includes: Traverse the multiple SerDes channels to determine the physical resource group or die to which each SerDes channel belongs, and check for clock domain conflicts or input / output resource contention.

5. The method according to claim 1, characterized in that, The parallel execution of single-channel symbol alignment and multi-channel collaborative alignment also includes: Based on the SerDes channel that was aligned first in the single-channel symbol alignment and the multi-channel collaborative alignment, the subsequent group binding steps are performed.

6. The method according to claim 1, characterized in that, The step of determining the group binding strategy for the SerDes channel with implemented symbol alignment based on historical initialization data includes: Based on historical initialization data, a pre-trained adaptive learning optimization model is used to generate a grouping binding strategy for the SerDes channels that have achieved symbol alignment.

7. The method according to claim 6, characterized in that, The execution binding includes: When the grouping binding strategy involves SerDes channels spanning different physical resource groups or chip dies, clock or delay compensation logic is automatically inserted based on the detection results.

8. A SerDes link initialization device, characterized in that, The device includes: The grading module is used to detect multiple SerDes channels and mark each SerDes channel with a corresponding health level based on the detection results. The alignment module is used to perform single-channel symbol alignment and multi-channel collaborative alignment in parallel for SerDes channels that meet the preset health level, so as to obtain SerDes channels with achieved symbol alignment; the single-channel symbol alignment includes establishing a symbol lock for each SerDes channel individually; the multi-channel collaborative alignment includes initiating multi-channel binding preparation, which includes group pre-synchronization or group pre-alignment. The grouping module is used to determine the grouping binding strategy for the SerDes channel with implemented symbol alignment based on historical initialization data, and to perform binding according to the grouping binding strategy; the historical initialization data includes at least one of fault type, grouping power and bandwidth performance.

Citation Information

Patent Citations

  • Method and system for aligning high speed serial communication channels

    CN102708080A

  • Multi-link data real-time alignment method and system and storage medium

    CN116149598A