Data channel time bias removal and rate adaptation within package structure containing multiple circuit dies

Through the PCIe retimer circuit, the alignment code symbols in the multi-channel FIFO are detected and synchronized, and the de-time bias and rate adaptation between channels are achieved, which solves the problem of shortening the channel range in PCIe 5.0, and improves data transmission efficiency and system flexibility.

CN120303650APending Publication Date: 2025-07-11KANDOU LABS SA

Patent Information

Application Number
CN202380083351.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-18
Filing Date
2023-10-18
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

With the increase in PCIe 5.0 data rate, shortening of channel range has led to an increase in the demand for retimers, and the prior art is difficult to effectively compensate for signal attenuation and reduce noise and jitter. At the same time, there are challenges in achieving channel-to-channel time bias and rate adaptation in multi-chip systems.

Method used

The PCIe retimer circuit is adopted to detect the alignment code symbols in the FIFO of multiple data channels, generate the core write clock, and synchronize logic under the common reference clock to realize the de-time bias and rate adaptation between channels, and use FIFO for data storage and processing, and combine high-speed inter-chip interconnection to achieve efficient communication.

Benefits of technology

Effectively compensate time deviation between channels, reduce noise and jitter, improve data transmission efficiency, adapt to different speed requirements, reduce system power consumption, and realize highly flexible data routing in multi-chip systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120303650A_ABST
    Figure CN120303650A_ABST
Patent Text Reader

Abstract

Methods and systems described herein are used for multi-channel alignment and rate adaptation between cores (1304, 1302) of a multi-core package structure (1300), in particular for exchanging alignment information (algncount, rpcsalgnctl) across clock domains of different cores (1304, 1302) based on a core write clock (wrtileclk) generated from a local system clock (txclk) within a main core (1302), the period of the core particle write-in clock (wrticuclock) is equal to the period of a common reference clock (refclk), and the position of a pulse corresponding to the core particle write-in clock (wrticuclock) in the period of the common reference clock (refclk) is determined by the effective period of a counter.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims the benefit of U.S. Application No. 63 / 380,045, filed on October 18, 2022, entitled "Data Channel De-Skewing and Rate Adaptation in Asynchronous FIFOs of Multi-Channel PCIe Retimers", and claims the benefit of U.S. Application No. 63 / 380,042, filed on October 18, 2022, entitled "Data Channel De-Skewing and Rate Adaptation in a Package Structure Containing Multiple Die", the entire contents of both of which are incorporated herein by reference for all purposes.

[0003] References

[0004] The following references are incorporated herein by reference for all purposes:

[0005] U.S. Patent No. 9,100,232, issued on August 4, 2015, entitled "Method for Evaluating Codes by ISI Ratio", inventor Amin Shokrollahi, application number 14 / 612,241, filed on February 2, 2015, publication number 2015 / 0222458, published on August 6, 2015, hereinafter referred to as [Shokrollahi]. Background Art

[0006] As the data rate of PCIe 5.0 (32 Gbps) is increased compared to previous generations (e.g., the maximum data rate of PCIe 4.0 is 16 Gbps), the channel range becomes shorter than before, and the need for retimers becomes more prominent. Common channels include system boards, backplanes, cables, adapter cards, and add-in cards. The losses generated by connections across such channels (usually in the form of combinations of such channels and their slots) often exceed the target loss requirement of -36 dB at 16 GHz. A retimer can extend the channel range beyond the boundary of the operating range without using a retimer.

[0007] The retimer divides the link between the host (root complex, abbreviated as RC) and the device (endpoint) into two independent segments. Therefore, the retimer can re-establish a new forward PCIe link, and this re-establishment process includes re-training and appropriate equalization processing for implementing the physical layer and the link layer.

[0008] The re-timer is a pure analog amplifier that compensates for signal attenuation by boosting the signal. However, while boosting the signal, the re-timer increases noise, which often exacerbates jitter. In contrast, the re-timer incorporates both analog and digital logic, equalizes the signal, extracts its clock signal, and outputs a signal with a large amplitude and low noise and jitter. In addition, the re-timer can maintain the power state to keep the system power consumption at a low level.

[0009] The specifications of the re-timer were first established in PCIe 4.0 and are expected to continue to be used in PCIe 5.0. Figure 1A and Figure 1B The following shows common application scenarios of the re-timer in some embodiments. Figure 1A Employ one re-timer, place it on the motherboard, and logically position it between the PCIe root complex (RC) and the PCIe endpoint device.

[0010] Figure 1B The following shows the case of using two re-timers. Among them, the first re-timer is also placed on the motherboard, while the second re-timer is placed on an adapter card that connects the motherboard and an additional card, and the PCIe endpoint device is included in the additional card.

[0011] In a complex PCIe system, the number of PCIe endpoint devices may be much larger than the number of idle PCIe ports. In such cases, a switching device can be used to increase the number of PCIe ports. The switch can connect multiple endpoint devices to the same root node and route data packets only to the specified destination instead of mirroring the data to all ports. An important feature of the switch is bandwidth sharing, and all endpoint devices can share the bandwidth of the root node. SUMMARY OF THE INVENTION

[0012] This "Summary of the Invention" section is a brief description of a series of concepts detailed in the following "Detailed Description" section. The purpose of this "Summary of the Invention" section is not to identify the key or main features of the claimed technical solution, nor to assist in determining the scope of the claimed technical solution. For those of ordinary skill in the art, other objectives and / or advantages of the embodiments of the present invention will become readily apparent by referring to the "Detailed Description" section and the accompanying drawings.

[0013] In the methods and systems described herein: Detect alignment characters in FIFOs of multiple data channels of multiple die, the multiple die including a master die and one or more slave dies; Determine that alignment characters have been detected in FIFOs of all channels of all die, and generate an alignment character found signal accordingly; Generate a die write clock according to a local system clock, the period of the die write clock being equal to the period of a common reference clock, the die write clock corresponding to a pulse, the position of the pulse within the period of the common reference clock being determined by the active period of a counter; Send the alignment character found signal to the synchronization logic of each of the slave dies according to the die write clock; Sample the alignment character found signal in each slave die and the master die according to the common reference clock by the synchronization logic; For each die of the multiple die, synchronize the alignment character found signal with a locally generated system clock, and set the read pointer of the FIFO to the position containing the alignment character accordingly; And output data from each FIFO according to the locally generated system clock. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1A and Figure 1B Shows two usages of a retimer in some embodiments.

[0015] Figure 1C Is a block diagram of a retimer data path in some embodiments.

[0016] Figure 2 Is a block diagram of three configurations of a routing channel between retimer ports in some embodiments.

[0017] Figure 3 Is a block diagram of two possible dual-die combinations within the same package structure in some embodiments.

[0018] Figure 4 Is a block diagram of a four-die combination within the same package structure in some embodiments.

[0019] Figure 5 Is a block diagram of another four-die combination within the same package structure and using die-to-die communication in some embodiments.

[0020] Figure 6 Is a block diagram of a high-speed die-to-die interconnect in some embodiments.

[0021] Figure 7 Is a block diagram of a crossbar switch in some embodiments.

[0022] Figure 8 Is a block diagram of a system for channel de-skewing in some embodiments.

[0023] Figure 9 The de-skewing timing diagram for the minimum skew scenario in some embodiments.

[0024] Figure 10 The de-skewing timing diagram for the common skew scenario in some embodiments.

[0025] Figure 11 The de-skewing timing diagram for the maximum skew scenario in some embodiments.

[0026] Figure 12 The system block diagram for rate adaptation in some embodiments.

[0027] Figure 13 The block diagram of a multi-die communication system for inter-channel alignment between dies in some embodiments.

[0028] Figure 14 The timing schematic diagram of a die clock generator in some embodiments.

[0029] Figure 15 The FIFO filling level diagram during the rate adaptation process.

[0030] Figure 16 The schematic diagram of a multi-die communication system for rate adaptation in some embodiments.

[0031] Figure 17 The information exchange timing diagram in multi-die rate adaptation in some embodiments.

[0032] Figure 18 The flowchart of method 1800 in some embodiments. Detailed implementation manners

[0033] Although the technical ability to fully integrate multiple systems into the same integrated circuit is increasing day by day, the practice of separately maintaining multiple chip systems and subsystems still has significant advantages. For non-limiting description purposes, in at least some aspects of the present invention described herein, exemplary embodiments employ a system environment consisting of at least one point-to-point communication interface that connects two integrated circuit chips, respectively representing a root complex (i.e., a host) and an endpoint device, wherein the communication interface is supported by a number of data channels, and each channel consists of four high-speed transmission line signal conductors.

[0034] A retimer generally includes a PHY and a retimer core logic. The PHY includes a receiver part and a transmitter part. The PHY receiver is used for data recovery, data deserialization, and clock recovery, while the PHY transmitter is used for data serialization and amplification processing before output transmission. The retimer core logic is used for de-skewing (of a multi-channel link) and rate adaptation to compensate for the frequency difference between the two side ports.

[0035] Since the retimer is located in the path between the root complex (such as a CPU) and the endpoint device (such as a cache block), the retimer has additional value. An integrated processing unit (such as an accelerator) can be integrated within the retimer to perform data processing in the root complex to endpoint device path.

[0036] To achieve a highly flexible solution, the PCIe retimer uses a conventional PHY interface for the PCIe bus and a high-speed die-to-die interconnect for the data processing unit (DPU). The high-speed die-to-die interconnect can achieve an extremely high-speed communication link between chiplets within the same package structure. The PCIe retimer circuit is a chiplet (die) with a four-channel retimer and is capable of connecting to the DPU chiplet through the high-speed die-to-die interconnect. Among them, one, two, or four channels can form a multi-channel link through which data can be transmitted across all links. Alternatively, each channel can also be configured as a single-channel link. Each channel within the PCIe retimer has two PHYs, which are respectively located at both ends (i.e., the upstream port and the downstream port). Since the total number of channels is four, a PCIe retimer die has eight PHYs. In addition, the PCIe retimer die also has communication lines to enable the exchange of control information between two or more PCIe retimer dies.

[0037] Through one (or more) PCIe retimer chiplets, the following structures can be constructed (which will be further described in detail below):

[0038] - Four-channel retimer;

[0039] - A single die with fully flexible 4×4 static channel routing;

[0040] - Four-channel retimer with an accelerator (DPU);

[0041] - Two dies encapsulated within the same package structure: a retimer die and a DPU die;

[0042] - Eight-channel retimer;

[0043] - Two dies encapsulated within the same package structure, but only with limited static channel routing - high flexibility 4×4 routing within the same die, but without the function of enabling data to cross die boundaries;

[0044] - Eight-channel retimer with fully flexible channel routing;

[0045] - Two dies encapsulated within the same package structure, achieving data routing across different chiplets through the high-speed die-to-die interconnect, but resulting in additional latency;

[0046] - An eight-channel retimer with an accelerator (DPU);

[0047] - Three dies encapsulated within the same package structure: two retimer dies and one DPU die;

[0048] - A sixteen-channel retimer;

[0049] - Four dies encapsulated within the same package structure, but with only limited static channel routing - high flexibility 4×4 routing within the same die, but without the function of enabling data to cross die boundaries.

[0050] Figure 1C Shown are in some embodiments Figure 1A and Figure 1B the data paths of the PCIe retimer circuits shown. Figure 1C The retimer data paths are applicable to both single-tile and multi-tile embodiments. Two possible transfer schemes for transferring data from the receiver to the transmitter include: storing the encoded data in a FIFO; and storing the decoded data in a FIFO.

[0051] In the case of storing the encoded data in a FIFO, the received data of the PHY can be encoded in 8b10b or 128b130b. This data can be divided into 16-bit or 32-bit data blocks at any position in the data stream. In this mode, the received data is directly forwarded and stored in the FIFO. At the same time, the data is decoded by the data block detection and data block alignment circuit. Among them, the data block boundary that can accurately identify the position of the ordered set (i.e., the data block) in the received data stream is stored in the FIFO as sideband information. To compensate for the processing delay caused by data block alignment, a pipelining stage can also be added. In the de-skew processing after the FIFO, the barrel shifter aligns each data block so that they have the same starting position. Since the data stream contains synchronization header bits, it can be transmitted to the transmitter without further modification.

[0052] In the case of storing decoded data in a FIFO, the data block detection and alignment logic directly decodes the received data into 8B or 128B data blocks. Among them, after extracting overhead information such as control / data type identification information (8b10b) or synchronization header information (data block start position, ordered set type in 128b130b) from the data, it is used as sideband information together with the decoded data and stored in the FIFO. The data is aligned with the ordered set boundary according to its nature in the FIFO, and the corresponding de-skew processing includes moving the FIFO read pointer to the corresponding aligned symbol storage position. When forwarding the FIFO read data to the transmitter, it is also necessary to insert synchronization header bits into the data stream, and the removal and insertion of synchronization header bits generally result in idle cycles. It should be noted that there is not much difference in the insertion / removal of SKP ordered sets between the two 128b130b modes. The length of the input SKP ordered set is at least 12 symbols, among which the first eight symbols (64 bits) contain the same bytes. At any position, 32 bits can be extracted from these eight symbols. As Figure 1C shown, the encoded data is always stored in the FIFO. In the following sections, the description of data blocks all relates to this data transfer mode.

[0053] Chip Structure

[0054] Figures 2 to 5 Shows various structures of the PCIe retimer circuit from the perspective of the data stream in some embodiments. The maximum number of dies contained in each package structure shown is four. Figure 2 Shown is an optional three-way routing scheme for a package structure containing a single die. Such embodiments can be used as four-channel PCIe retimers. Within the same circuit die, all data is routed from one port to another port through the channel routing logic. The Raw MUX independently routes each channel between the ports. The package structure 200 shows a feed-through path, the package structure 205 shows a twisted path, and the package structure 210 shows a port mirror. Specifically, only one direction is shown in the package structure 210, and there is another mirror in the opposite direction. In some embodiments, the serializer / deserializer (SD) above each structure diagram can be connected to an upstream device such as a root complex, for example, and the serializer / deserializer below each structure diagram can be arranged to be connected to an endpoint device, for example; vice versa.

[0055] Figure 3 Shown are two possible combinations of dual dies within the same package structure. The package structure 305 can correspond to an eight-channel PCIe retimer with the minimum delay. The price of the minimum delay is that a routing structure is adopted in which each channel is routed between the upstream and downstream ports within the same die. The de-skew information is exchanged between the two dies through a communication link to implement channel de-skew across all eight channels.

[0056] Figure 3 The encapsulated structure 310 shown may correspond to an eight-channel PCIe retimer circuit with full routing flexibility across the circuit die. The cost of full flexibility is the additional latency and power consumption due to die-to-die (D2D) interconnects. The original data MUX in each PCIe retimer circuit die is either routed directly to the opposite port (as shown at 305) or routed via a high-speed die-to-die interconnect (as shown at 310). When routed via the high-speed die-to-die interconnect, data can be passed to an adjacent die. In this case, inter-channel de-skewing is performed directly on one of the dies without exchanging de-skewing information between the chips.

[0057] Figure 4 Shown is an encapsulated structure 400 with four dies. Such an encapsulated structure can serve as a sixteen-channel PCIe retimer circuit. In such embodiments, de-skewing information is exchanged between the four dies via communication links to perform inter-channel de-skewing across all 16 channels. In such embodiments, D2D interconnects are not used.

[0058] Figure 5 Shown is another four-die combined encapsulated structure 500. Such a configuration can implement a sixteen-channel PCIe retimer circuit with highly flexible 2×8 channel routing. As shown, data routing between the pair of circuit dies on the left and the pair of circuit dies on the right can achieve full routing within eight channels, but with increased latency.

[0059] Figure 6 A block diagram of the high-speed die-to-die interconnect for some embodiments is shown. As shown, the high-speed die-to-die interconnect uses eight transceiver paths, each operating at a speed of 25 GBd, for transmitting 5 bits over 6 wires, with a total throughput of 125 Gbps. In addition, the interface includes two differential clock channels operating at a speed of 6.25 GHz. The high-speed die-to-die interconnect can use the 5b6w code described by

Shokrollahi

[0060] Figure 7 A block diagram of a channel-switching multiplexer (MUX) (also referred to herein as a cross-switch 700 or channel routing logic) for performing channel routing in the retimer circuit die of an integrated chip module (ICM) in some embodiments is shown. Figure 7It includes a left block diagram and various channel routing diagrams on the right. In the top channel routing structure 705, after the data is fed in through a deserialiser, it enters the PHY. After passing through the core logic and the PHY, it is output to the bottom through a serializer. In the middle diagram 710, the data is fed into a port, processed by the core logic, and then fed out at the corresponding PHY at the bottom. In the bottom diagram 715, all the data is fed into the PHYs on the top side of a PCIe retimer circuit and directly forwarded from these PHYs to the high-speed die-to-die interconnect. Further, after passing through the core logic from the high-speed die-to-die interconnect, the data is fed to the PHYs on the bottom side of another PCIe retimer die. In all these cases, there are further data paths in the opposite direction.

[0061] Figure 7 The left side shows a schematic diagram of the original data multiplexer logic. Each serial data transceiver PHY is numbered from 0 to 7 and includes a receiver deserialiser (DES) and a transmitter serializer (SER). Among them, the top channels (PHY0 and PHY4) show three different data paths corresponding to the respective data paths shown on the right. Figure 7 The right data path 705 corresponds to the data being input from the PHY0 of the PCIe retimer circuit and output from this PHY0 along the left path. Path 710 corresponds to the data being received by PHY0 and passing through to Figure 7 The feed-through path on the left. Path 715 corresponds to the following path: all the received data is directly forwarded to the adaptation layer for transmission via the die-to-die data interface; the data from the die-to-die data interface is forwarded by a second PCIe retimer to the core logic; after being processed by the core logic, the data is output by the attached PHY.

[0062] The second channel (PHY1 and PHY5) is used for the multiplexing function. Among them, each core logic / transmitter path can receive data from each of the eight channels or obtain data from the die-to-die data interface. The other channels (PHY0 and PHY4, PHY2 and PHY6, PHY3 and PHY7) have the same switching function. The bottom shows the multiplexing between one of the channels and the die-to-die data interface. For each channel connected to the high-speed die-to-die interconnect, any input PHY can be selected. Therefore, in some embodiments, data mirroring can be achieved by selecting the same received PHY data for multiple adaptation layer physical ports. Details of the port mirroring implementation will be further described in detail below.

[0063] Data path switching in the raw data multiplexer includes: receiving deserialized channel-specific data codewords on a 32-bit data bus; enabling corresponding data lines; clock recovery; and corresponding reset operations. It should be noted that only the raw data is multiplexed, and the received data is not processed. The raw data multiplexer logic is statically configured by configuration bits, and the switching operation is an asynchronous operation. If the settings of the raw data multiplexer are changed in the task mode, it may result in invalid data and may cause glitches on the clock line. Therefore, the multiplexing logic can be set during the reset process.

[0064] Multi-channel de-time skew

[0065] The de-time skew processing and the rate adaptation processing are related to each other and are performed by the same module (de-time skew and rate adjustment control module). First, the inter-channel time skew is compensated. This processing is also called channel alignment processing and is generally completed by a FIFO. Among them, the alignment code symbols in the data stream are first detected. Due to the time skew between channels, the received times of the alignment code symbols of each channel are different. In the de-time skew processing, in addition to storing the received alignment code symbols in the FIFO, the positions of these alignment code symbols in the FIFO are further saved. All channels perform this processing separately with their respective recovered clocks. After the alignment code symbols of all channels are stored in the corresponding FIFOs, data is read from the FIFOs, where the reading starts from the reading pointer that determines the storage position of the alignment code symbols. At the reading end, the FIFO reading pointers of each line are set at the storage position of the alignment code symbols. The reading pointers of all channels are set simultaneously with the same clock so that the first data output by each FIFO corresponds to the alignment code symbol. For each channel, the filling degree of its FIFO is monitored, and according to this filling degree, the following operations are performed: when the FIFO filling degree is almost "empty", special data for rate adaptation is inserted; when the FIFO filling degree is almost "full", such data is removed. Among them, rate adaptation code symbols are used for this purpose. When such rate adaptation code symbols exist simultaneously in all channels (i.e., the case where the de-time skew processing is completed), the removal or copying (insertion) of this data can be performed simultaneously in all channels. Below, the rate adaptation processing is further described in detail.

[0066] The following solves the following problems: In the retimer mode, all transmitters are synchronized to the same common reference clock; however, within the retimer, each data channel often has its own read clock and there is no common read clock. The read clock basically corresponds to the transmit clock of the attached serializer. In addition, there is also the problem of exchanging alignment status and FIFO status between all dies in a multi-die system through low-speed I / O pads, which will be described in further detail below. The methods and systems described below achieve the fusion of inter-channel de-skewing and multi-channel rate adaptation in a common FIFO for each channel by synchronously changing the FIFO read pointer for each channel.

[0067] Figure 8Block diagram of channel alignment logic 800 for implementing the channel de-skew concept in a PCIe retimer circuit in some embodiments. In some embodiments, a method for performing channel de-skew includes: separately detecting, by alignment symbol detection logic 805, alignment symbols within the first-in-first-out (FIFO) buffer 820 of each data channel according to the recovered clock signal rx_clk; and generating a single-cycle pulse rx_algn as the alignment symbols are detected. After storing the alignment symbols, the position within the FIFO is also stored as a write pointer, which may further include: storing the bit-level start position of the alignment symbols in the 32-bit position of the FIFO. In some embodiments, the alignment symbols are 32-bit symbols. Note that since the encoded data is stored in the FIFO, the data block boundaries continuously change, and the data block boundaries of the alignment symbols need to be stored simultaneously. The method further includes: separately generating, for each data channel, a channel-specific alignment symbol present pulse rx_algn_str, e.g., by broadening the alignment detection pulse rx_algn with pulse broadening logic 815, to indicate that the corresponding alignment symbol is stored in the FIFO. In some embodiments, the pulse length depends on the required de-skew capability. For example, when the maximum de-skew capability (within a clock cycle) is N, the pulse length L = ceil(N + 2). In such embodiments, N is determined by the maximum input skew plus the skew introduced by the deserialiser (equal to the bit width, i.e., 1UI) and the skew introduced by the synchroniser. The broadened alignment pulses of all data channels are asynchronously combined by the die-specific AND gate 810 to indicate that alignment symbols have been stored in the FIFOs of all data channels of the die. Note that since the read clocks of the individual data channels are independent of each other, the AND combination operation is performed before the synchronisation operation. In some embodiments, to prevent glitches at the input of the synchroniser, the AND combination consists of instantiated tech cells. For each data channel, the ANDed signal rx_algn_comb is synchronised with the aligned transmit clock tx_clk. The rising edge of the synchronised signal is detected. The above alignment pulse rx_algn_str is broadened by an additional two clock cycles compared to the required de-skew capability to ensure that the duration corresponding to the remaining pulse width is at least two clock cycles even in the case of maximum skew between two data channels. This pulse width is sufficient to achieve a safe clock domain crossing since the clock domain needs to transition from the receiver's rx_clk to the transmitter's tx_clk. In some embodiments, in retimer mode, the FIFO read clocks of all channels are aligned with a common reference clock. Among them, the read pointer of the FIFO is set equal to the stored write pointer using a single-cycle rising-edge pulse output by the alignment control finite state machine (FSM) 825, thereby setting the current read position of the FIFO to the alignment symbol position.Since the rising edge pulses of all FIFOs are synchronized, the read pointers of all FIFOs are updated simultaneously. Since the encoded data is stored in the FIFOs, the alignment operation may include: adjusting an internal barrel shifter to address different data block boundaries for different channels. Additionally, since the read clocks are individually aligned with a common reference clock, there may still be a minimum time skew between data channels that is equal to a single clock cycle. According to the transmit time skew budget specified in the PCIe base specification, this time skew is within an acceptable range. The alignment operation may cause an interruption in the data stream transmitted downstream. In some embodiments, the interruption may be made acceptable by configuring a bit to select between outputting a fixed pattern (such as a high-speed 1010 pattern) and outputting the previously received data. After channel deskewing, the read operation continues for each FIFO, and all FIFOs output aligned codewords simultaneously. In some embodiments, the effective read position of the FIFO may be adjusted by a barrel shifter such that reading starts from the sync header bit of the aligned codeword. Since the encoded data is stored in the FIFO, the starting position of the aligned ordered set can be any position in the FIFO. For example, in one channel, the starting position of the aligned ordered set may be bit 3; in another channel, the starting position of the aligned ordered set may be bit 19; in yet another channel, the starting position of the aligned ordered set may be bit 11. After alignment, the first bit of the aligned ordered set necessarily starts from bit 0. The barrel shifter can shift all bits in the codeword by a number of bits. In the above example, the data of the first channel may be shifted by 3 bits, the data of the second channel may be shifted by 19 bits, and the data of the third channel may be shifted by 11 bits. It should be noted that since the sync header bit itself is part of the data stream, no further action is required in the case of a 128b130b encoded data stream.

[0068] Figures 9 to 11 Three timing diagrams for channel deskewing in various time skew amount scenarios. In Figure 9 , there is a minimum time skew amount between two data channels. In Figure 10 , there is a common or medium time skew amount (in this case, specifically about 2.7 clock cycles) between two data channels. In Figure 11 , there is a large time skew amount (in this case, about five clock cycles) between two data channels.

[0069] Rx_clkX and rx_dataX are the recovered clocks and received data lines for channels 1 and 2 respectively, which can specifically be the FIFO write clock and data. Rx_algnX is a pulse indicating the presence of an alignment character (A). Rx_algnX is also used to trigger the storage operation of the FIFO write pointer. Rx_algnX_str is the widened pulse, which in these examples is widened by an additional six clock cycles. Rx_algn_comb is the "AND" combination of all rx_algnX_str signals for all data channels. Tx_clk is the transmission clock, i.e., the FIFO read clock. Tx_algn_comb_g1 and Tx_algn_comb_g2 are the "AND" combination signals after synchronization (by the first and second synchronous flip-flops). Tx_algn_found is the decoded rising edge of tx_algn_comb_g1, which is used to set the FIFO read pointers for channels 1 and 2. Tx_data1 and Tx_data2 signals are the FIFO output data sent to the transmission logic.

[0070] Rate adaptation

[0071] After de-time skew, rate adaptation is performed. During the rate adaptation process, the FIFO filling level is monitored, and based on this filling level, the following operations are carried out: when the FIFO filling level is "empty", skip (SKP) ordered set characters for rate adaptation are inserted; when the FIFO filling level is "full", the SKP ordered set characters are removed. Since each data channel has been de-time skewed, the rate adaptation characters for all channels can be examined simultaneously, and they can be removed or copied (inserted) simultaneously within all channels. The purpose of rate adaptation can be to keep the current filling level of each data channel FIFO within an acceptable range to prevent overflow or underflow.

[0072] Figure 12 Is a block diagram of the rate adaptation logic 1200 in some embodiments. In Figure 12In response to the skip code symbol detection logic 1205 detecting a skip code symbol, a single-cycle pulse wr_skp is issued. This pulse is issued separately for each channel within the recovered clock (FIFO write clock rx_clk). This skip pulse is fed as sideband information to the FIFO 820 and stored at a location one storage position ahead of the corresponding skip code symbol (rather than stored together with the skip code symbol itself). As shown, rate adaptation uses the same FIFO as the channel de-time skew 820. However, in some embodiments, these two functions may also use different FIFOs respectively. In some embodiments, by performing both inter-channel de-time skew and rate adaptation operations with the same buffer, the overall delay of the retimer path can be reduced. At the FIFO read end, the removal (also referred to as "clipping") of the skip code symbol can be achieved by knowing whether there is skip information at a time point one clock cycle earlier than the read time of the skip code symbol, or the insertion (also referred to as "padding") of the code symbol can be achieved, for example, by reading the skip code symbol twice.

[0073] As described above, the rate adaptation FSM 1210 monitors the fill level of all FIFOs. When any FIFO gives an indication that the FIFO is "full", a "clipping" operation is performed to simultaneously remove the skip code symbols of all FIFOs. In some embodiments, the removal of the skip code symbol is achieved by doubling and incrementing the FIFO read pointer within one clock cycle. Similarly, when any FIFO gives an indication that the FIFO is "empty", a "padding" operation is performed. In this case, skip code symbols are inserted into all FIFOs simultaneously. In some embodiments, the existing code symbol is read twice and the FIFO read pointer is not incremented within one clock cycle.

[0074] The insertion or removal of the skip code symbol is performed only when there are skip code symbols in the FIFO memory. The skip sideband information takes effect at a time point one clock cycle earlier than the actual read and output time of the skip code symbol to trigger the padding or clipping operation. The skip indication information dec_ptr, inc_ptr appears in all FIFOs simultaneously. If the skip indication information does not appear in all FIFOs simultaneously, a rate adaptation error message ( Figure 12 ra_err in

[0075] In some embodiments, when the FIFO pointer wraps around to the FIFO start position, a flag information is issued. After synchronizing the flag information to the FIFO read end, the value of the FIFO read pointer is evaluated. In some embodiments, the value of the read pointer is evaluated by synchronizing the MSB of the FIFO write pointer to the FIFO read end that performs a rising edge detection on the synchronized signal. To prevent data loss, the FIFO always stores all data until it can perform rate adaptation. Among them, in the worst case, just after the skip symbol is passed into the FIFO, an indication information of "FIFO full" or "FIFO empty" is issued, so that at least one more codeword is stored before the next skip symbol arrives. In some embodiments, the skip symbols are not equally spaced, and the size of the FIFO is increased accordingly. To avoid the FIFO write pointer and the read pointer converging to one place with each other (thus causing the data read by the FIFO read end to be unstable), an indication information of the FIFO filling degree can be further provided. In one of the cases, if the FIFO is "full" and no rate adaptation operation for reducing the FIFO filling degree is performed, a FIFO overflow indication information as an error flag information is issued. In another case, if the FIFO is "empty" and no rate adaptation operation for increasing the FIFO filling degree is performed, a FIFO underflow indication information as an error flag information is issued

[0076] In each 128b130b mode (the third / fourth / fifth generation PCIe), the object of rate adaptation is a 32-bit data block. Since the synchronization header bits themselves are part of the data stream, the length of the ordered set is not an integer multiple of 16 or 32, resulting in the exact position of the skipped ordered set being variable. Accordingly, the ordered set boundary problem can be solved by inserting or removing 32-bit data blocks. In some embodiments, by storing the synchronization header bits as sideband information, the ordered set boundary can be kept unchanged.

[0077] As described above, by integrating the inter-channel de-time skew function and the rate adaptation function within the same FIFO, data transmission within each channel can be achieved through a single FIFO rather than multiple FIFOs, thereby reducing latency. A device includes alignment symbol detection logic 805 for detecting alignment symbols within first-in-first-out buffers (FIFOs) 820 of multiple data channels of a data link and storing FIFO addresses corresponding to the positions of the alignment symbols within each FIFO. The device further includes an alignment control finite state machine (FSM) 825 for synchronously adjusting the read pointer positions of each FIFO 820 to the stored FIFO addresses corresponding to the positions of the alignment symbols within the FIFO upon detection of alignment symbols in all data channels. The device further includes skip symbol detection logic 1205 for detecting skip ordered sets (SKPs) within each FIFO 820 and storing SKP pulses within each FIFO 820 at a position one address earlier than the SKP, each SKP including two or more SKP symbols. The device further includes a rate adaptation FSM 1210 for: monitoring the fill levels of each FIFO of the multiple data channels; queuing rate adaptation events upon the fill level of at least one FIFO exceeding a threshold; and executing rate adaptation events by operating the read pointer according to rate adaptation events upon reading of the SKP pulses of all data channels.

[0078] In some embodiments, a method includes: detecting alignment symbols within first-in-first-out buffers (FIFOs) 820 of multiple data channels of a data link and storing FIFO addresses corresponding to the positions of the alignment symbols within each FIFO. Upon detection of alignment symbols in each data channel, synchronously adjusting the read pointer positions of each FIFO to the stored FIFO addresses corresponding to the positions of the alignment symbols within the FIFO. The method further includes: detecting skip ordered sets (SKPs) within each FIFO 820 and storing SKP pulses within each FIFO 820 at a position one address earlier than the SKP, each SKP including two or more SKP symbols; monitoring the fill levels of each FIFO of the multiple data channels; queuing rate adaptation events upon the fill level of at least one FIFO exceeding a threshold; and executing rate adaptation events by rate adaptation logic operating the read pointer according to rate adaptation events upon reading of the SKP pulses of each data channel.

[0079] In some embodiments, the fill level of at least one FIFO exceeds the "full" threshold, and the rate adaptation event is a skip event for incrementing the read pointer of each FIFO of the plurality of data channels as the read pointer of each FIFO reaches the SKP address to remove SKY symbols from all data channels. Similarly, the fill level of at least one FIFO exceeds the "empty" threshold, and the rate adaptation event is a padding event for keeping the read pointer of each FIFO of the plurality of data channels unchanged within a clock cycle as the read pointer of each FIFO reaches the SKP address to insert SKP symbols into all data channels. In some embodiments, SKP pulses are stored as sideband information in each FIFO.

[0080] In some embodiments, synchronously adjusting the read pointer position of each FIFO to the stored FIFO address corresponding to the aligned symbol position within that FIFO further comprises: receiving an aligned symbol found signal.

[0081] Multi-die deskew

[0082] The embodiments described herein provide an efficient PCIe retimer circuit that can configure a multi-die package structure into one of the several configurations described above. Accordingly, the methods and systems described herein provide solutions for performing both channel deskewing and rate adaptation across multiple dies according to the configuration under constraints such as transmitting signals through low-speed I / O pads. In a single-die implementation, the deskewing information and FIFO status information for rate adaptation can be exchanged at the maximum speed between two or more channels (in a single-die implementation, the maximum number of channels is four) (clock frequency is 1 GHz). However, a multi-die implementation uses another method. In a multi-die implementation, deskewing and FIFO status / rate adaptation information are exchanged between two or four dies through low-speed I / O pads. That is, the number of information exchange lines needs to be minimized. Figure 13 Shown is an integrated multi-die circuit module for multi-die channel deskewing by exchanging deskewing information in some embodiments. To minimize the connection lines, several factors can be considered. First, since the branching requirements of the multi-die retimer are limited, the alignment requirements of the multi-die structure are also limited. The object of the alignment operation between multi-dies is channels whose number is a multiple of four. For an eight-channel retimer, there is no need to support the 2-4-2 channel branching scheme, and only the 8, 4-4, 4-2-2, or 2-2-4 scheme (i.e., the operation amount is three times that of the single-die scheme) or eight channels (i.e., 4 channels distributed on two dies) need to be supported. Similarly, for a sixteen-channel retimer, the cross-die branching methods supported are 16, 8-8, 8-4-4, and 4-4-8.

[0083] For a given die, since all data channels either operate separately or are combined with each other into larger links, the amount of alignment information exchanged in each direction for each pair of master and slave dies is one bit. One bit from the slave die to the master die indicates that alignment codewords exist in the de-skew FIFOs of all channels of the slave die. Figure 13 In Figure 13 , the RPCS alignment status signal “rpcs_algn_sts” is the AND operation alignment result for all four channels of the die (such as Figure 8 the output result of the AND gate). In the opposite direction, i.e., from the master die to the slave die, one bit indicates that alignment codewords exist in all channels of the link, and all channels need to set the FIFO read pointer to continue reading the storage location of the alignment codewords (RPCS alignment control signal “rpcs_algn_ctl”). All interface signals involved are as follows:

[0084] -rpcs_algn_sts_o[1:0] (output from “slave”, one combination for the eight-channel mode and another combination for the sixteen-channel mode)

[0085] -rpcs_algn_sts_i[2:0] (input to “master”)

[0086] -rpcs_algn_ctl_o (output from “master”, distributed to 3 slave dies)

[0087] -rpcs_algn_ctl_i (input to “slave”)

[0088] The multi-die de-skew operation is similar to the above single-die mode. The corresponding methods include: detecting alignment codewords in each data channel of each die; and storing the write pointer position as sideband information. Figure 13 Shown is an apparatus 1300 for inter-channel de-skew within a chip package structure including multiple circuit dies (i.e., dies). As shown, the apparatus 1300 includes channel alignment logic 800, which may include, for example, codeword detection logic 805. The codeword detection logic 805 is used to detect alignment codewords in the FIFOs of multiple data channels of multiple dies. The multiple dies include a master die 1302 and one or more slave dies 1304. Figure 13 An embodiment with three slave dies 1304 is shown, but such an embodiment should not be considered limiting.

[0089] Figure 13It further includes a die write clock generator 1306 disposed in the main die, which is used to generate a die write clock wr_tile_clk according to the local system clock tx_clk[0]. The period of the die write clock is equal to the period of the common reference clock refclk, and the position of the pulse corresponding to the die write clock within the common reference clock period is determined by the effective period of the counter. In some embodiments, the position of the pulse of the die write clock is related to the inter-die propagation time. In some embodiments, the position of the pulse of the die write clock can be programmed by adjusting the effective period of the counter.

[0090] Figure 13 It further includes a multi-channel controller 1308 disposed in the main die. The controller is used to: determine that alignment characters have been detected in the FIFOs of all channels of all dies; generate an alignment character found signal; and send the alignment character found signal to the synchronization logic of each die in the plurality of dies according to the die write clock. In some embodiments, the multi-channel controller includes a logical AND gate 1310, which is used to determine that alignment characters have been detected in the FIFOs of all channels of all dies by performing a logical "AND" operation on the die-specific alignment character found signals rpcs_algn_sts generated by each die. In some embodiments, the die-specific alignment character found signal is generated by a die-specific logical AND gate in each die. Each die-specific logical AND gate is used to generate the die-specific alignment character found signal by performing a logical "AND" operation on the channel-specific alignment character found signals associated with the respective data channels of a given die. Such die-specific AND gates 810 are shown in Figure 8 the channel alignment logic 800 of. As Figure 8 shown, the alignment character detection logic 805 is used to generate each channel-specific alignment character found signal as a pulse as an alignment character is detected in the data channel. The channel alignment logic 800 further includes a pulse broadening logic 815, which is used to broaden the pulse for a predetermined number of locally generated receive clock cycles.

[0091] As Figure 13 shown, each die further includes synchronization logic 1315. The synchronization logic 1315 of each die is used to sample the alignment character found signal according to the common reference clock and synchronize the alignment character found signal with the locally generated system clock tx_clk[n]. In each die, an alignment control state machine (such as Figure 8 825 in) is used to set the read pointer of each FIFO 820 in the die to the position containing the corresponding alignment character, and the plurality of FIFOs 820 are used to output data according to the locally generated system clock.

[0092] In some embodiments, channel alignment logic is used to store the position of each alignment symbol with each alignment symbol detected by the FIFOs of the plurality of data channels. Figure 8 stores the write pointer address "store_wr_ptr" to represent the address containing the corresponding alignment symbol.

[0093] In some embodiments, the maximum time skew between the data output by each FIFO according to the locally generated system clock is at most one period of the locally generated system clock.

[0094] In some embodiments, each die further includes a ring counter 1605 whose count value is synchronized by the alignment symbol found signal tx_algn_found. In some embodiments, each die may further include the rate adaptation logic 1200 described above in connection with Figure 12 The rate adaptation logic 1200 can be used to monitor the FIFO fill level "fill_level" of each FIFO of the plurality of data channels through the rate adaptation FSM 1210, and generate a FIFO fill level status signal "rpcs_fifo_sts" as the FIFO fill level of one of the FIFOs exceeds a threshold. The skip symbol detection logic 1205 is used to detect the skip ordered set of the FIFO of each data channel, and the rate adaptation FSM 1210 pads or truncates the skip ordered set of each FIFO according to the FIFO fill level status signal and a pre-designed value of the ring counter of each die. In some embodiments, padding and truncation are performed by the ways of not incrementing and doubling incrementing the read pointer, respectively.

[0095] Referring again to Figure 13 , the die-specific alignment symbol found signal rx_algn_str is a "logical AND" combination of all widened channel-specific alignment indication signals of the data channels within the same link. The combined die-specific alignment symbol found signal of each slave die is provided to the master die separately. It should be noted that since all rpcs_algn_sts signals are output by flip-flops and the widening degree is large enough, there is no risk of generating glitches during the "logical AND" combination process and the synchronization process.

[0096] In the master die, a signal "rx_algn_comb" is generated by performing a "logical AND" combination operation on the common alignment indication signals of all dies (including the master die itself). If the output is a valid high level, alignment symbols exist in all channels of all dies, thereby starting the de-skewing process. Since the "logical AND" combination signal is asynchronous, it needs to be synchronized by a two-flip-flop synchronization logic first.

[0097] The synchronized common alignment signal indicates that all data channels have alignment codewords and sets the FIFO read pointer of each channel to the alignment codeword storage location. Subsequently, all channels read data from the FIFO simultaneously. In the fifth-generation mode (32 GTps), the FIFO read pointer is synchronously updated at a clock frequency of 1 GHz, and the transmit timing budget allows for an uncertainty of one clock cycle. After the initial alignment, there is no further communication between the dies regarding misalignment. Channel alignment occurs at the start of the initialization, and a hysteresis is preferably present. Additionally, since there is no alignment indication signal after training, the dies do not send alignment loss indication information to the master die.

[0098] Multi-die clock timing

[0099] One difficulty in multi-die channel de-timing is that a 1 GHz signal that needs to be processed at a switching frequency of 500 MHz (rise / fall time < 1 ns) needs to be sent through I / O pads that can only handle switching frequencies within 200 MHz (corresponding to a rise / fall time of approximately 2.5 ns). To bypass this limitation, the die clock timing concept is adopted below. Generally speaking, this concept means that the master die distributes a balanced 100 MHz synchronous reference clock to all dies (master die and slave dies), which can achieve synchronization among all dies. Both the master die and the slave dies set the read pointer according to the position determined by the corresponding write pointer stored, and generate a 1 GHz local clock tx_clk[n] based on this 100 MHz common reference clock.

[0100] This clock timing mechanism is shown at the Figure 13 bottom to demonstrate the corresponding cross-clock domain scheme. The master die has a locally generated die write clock "wr_tile_clk" based on 1 GHz, which has an active cycle every 10 ns to match the period of the above 100 MHz reference clock. The output alignment control signal (such as algn_found) is timed by the die write clock and then sampled by the synchronized 100 MHz common reference clock in each of the master die and the slave dies. In this way, by properly timing the die write clock, the setup time from the die write clock to the reference clock ( Figure 14 twr in the timing diagram) can be made greater than 4 ns, which is sufficient to span the path routed from one die to another through the I / O pads and the substrate. After passing through the flip-flop timed by the refclk clock, the alignment codeword found signal is synchronized with the tx_clk of the locally generated 1 GHz channel by the synchronization processing stage.

[0101] The die write clock is generated by the die clock generator. Figure 14The logic and timing diagram of such a die clock generator are shown. Among them, in order to synchronize the die write clock with the reference clock, the reference clock is further synchronized with the 1GHz operating clock. The counter is triggered by the rising edge and counts from 0 to 9. The programmable decoder enables a sequence to be valid within any selected one cycle (i.e., the counter value). During this valid cycle, a gated clock is created according to the 1GHz operating clock, so as to obtain a die clock that is valid in one cycle out of every ten cycles.

[0102] Figure 14 The upper part of the timing diagram shows the synchronization of the reference clock and the generation of the enable pulse of the gated clock (i.e., the die write clock). twr represents the time between the generated clock in the master die and the 100MHz sampling clock refclk in the slave die. The twr time is programmable, and according to the timing requirements, it should be greater than 4ns. In some embodiments, the twr is programmed according to the setup and / or hold requirements, and can be adjusted by selecting one of the counter values that makes wr_tile_clk valid. For example, if the counter value is selected as "0", the setup time can be extended (but the hold time before the next cycle will be shortened), and if the counter value is selected as "4", the setup time can be shortened (and thus the hold time before the next cycle will be extended). The clock source of the die write clock is the 1GHz common clock in the master die, such as tx_clk[0] (PHY transmit clock channel 0).

[0103] Since all "algn_found" signals are synchronized with the 100MHz common reference clock and are respectively synchronized with the channel-based transmit clock tx_clk[n], there can be at most an uncertainty of one 1GHz clock cycle (1ns), thus very well falling within the 1.25ns timing tolerance required by the fifth-generation mode.

[0104] As Figure 14 shown in the schematic diagram, in addition to the above, another counter control logic is provided, which serves at least the following two purposes: (1) when the refclk synchronization processing stage is not in use, disable it to extend its service life; (2) disable the counter restart function. The counter control unit monitors the start_cnt pulse at all times to check whether it drifts.

[0105] Regarding the delay aspect of the above multi-die de-timing algorithm, the following considerations can be made:

[0106] - The slave die sends rx_algn_str to the master die;

[0107] - Synchronize the rx_algn_comb combined signal;

[0108] - Further synchronize the synchronized signal with the wr_tile_clk;

[0109] - The master die sends the obtained algn_found signal to the slave die;

[0110] - Sample the algn_found signal first with refclk and then with rd_tile_clk.

[0111] Generally speaking, compared with the delay of about 7 - 12 ns in the single-die alignment case, the total delay of the above case is about 20 - 25 ns. By adjusting the FIFO filling degree for rate adaptation, this delay can be shortened to 5 - 12 ns. Hereinafter, this will be described in further detail. When the target FIFO filling degree is set to the minimum value, the skip (SKP) ordered set will be removed from the data stream, thereby reducing the delay. Figure 15 It is a schematic diagram of the FIFO filling degree directly related to the delay. To minimize the area, the FIFO depth needs to be properly set, that is, the minimum depth is at least 32 codewords, and a FIFO based on dual-port SRAM needs to be considered. For example, for a FIFO based on flip-flops, the number of flip-flops required per die is 32 (bits) × 32 (depth) × 4 (channels) = 4096.

[0112] Figure 18 It is a flowchart of method 1800 in some embodiments. As shown, method 1800 includes: detecting 1805 alignment codewords in the FIFOs of multiple data channels of multiple dies, where the multiple dies include a master die and one or more slave dies. The method further includes: determining 1810 that alignment codewords have been detected in the FIFOs of all channels of all dies, and then generating an alignment codeword found signal. The method further includes: generating 1805 a die write clock according to the local system clock, where the period of the die write clock is equal to the period of the common reference clock, and the position of the pulse corresponding to the die write clock within the common reference clock period is determined by the valid period of the counter. The method further includes: sending 1820 the alignment codeword found signal to the synchronization logic in each slave die according to the die write clock. The method further includes: sampling 1825 the alignment codeword found signal in each slave die and the master die according to the common reference clock by the synchronization logic, so that the alignment codeword found signal is synchronized with the locally generated system clock of each die in the multiple dies, and then setting the read pointer of the FIFO to the position containing the alignment codeword. The method further includes: outputting 1830 data from each FIFO according to the locally generated system clock.

[0113] Multi-die rate adaptation

[0114] Figure 16 It is a block diagram of information exchange for multi-die rate adaptation in some embodiments. The information to be exchanged in each direction includes two bits: two of them represent the FIFO filling degree (status); the other two bits represent the obtained FIFO adjustment operation (control). Each die monitors the FIFO filling degree of all channels in the link. If any FIFO gives an indication of "FIFO full", the corresponding die reports "FIFO full" to the master die. Similarly, if any FIFO in the link gives an indication of "FIFO empty", the corresponding die reports "FIFO empty" to the master die. If one FIFO gives an indication of "full" while the other gives an indication of "empty", it indicates an error, and this error is reported to the master die.

[0115] As Figure 16 shown, a device 1600 for multi-die rate adaptation includes a plurality of ring counters 1605, each ring counter is respectively contained in the corresponding die of the multi-die package structure, and the plurality of ring counters are used to increment and output synchronization pulses. As Figure 16 shown, the multi-die package structure includes a master die 1610 and three slave dies 1615, but this does not imply any limitation. The master die 1610 in the multi-die package structure is used to synchronize the synchronization pulse with the count values of the plurality of ring counters 1605 according to the alignment symbol find signal "tx_algn_found". As described above, after generating the alignment symbol find signal according to the die write clock, it is synchronized to each die according to the common reference clock, so that the time offset between dies is not greater than one clock pulse of the locally generated system clock.

[0116] The FIFO filling degree detection logic in the master die is used to monitor that the FIFO filling degree of the FIFO in a certain die of the multi-die package structure exceeds the threshold after the first synchronization pulse, and output the rate adaptation control signal "rpcs_fifo_ctl" to each die of the multi-die package structure. As Figure 16 shown, the FIFO filling degree detection logic in the master die includes two OR gates. One OR gate 1620 is used to detect that a channel's FIFO is "full", and the other OR gate 1625 is used to detect that a channel's FIFO is "empty". Each die includes a rate adaptation FSM 1210, which is used to adjust the read pointer of each FIFO according to the rate adaptation control signal after subsequent synchronization pulses, so as to pad or truncate the stored skip symbols according to the rate adaptation control signal.

[0117] The information exchange operation is as described below, and the corresponding timing diagram is shown in Figure 17。The synchronous ring counter of each die continues to increment after being started by the alignment control signal. Among them, the counters of all channels operate in the same way, with a maximum difference of one clock cycle between dies. All rate adaptation operations are synchronized with this counter, which will be described in further detail below. The multi-die rate adaptation operation makes use of the above multi-die channel de-time skew operation by synchronizing the counter according to the alignment symbol generated during the channel de-time skew process to find the signal tx_algn_found.

[0118] The ring counter counts from 0 to N - 1, where N is programmable. Assuming N = 16, the counter repeats every 16 clock cycles. While the ring counter effectively counts the clock cycles, the corresponding logic is programmed to perform various operations at specific count values of the ring counter. As described in the context of multi-die de-time skew processing above, the ring counter of each channel of each die is started by the alignment pulse. In Figure 17 the timing diagram, the alignment pulse is the signal ra_sync. Since the counters of each die have been synchronized according to the tx_algn_found signal, the ring counters of each die generate synchronization pulses within the time range of a single 1 GHz system clock cycle. When the FIFO becomes "full" or "empty", this condition can be reported to the master die that is only synchronized with the ring counter. The FIFO fill level status signal fifo_stat is sent to the master die, for example, within a time of N / 2 = 8 cycles, which is sufficient for the signal to be transmitted from the slave die to the master die.

[0119] Within the master die, the FIFO status signal is synchronized by the tech_sync2 unit and is monitored after several cycles. Within the master die, after M clock cycles, the FIFO fill level is evaluated according to the synchronization pulse. M is programmable. In Figure 17 the timing diagram, M = 4. When programming M, the die-to-die transmission delay and synchronization delay need to be considered. Subsequently, the corresponding control signal fifo_ctrl is generated, namely "padding" (inserting skip symbols) or "truncating" (removing skip symbols). To achieve synchronization within each slave die, the control signal is widened by a programmable number of clock cycles. The widened control signal is fed back to all slave dies. As Figure 16 shown, the multi-channel controller can receive the ctrl_act signal from the ring counter of the master die, and this signal can correspond to the count value used to initiate Figure 17 the next operation in the series of operations shown.

[0120] In all slave dies, the information is synchronized by the tech_sync2 unit and is evaluated according to the synchronization pulse after K clock cycles. Similarly, K is programmable and is selected to compensate for the die-to-die transmission delay and synchronization delay. In Figure 17In the timing diagram, the obtained control signal is fifo_ra_plan (where "plan" represents the rate adaptation to be implemented). This signal stabilizes before the next synchronization pulse (ra_sync) from the ring counter arrives. The control logic then waits for the next synchronization pulse of the ring counter and, for example, in the case of N-1 or 0, activates the padding or truncation logic. In Figure 17 In the timing diagram, the corresponding signal is fifo_ra_action

[0121] (rate adaptation valid signal). Since the ring counters in all die run synchronously with the data stream in response to the tx_algn_found signal from the deserialization operation, any clock skew can be automatically compensated. With the arrival of the next skip ordered set ( Figure 17 SKP-OS in the timing diagram), the rate adaptation operation starts. The interval time (i.e., the number of clock cycles counted) between the skip ordered set and the synchronization pulse is the same in all die, so all die remove or insert skip symbols simultaneously. If the SKP-OS needs to be truncated, the control logic can double-increment the read pointer in one clock cycle, effectively skipping the position of the SKP-OS. If padding of the SKP-OS is required, the control logic can not increment the read pointer in one clock cycle, effectively causing the SKP-OS to be read twice.

[0122] In the above text, the end value (N) of the ring counter is a precisely known value. N can be programmed to a value sufficient to achieve a full round-trip transmission, thus resolving all inter-die transmission delays and synchronization uncertainties. Repeated FIFO "full" or FIFO "empty" indications do not cause any problems. Two possible solutions are given below.

[0123] When fifo_ra_action is originally valid, keep it valid. Additionally, when an indication requesting the opposite operation appears (e.g., initially a FIFO "full" indication and then a FIFO "empty" indication), fifo_ra_action can become invalid again. In one of the cases, fifo_ra_action can remain valid until the skip ordered set used to initiate the rate adaptation appears. When fifo_ra_action is originally valid, the control logic requests a rate adaptation operation. If the FIFO fill level continues to change (e.g., due to loss of skip ordered sets during long-distance packet transmission), a second or third rate adaptation request can be issued continuously. After being processed in the same way as above, such requests can be further saved by the padding and truncation control logic. After a period of time, when one or more skip ordered sets arrive, multiple rate adaptation steps can be executed sequentially without further interaction.

[0124] In the multi-die rate adaptation, the following signals are used:

[0125] - rpcs_fifo_sts_o[1:0][1:0] (output from the "slave", one combination for the eight-channel mode and another for the sixteen-channel mode)

[0126] - rpcs_fifo_sts_i[2:0][1:0] (input to the "master")

[0127] - rpcs_fifo_ct_i[1:0] (input to the "slave")

[0128] The status information uses the following encoding scheme:

[0129] - 2'b 00: All FIFOs of the slave die are within the upper and lower limits (no action required)

[0130] - 2'b 01: Any one of the FIFOs of the slave die is "full" (requiring SKP removal)

[0131] - 2'b 10: Any one of the FIFOs of the slave die is "empty" (requiring SKP insertion)

[0132] - 2'b 11: An error has occurred, the FIFO operation is abnormal, notify the master die

[0133] The control information uses the following encoding scheme:

[0134] - 2'b 00: No action required, keep the FIFO unchanged

[0135] - 2'b 01: Remove one SKP ordered set from the data stream

[0136] - 2'b 10: Insert one SKP ordered set into the data stream

[0137] - 2'b 11: An error has occurred, insert an ERROR ordered set into the data stream

[0138] Since the multi-channel rate adaptation and alignment information exchange adopt the same synchronization concept, that is, both use the reference clock and the die write clock, there is no need to use Gray coding. One of the problems may be the long turnaround time. When the FIFO is "full" or "empty", a request to insert or remove the SKP ordered set can be issued quickly. However, it takes a long time for the next SKP ordered set to arrive. First, when processing (inserting or removing) the SKP ordered set, the FIFO fill level is updated accordingly. However, at the same time, the multi-channel controller module may have issued the next FIFO control request before, resulting in the problem of performing one more SKP insertion or removal. For this problem, a solution that can be adopted is to immediately change the FIFO fill level indication information once there is a FIFO fill level change control request, and update the FIFO fill level again at the first time after the change request is executed. When the FIFO is "full" or "empty", this information is forwarded to the main die via the "rpcs_fifo_sts" line. The main die then issues an "insert_skp" request or a "remove_skp" request. At the same time, any FIFO "full" or "empty" indication information from the slave die is blocked inside the main die for N clock cycles, where N is programmable. In this way, subsequent non-desired FIFO change requests can be blocked before the actual request is processed. After the FIFO change request is synchronized, it is forwarded to all slave dies via the "rpcs_fifo_ctl" line. After the FIFO controller (for each channel) is addressed, it stores the request (separately and individually), changes the FIFO fill level indication information to "normal", and maintains this state until the request can be finally processed. Once the SKP ordered set is detected, the FIFO update request can be immediately executed, and the corresponding SKP insertion or SKP removal can be performed. After the request is processed, the FIFO fill level is updated. If the FIFO fill level status is still not "normal", the FIFO fill status is sent to the main die again via the "rpcs_fifo_sts" line.

[0139] In some embodiments, a method includes: detecting skip ordered sets in a plurality of data channels, and storing skip pulses in corresponding FIFO locations of each data channel as each skip ordered set is detected. The method further includes: finding a signal according to an alignment symbol, synchronously starting a ring counter in each die of a multi-die package structure, each ring counter synchronously maintaining a count value, and periodically outputting a synchronization pulse. As described above, after generating the alignment symbol find signal according to the die write clock, it is synchronized to each die according to a common reference clock, so that the time offset between the dies is no greater than one clock pulse of the locally generated system clock.

[0140] In response to the first synchronization pulse, by monitoring the status signal with the logic in the master die, the filling degree of each FIFO in the multi-die package structure is monitored according to the preset value of the ring counter. Subsequently, when it is determined that the filling degree of the FIFO in a certain die of the multi-die package structure exceeds the threshold, a rate adaptive control signal is output. After each ring counter reaches the second preset value, the rate adaptive control signal is respectively evaluated through the corresponding logic in each die of the multi-die package structure. With the arrival of the second synchronization pulse, the rate adaptive logic is started so that it operates on the SKP ordered set in the FIFO of each die according to the rate adaptive control signal.

[0141] Time skew budget

[0142] For retimers, the PCIe basic specification has different requirements for the inter-channel input time skew and the inter-channel output time skew: the inter-channel input time skew must be compensated, while the inter-channel output time skew is acceptable. The input and output time skews vary with different data rates.

[0143] Input (RX) time skew

[0144] The input time skew requirements are shown in Table I below. When converting the time / UI requirements into the corresponding clock cycle requirements, the de-skew requirements can be extracted first. The de-skew logic obtains the data sufficient to read all channels in a de-skewed manner by querying the memory, or it stores these data itself. In this way, the fastest channel will slow down its speed relative to the slowest channel, resulting in an increase in delay. The number of clock cycles required for this process is listed in the column of "de-skew requirements". It can be seen that synchronizing the de-skew information and the rate adaptive information of all (asynchronous) channels requires an additional three clock cycles (the column of "CDC overhead"), and the corresponding de-skew budget is listed in the right column.

[0145] Input time skew

[0146]

[0147] Table I: Input (Rx) time skew requirements

[0148] Output (TX) time skew

[0149] The output skew requirements are shown in Table II below. In the PCIe basic specification, the skew unit is ns, which is converted to the corresponding unit intervals and number of clock cycles (right column) here. Since the output skew is greater than one clock cycle (16 / 32 GTps), the clock synchronization requirements are easier to meet because the uncertainty of one clock cycle is acceptable between channels of multiple dies. In the low data rate mode, the output skew requirements are more difficult to meet, so appropriate synchronization operations are required. However, since the clock frequency is below 500 MHz, achieving synchronization with a 1 GHz clock (i.e., + / - 1 GHz clock cycles) is sufficient to meet the PCIe output skew requirements.

[0150] Output skew

[0151]

[0152] Table II: Output (Tx) skew requirements

Claims

1. A method, characterized in that, Comprising: Detecting alignment characters in a first-in, first-out (FIFO) buffer of multiple data channels of multiple dielets, wherein the multiple dielets include a master dielet and one or more slave dielets; Determining that an alignment character has been detected in the FIFO buffer of each channel of each dielet, and generating an alignment character found signal in response; Generating a dielet write clock based on a local system clock, wherein a period of the dielet write clock is equal to a period of a common reference clock, and wherein the dielet write clock corresponds to a pulse, and a position of the pulse within the period of the common reference clock is determined by an active period of a counter; In response to the dielet write clock, sending the alignment character found signal to synchronization logic of each of the slave dielets; Sampling, by the synchronization logic in each slave dielet and each master dielet, the alignment character found signal based on the common reference clock; For each dielet of the multiple dielets, synchronizing the alignment character found signal with a locally generated system clock, and setting a read pointer of the FIFO buffer to a position including the alignment character in a responsive manner; and Outputting data from each FIFO buffer based on the locally generated system clock.

2. The method according to claim 1, wherein Further comprising: Storing a position including the alignment character in response to detecting each alignment character in the FIFO buffer of the multiple data channels.

3. The method according to claim 1, wherein The position of the pulse of the dielet write clock is related to an inter-dielet propagation time.

4. The method according to claim 1, wherein The position of the pulse of the dielet write clock can be programmed by adjusting the active period of the counter.

5. The method according to claim 1, wherein A maximum time skew between data output from each FIFO buffer based on the locally generated system clock is at most one period of the locally generated system clock.

6. The method according to claim 1, characterized in that, Determining that an alignment character has been detected in the FIFO buffer of each channel of each dielet includes performing a logical AND operation on dielet-specific alignment character found signals generated by each dielet.

7. The method according to claim 6, wherein Further comprising: Generating the dielet-specific alignment character found signal by performing a logical AND operation on channel-specific alignment character found signals associated with each data channel of a given dielet.

8. The method according to claim 7, wherein Each channel-specific alignment character found signal corresponds to a pulse, wherein the pulse is generated in response to detecting the alignment character in the data channel, and wherein the pulse is widened by a preset number of locally generated receive clock periods.

9. The method according to claim 1, wherein Further comprising: Synchronizing count values of multiple ring counters, wherein each dielet has one of the multiple ring counters.

10. The method according to claim 9, characterized in that, Further comprising: Monitoring a FIFO buffer fill level of each FIFO buffer of the multiple data channels; Generating a FIFO buffer fill level status signal in response to the FIFO buffer fill level of one of the FIFO buffers exceeding a threshold; Detecting a skipped ordered set in the FIFO buffer of each data channel; And In response to the FIFO buffer fill level status signal, pad or truncate the skip ordered sets within each FIFO buffer, wherein the padding or truncation is performed according to a preset value of the ring counter of each die.

11. A device, characterized in that, Comprising: Alignment symbol detection logic, wherein the alignment symbol detection logic is used to detect alignment symbols in the FIFO buffers of multiple data channels of multiple dies, wherein the multiple dies include a master die and one or more slave dies; A die write clock generator located within the master die, wherein the die write clock generator is used to generate a die write clock according to a local system clock, wherein the period of the die write clock is equal to the period of a common reference clock, wherein the die write clock corresponds to a pulse, and the position of the pulse within the period of the common reference clock is determined by the effective period of a counter; A multi-channel controller located within the master die, wherein the multi-channel controller is used to: determine that alignment symbols have been detected in the FIFO buffer of each channel of each die; generate an alignment symbol found signal; and in response to the die write clock, send the alignment symbol found signal to the synchronization logic of each of the multiple dies; The synchronization logic located within each slave die, wherein the synchronization logic is used to: sample the alignment symbol found signal according to the common reference clock; and synchronize the alignment symbol found signal with the locally generated system clock for each die of the multiple dies; An alignment control state machine located within each die, wherein the alignment control state machine is used to set the read pointer of each FIFO buffer of the die to the position including the alignment symbol; and Multiple of the FIFO buffers, for outputting data according to the locally generated system clock.

12. The device according to claim 11, characterized in that, In response to detecting each alignment symbol in the FIFO buffers of the multiple data channels, store the position including the alignment symbol.

13. The device according to claim 11, characterized in that, The position of the pulse of the die write clock is related to the propagation time between dies.

14. The device according to claim 11, characterized in that, The position of the pulse of the die write clock can be programmed by adjusting the effective period of the counter.

15. The device according to claim 11, wherein The maximum time skew between the data output from each FIFO buffer according to the locally generated system clock is at most one period of the locally generated system clock.

16. The device according to claim 11, characterized in that, The multi-channel controller located within the master die includes a logic AND gate, wherein the logic AND gate is used to determine that alignment symbols have been detected in the FIFO buffer of each channel of each die by performing a logical "AND" operation on the die-specific alignment symbol found signals generated by each die.

17. The device according to claim 16, wherein Further comprising a die-specific logic AND gate located within each die, wherein the die-specific logic AND gate is used to generate the die-specific alignment symbol found signal by performing a logical "AND" operation on the channel-specific alignment symbol found signals associated with each data channel of a given die.

18. The device according to claim 17, characterized in that, The alignment symbol detection logic is used to generate each channel-specific alignment symbol found signal as a pulse, where the pulse is generated in response to detecting the alignment symbol within the data channel, and each data channel includes pulse stretching logic for stretching the pulse by a preset number of locally generated receive clock cycles.

19. The device according to claim 11, characterized in that, It further includes a ring counter located within each die, where the ring counter has a count value synchronized by the alignment symbol found signal.

20. The device according to claim 19, characterized in that, It further includes: Skip symbol detection logic for detecting a skip ordered set within the FIFO buffer of each data channel; A rate adaptation finite state machine, where the rate adaptation finite state machine is used to: monitor the FIFO buffer fill level of the FIFO buffer of each of the plurality of data channels; generate a FIFO buffer fill level status signal in response to the FIFO buffer fill level of one of the FIFO buffers exceeding a threshold; and in response to the FIFO buffer fill level status signal, pad or truncate the skip ordered set within each FIFO buffer, where the padding or truncation is performed according to a preset count value of the ring counter of each die.

Citation Information

Patent Citations

  • Method for Code Evaluation Using ISI Ratio

    US20150222458A1

Cited By

  • Data transmission method based on UCIe, storage medium and artificial intelligence chip

    CN120934696A

  • Uci e-based data transmission method, storage medium and artificial intelligence chip

    CN120934696B