Lane Repair and Lane Reversal Implementation for Die-to-Die (D2D) Interconnects

By introducing redundant diamond drills and circuits into chip assembly technology and adopting a multi-protocol-capable packaged interconnect protocol, the waste and cost problems caused by failure of a single diamond drill connection are solved, achieving higher reliability and flexibility.

JP2024530062A5Active Publication Date: 2025-05-09INTEL CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023567003
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-06-20
Filing Date
2022-11-22
Publication Date
2025-05-09
Estimated Expiration
2042-11-22

AI Technical Summary

Technical Problem

In existing chip assembly technology, failure to connect a single diamond will cause the entire package to be discarded, causing waste and cost issues.

Method used

Adopt advanced interconnection technologies, including the design of redundant diamonds and other redundant circuits and recovery mechanisms in the physical layer, and are managed through multi-protocol-capable packaged interconnect protocols such as UCIe for failure recovery and data channel remapping.

Benefits of technology

Improves the inefficiency of chip assembly production and design costs, while reducing waste and costs caused by diamond connection failures, achieving higher reliability and flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

In one embodiment, an apparatus includes a first die having a die-to-die adapter for communicating with a protocol layer circuit and a physical layer circuit, where the die-to-die adapter receives first information of a first interconnect protocol; and a physical layer circuit coupled to the die-to-die adapter. The physical layer circuit is configured to receive and output the first information to the second die via the interconnect, and has a first plurality of transmitters for transmitting data over a first plurality of data lanes; and at least one redundant transmitter. The physical layer circuit may be configured to remap a first data lane of the first plurality of data lanes to the at least one redundant transmitter. Other embodiments are described and claimed.
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] As chiplets are integrated using advanced packaging techniques, yields are high and design costs remain low, but a single bump connection failure on a manufactured die can result in the entire packaged part being scrapped, leading to waste and cost issues. [Brief description of the drawings]

[0002] [Figure 1] FIG. 2 is a block diagram of a package according to one embodiment.

[0003] [Figure 2A] 1A-1C are cross-sectional views of different packaging options incorporating embodiments. [Figure 2B] 1A-1C are cross-sectional views of different packaging options incorporating embodiments. [Figure 2C] 1A-1C are cross-sectional views of different packaging options incorporating embodiments. [Figure 2D] 1A-1C are cross-sectional views of different packaging options incorporating embodiments.

[0004] [Figure 3A] FIG. 1 is a block diagram of a layered protocol in which one or more embodiments may be implemented. [Figure 3B] FIG. 1 is a block diagram of a layered protocol in which one or more embodiments may be implemented.

[0005] [Figure 4A] FIG. 1 is a block diagram of a multi-die package in accordance with various embodiments. [Figure 4B] FIG. 1 is a block diagram of a multi-die package in accordance with various embodiments.

[0006] [Diagram 5] FIG. 2 is a schematic diagram illustrating a die-to-die connection according to one embodiment.

[0007] [Figure 6A] FIG. 1 is a timing diagram illustrating sideband signaling according to one embodiment. [Figure 6B] FIG. 1 is a timing diagram illustrating sideband signaling according to one embodiment.

[0008] [Figure 7] FIG. 2 is a flow diagram illustrating a bring-up flow for an on-package multi-protocol interconnect according to one embodiment.

[0009] [Figure 8] FIG. 13 is a flow diagram of a link training state machine according to one embodiment.

[0010] [Figure 9] FIG. 11 is a flow diagram of further details of main band initialization according to one embodiment.

[0011] [Figure 10] FIG. 13 is a flow diagram of main band training according to one embodiment.

[0012] [Figure 11A] 1 is an illustration of data lane remapping possibilities according to one embodiment.

[0013] [Figure 11B] 1 is an exemplary remapping configuration for a single lane failure within a module according to one embodiment.

[0014] [Figure 11C] 1 is an exemplary remapping configuration for a two-lane failure in a module according to one embodiment.

[0015] [Figure 12A] FIG. 2 is a schematic diagram of a portion of a die-to-die connection according to one embodiment.

[0016] [Figure 12B] FIG. 1 is a schematic diagram of a portion of a die-to-die connection according to one embodiment illustrating a single lane remapping operation.

[0017] [Figure 12C] FIG. 2 is a schematic diagram of a portion of a die-to-die connection illustrating a two lane remapping operation according to one embodiment.

[0018] [Figure 13] FIG. 2 is a flow diagram of a method according to one embodiment.

[0019] [Figure 14] FIG. 2 is a block diagram of another exemplary system according to one embodiment.

[0020] [Figure 15] FIG. 1 is a block diagram of a system according to another embodiment, such as an edge platform.

[0021] [Figure 16] FIG. 2 is a block diagram of a system according to another embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0022] In various embodiments, in a semiconductor package Implemented in One or more die implementations provide a recovery scheme do , Using advanced interconnect technology The yield of the packaged parts Revised In implementation, the distribution of redundant bumps and other redundant circuits and corresponding recovery mechanisms within the physical layer may be carefully designed to achieve maximum coverage with minimum overhead. Also, embodiments provide the ability to perform lane remapping and / or lane reversal between transmit and receive lanes when these lanes are not connected in a bit-lane-matching manner due to, for example, floorplan constraints.

[0023] The embodiments provide a very general means to solve all types of lane remapping issues, such as repair, remapping, etc., with minimal overhead to bump area and performance. To this end, redundant lanes can be provided at the beginning and end of each module, with enough flexibility to shift per-lane data forward and backward independently at the transmit and receive sides. Full utilization of redundant lanes can be used to repair and recover from any type of bump connection issue.

[0024] In various embodiments, a multi-protocol on-package interconnect may be used to communicate between separated dies of a package. Initializing and training this interconnect with an ordered bring-up flow may enable independent reset of different dies, detection of "end of reset" of the partner die, and ordered initialization and training (in that order) of the sideband and mainband interfaces of the interconnect. More specifically, sideband initialization may be performed to detect that the link partner die has exited reset and initialize and train the sideband. The mainband may then be initialized and trained, which may include any lane reversal and / or repair operations as further described herein. Such mainband operations may utilize the already brought-up sideband to communicate synchronization and status information.

[0025] In embodiments that perform lane reversal and / or repair, yield loss due to lane connection issues in advanced packaging multi-chip packages (MCPs) can be recovered. Furthermore, lane repair techniques according to an embodiment can provide both left-shift and right-shift techniques to cover the entire bump map for efficient lane repair. Still further, lane reversal detection can enable die rotation and die mirroring to enable multiple on-package instantiations on the same die. In this manner, lane reversal can eliminate multiple tape-ins of the same die.

[0026] Embodiments may be implemented in the context of a multi-protocol on-package interconnect protocol that may be used to connect multiple chiplets or dies on a single package, where a powerful ecosystem of separated die architectures may be interconnected together. The on-package interconnect protocol may be referred to as the "Universal Chiplet Interconnect Express (UCIe) interconnect protocol" and may conform to the UCIe specification, which may be published by a Special Interest Group (SIG) or other promoter or entity. Although referred to herein as "UCIe," it should be understood that the multi-protocol on-package interconnect protocol may employ other terminology.

[0027] The UCIe interconnect protocol may support multiple underlying interconnect protocols, including a flit-based mode of a particular communication protocol. In one or more embodiments, the UCIe interconnect protocol may support a flit mode of the CXL protocol, for example according to a given version of the Compute Express Link (CXL) specification (such as the CXL Specification Version 2.0 (published in November 2020) or any future updates, versions or variations thereof); a PCIe flit mode, for example according to a given version of the Peripheral Component Interconnect Express (PCIe) specification (such as the PCIe Base Specification Version 6.0 (published in 2022) or any future updates, versions or variations thereof); and a raw (or streaming) mode used to map any protocol supported by the link partner. It should be noted that in one or more embodiments, the UCIe interconnect protocol may not be backward compatible and may instead accommodate current and future versions of the above-mentioned protocols or other protocols that support a flit mode of communication.

[0028] Embodiments may be used to provide compute, memory, storage and connectivity across the compute continuum, spanning cloud, edge, enterprise, 5G, automotive, high performance computing and handheld segments. Embodiments may be used to package or otherwise combine die from different sources, including different fabs, different designs and different packaging technologies.

[0029] Chiplet integration on package also allows customers to make different tradeoffs for different market segments by choosing different numbers and types of dies. For example, depending on the segment, one may choose different numbers of compute, memory and I / O dies. Thus, lower product stock keeping unit (SKU) costs are achieved since different die designs for different segments are not required.

[0030] Referring now to FIG. 1, a block diagram of a package according to one embodiment is shown. As shown in FIG. 1, package 100 may be any type of integrated circuit package. In the particular diagram shown, package 100 includes a central processing unit (CPU) die 110. 0-n , accelerator die 120, input / output (I / O) tiles 130 and memory 140 1-4 At least certain of these dies may be coupled together via on-package interconnects according to one embodiment. As shown, interconnects 150 1-3 may be implemented as a UCIe interconnect. CPU 110 may be coupled via another on-package interconnect 155, which may potentially provide CPU-to-CPU connectivity on-package with the UCIe interconnect implementing a coherency protocol. As one such example, this coherency protocol may be Intel® Ultra Path Interconnect (UPI); of course, other examples are possible.

[0031] Protocols mapped to the UCIe protocol described herein include PCIe and CXL, although it should be understood that the embodiments are not limited in this respect. In an exemplary embodiment, the mapping of any underlying protocol may be done using flit formats, including raw modes. In one implementation, these protocol mappings may enable more on-package integration by replacing certain physical layer circuitry (e.g., PCIe SERDES PHY and PCIe / CXL LogPHY with link level retry) with a UCIe die-to-die adapter and PHY according to one embodiment, improving power and performance characteristics. In addition, raw modes may be protocol agnostic to allow other protocols to be mapped while enabling usage such as standalone SERDES / transceiver tile (e.g., Ethernet) on-package integration. As further shown in FIG. 1, the off-package interconnect may follow a variety of protocols, including CXL / PCIe protocols and double data rate (DDR) memory interconnect protocols, etc.

[0032] In an exemplary implementation, accelerator 120 and / or I / O tile 130 may be connected to CPU 110 using CXL transactions operating over UCIe interconnect 150, leveraging CXL I / O, coherency, and memory protocols. In the embodiment of FIG. 1, I / O tile 130 may provide an interface to CXL, PCIe, and DDR pins outside the package. Accelerator 120 may also be connected to CPU 110 using PCIe transactions operating over UCIe interconnect 150, either statically or dynamically.

[0033] A package according to an embodiment may be implemented in many different types of computing devices, from small portable devices such as smartphones to larger devices including client computing devices and servers or other data center computing devices. Thus, the UCIe interconnect may enable local and long distance connections at the rack / pod level. Although not shown in FIG. 1, it should be understood that at least one UCIe retimer may be used to extend the UCIe connection beyond the package with an off-package interconnect. Examples of off-package interconnect include electrical cables, optical cables, or any other technology for connecting packages at the rack / pod level.

[0034] Embodiments may further be used to support rack / pod level disaggregation using the CXL 2.0 (or later) protocol. In such a configuration, multiple compute nodes (e.g., a virtual hierarchy) from different compute chassis couple to a CXL switch that may couple to multiple CXL accelerators / Type-3 memory devices that may be located in one or more separate drawers. Each compute drawer may couple to this switch using an off-package interconnect that runs the CXL protocol through a UCIe retimer.

[0035] 2A-2D, cross-sectional views of different packaging options incorporating embodiments are shown. As shown in FIG. 2A, package 200 may be an advanced package offering advanced packaging techniques. In one or more embodiments, advanced package implementations may be used for performance optimization applications, including power-saving performance applications. In some such exemplary use cases, the channel reach may be short (e.g., less than 2 mm) and the interconnects may be optimized for high bandwidth and low latency with the best performance and power efficiency characteristics.

[0036] As shown in FIG. 2A, the package 200 includes multiple dies 210. 0-2 2A , three particular dies are shown, it should be understood that there may be more dies in other implementations. The die 210 is fitted onto a package substrate 220. In one or more embodiments, the die 210 may be fitted onto the substrate 220 via bumps. As shown, the package substrate 220 includes on-package interconnects 226 1-2 Multiple Silicon Bridges 225 1-2 In one embodiment, interconnect 226 may be implemented as a UCIe interconnect and silicon bridge 225 may be implemented as an Intel® EMIB bridge.

[0037] 2B, another embodiment of an advanced package is shown in which the package configuration is implemented as a chip-on-wafer-on-substrate (CoWoS). In this figure, package 201 includes a die 210 fitted onto an interposer 230, which includes corresponding on-package interconnects 236. As a result, interposer 230 fits onto package substrate 220 via bumps.

[0038] 2C, another embodiment of an advanced package is shown in which the package configuration is implemented using a fan-out organic interposer 230. In this figure, package 202 includes a die 210 fitted onto an interposer 230 that includes corresponding on-package interconnects 236. As a result, interposer 230 fits onto a package substrate 220 via bumps.

[0039] Referring now to FIG. 2D, another package diagram is shown. Package 203 may be a standard package offering standard packaging technology. In one or more embodiments, a standard package implementation may be used for low-cost and long-distance (e.g., 10 mm to 25 mm) interconnect using traces on an organic package / substrate while still offering significantly better BER performance compared to off-package SERDES. In this implementation, package 203 includes a die 210 fitted to a package substrate 220 with on-package interconnect 226 fitted directly into the package substrate 220 without including a silicon bridge or the like.

[0040] 3A / 3B, a block diagram of a layered protocol that may implement one or more embodiments is shown. As shown at a high level in FIG. 3A, multiple layers of the layered protocol implemented in the circuit 300 may implement an interconnect protocol. The protocol layer 310 may communicate information of one or more application-specific protocols. In one or more implementations, the protocol layer 310 may operate according to one or more of PCIe or CXL flit modes and / or streaming protocols to provide a generic mode for user-defined protocols to be transmitted. For each protocol, different organizations and associated flit transfers are available.

[0041] As a result, the protocol layer 310 couples to a die-to-die adapter (D2D) adapter 320 via an interface 315. In one embodiment, the interface 315 may be implemented as a flit-aware D2D interface (FDI). In one embodiment, the D2D adapter 320 is a multi-layered interface that includes the protocol layer 310 and the physical layer 330. Cooperation with to ensure successful data transfer across the UCIe link 340. The adapter 320 may be configured to provide a low-latency, optimized data path for the protocol flits by minimizing logic on the main data path as much as possible.

[0042] FIG. 3A illustrates various functions performed within the D2D adapter 320. The D2D adapter 320 may provide link state management and parameter negotiation for connected dies (also referred to as "chiplets"). Still further, the D2D adapter 320 may optionally ensure reliable delivery of data through cyclic redundancy checks (CRCs) and link level retry mechanisms, for example, when the raw BER is less than 1e-27. When multiple protocols are supported, the D2D adapter 320 may define an underlying arbitration mechanism. For example, when transporting CXL protocol communications, the adapter 320 may provide an arbitrator / multiplexer (ARB / MUX) function that supports communications of multiple simultaneous protocols. In one or more embodiments, when the D2D adapter 320 is responsible for reliable transfers, a flow control unit (flit) of a given size, for example 256 bytes, may define the underlying transfer mechanism.

[0043] If the operation is in flit mode, the die-to-die adapter 320 may insert and check CRC information. In contrast, if the operation is in raw mode, all information (e.g., bytes) of the flit is populated by the protocol layer 310. If applicable, the adapter 320 may also perform retries. The adapter 320 may further be configured to coordinate higher level link state machine management as well as bring up protocol option related parameter exchange with the remote link partner and power management coordination with the remote link partner, if supported. Depending on the usage model, different underlying protocols may be used. For example, in one embodiment, data transfer using direct memory access, software discovery and / or error handling, etc. may be handled using PCIe / CXL.io; memory use cases may be handled through CXL.Mem; and cache requirements for applications such as accelerators may be handled using CXL.cache.

[0044] As a result, the D2D adapter 320 couples to a physical layer 330 via an interface 325. In one embodiment, the interface 325 may be a raw D2D interface (RDI). As shown in FIG. 3B, the physical layer 330 includes circuitry for interfacing with a die-to-die interconnect 340 (which in one embodiment may be a UCIe interconnect or another multi-protocol capable on-package interconnect). In one or more embodiments, the physical layer 330 may be responsible for electrical signaling, clocking, link training, sideband, etc.

[0045] The interconnect 340 may include sideband and mainband links that may be in the form of so-called "lanes," which are physical circuits for carrying signaling. In one embodiment, a lane may constitute a circuit for carrying a pair of signals (one for transmit and one for receive) that are mapped to physical bumps or other conductive elements. In one embodiment, xN A UCIe link consists of N lanes.

[0046] 3B, the physical layer 330 includes three subcomponents: physical (PHY) logic 332, an electrical analog front-end (AFE) 334, and a sideband circuit 336. In one embodiment, the interconnect 340 includes a main band interface that provides a main data path over multiple physical bumps that may be organized into groups of lanes called modules or clusters.

[0047] The unit of construction of the interconnect 340 is referred to herein equally as a "cluster" or a "module." In one embodiment, a cluster may include N single-ended unidirectional full duplex data lanes, one single-ended lane for valid, one lane for trace, a differential forwarding clock per direction, and two lanes per direction for sideband (single-ended clock and data). Thus, the module (or cluster) forms the atomic grain of the structural design implementation of the AFE 334. There may be different numbers of lanes provided per module for standard and advanced packages. For example, for a standard package, 16 lanes make up a single module, while for an advanced package, 64 lanes make up a single module. Although embodiments are not limited in this respect, the interconnect 340 is a physical interconnect that may be implemented using one or more of conductive traces, conductive pads, bumps, and the like that provide interconnections between PHY circuits present on link partner dies.

[0048] A given instance of protocol layer 310 or D2D adapter 320 can transmit data through multiple modules where bandwidth scaling is implemented. The physical link of the interconnect 340 between the dies may include two separate connections: (1) a sideband connection; and (2) a mainband connection. In an embodiment, the sideband connection is used for parameter exchange, register access for debug / link training and compliance with and coordination with the remote partner for management.

[0049] In one or more embodiments, the sideband interface is formed with at least one data lane and at least one clock lane in each direction. Stated another way, the sideband interface is a two-signal interface for transmit and receive directions. In the use of advanced packages, additional data and clock pairs in each direction can be provided for redundancy, for repair or bandwidth increase. The sideband interface can include a forwarding clock pin and a data pin in each direction. In one or more embodiments, the sideband clock signal can be generated by an auxiliary clock source configured to operate at 800 MHz regardless of the main data path speed. The sideband circuitry 336 of the physical layer 330 can be provided with an auxiliary power supply and can be included in an always-on domain. In one embodiment, the sideband data can be communicated with a single data rate signal (SDR) of 800 megatransfers per second (MT / s). The sideband can be configured to operate on an always-on power supply and an auxiliary clock source. Each module has its own set of sideband pins.

[0050] The main band interface, which constitutes the main data path, may include a forwarding clock, a data valid pin, and N data lanes per module. In the advanced package option, N=64 (also referred to as ×64), with a total of four additional pins for lane repair provided in the bump map. In the standard package option, N=16 (also referred to as ×16), with no additional pins for repair provided. The physical layer 330 may be configured to coordinate the different functions and their relative sequencing for proper link bring-up and management (e.g., sideband forwarding, main band training and repair, etc.).

[0051] In one or more embodiments, an advanced package implementation may support redundant lanes (also referred to herein as "spare lanes") to handle defective lanes (including clock, enable, sideband, etc.). In one or more embodiments, a standard package implementation may support lane width reduction to handle failures. In some embodiments, multiple clusters may be aggregated to provide higher performance per link.

[0052] Referring now to FIG. 4A, a block diagram of a multi-die package according to one embodiment is shown. As shown in FIG. 4A. The package 400 includes at least a first die 410 and a second die 450. It should be understood that the dies 410 and 450 may be various types of dies including CPUs, accelerators, or I / O devices, etc. In the high-level diagram shown in FIG. 4A, an interconnect 440 coupling the dies together is shown as a dashed line. The interconnect 440 may be an instantiation of an on-package multi-protocol capable interconnect, such as a UCIe interconnect as described herein. Although not shown in detail in FIG. 4A, it should be understood that the interconnect 440 may be implemented using conductive bumps fitted to each die that may couple together to provide an interconnect between the dies. In addition, the interconnect 440 may further include in-package circuitry, such as conductive lines on or within one or more substrates. It should be understood that the term "lane" as used herein refers to any interconnect circuitry coupling one die to another die.

[0053] In a particular embodiment, the interconnect 440 may be a UCIe interconnect having one or more modules, each module including a sideband interface and a main band interface. In this high-level view, the main band interface couples to main band receiver and transmitter circuitry within each die. Specifically, the die 410 includes a main band receiver circuit 420 and a main band transmitter circuit 425, while the die 450, in turn, includes a main band receiver circuit 465 and a main band transmitter circuit 460.

[0054] FIG. 4A further illustrates the connections of the sideband interface. Generally, the sideband includes a data lane and a clock lane in each direction, and in the use of advanced packaging, additional data and clock pairs in each direction may be provided for redundancy. Thus, FIG. 4A illustrates a first possible connection implementation between the sideband circuits of these two dies. The die 410 includes a sideband circuit 430, which includes a first sideband circuit 432, which includes corresponding sideband clock and data receivers (R_C and R_D) and sideband clock and data transmitters (T_C and T_D) that respectively couple to corresponding sideband transmitter and receiver circuits of the sideband circuit 470 of the second die 450. The sideband circuit 430 includes redundant sideband clock and data transmitters and receivers (see above). To what is shown Transmitter and Receiver Abbreviation "R" is added to the end of ) also includes a second sideband circuit 434 having similar circuitry for the first sideband 432 and the second sideband 434 having similar circuitry for the second sideband 432.

[0055] 4A, a first sideband connection instantiation is shown, in which sideband circuits 432 and 472 function as functional sidebands and sideband circuits 434 and 474 function as redundant sidebands.

[0056] Depending on the sideband detection performed during sideband initialization, it may be determined that one or more of the sideband lanes and / or associated sideband circuits are defective and therefore at least a portion of the redundant sideband circuitry may be used as part of the functional sideband. More specifically, FIG. 4B shows a second possible connection implementation between the sideband circuits of the two dies. In this example, redundant sideband data transmitters and receivers are present in the sideband circuitry 472 and function as part of the functional sideband.

[0057] In different implementations, the initialization and bring-up flow may allow any connection as long as data-to-data and clock-to-clock connections are maintained. If redundancy is not required based on such initialization, both sideband circuit pairs may be used to extend the sideband bandwidth, allowing faster message exchange. Note that while Figures 4A and 4B are shown in the context of an advanced package configuration, similar sideband circuitry may be present in a die used in a standard package. However, in certain implementations, redundant sideband circuitry and redundant sideband lanes may not be present in a standard package, as the standard package may not provide redundancy and lane repair support.

[0058] Referring now to FIG. 5, a schematic diagram illustrating a die-to-die connection according to one embodiment is shown. As shown in FIG. 5, a package 500 includes a first die 510 and a second die 560. An interconnect 540, e.g., a UCIe interconnect, includes a plurality of sideband lanes, i.e., sideband lanes 541-544. Although a single direction of the sideband lanes is shown, it should be understood that a corresponding set of sideband lanes may be provided in other directions as well. The first die 510 includes a sideband data transmitter and a sideband clock transmitter, i.e., sideband data transmitters 511, 512 (sideband data transmitter 512 is a redundant transmitter). The first die 510 further includes sideband clock transmitters 514, 515 (sideband clock transmitter 515 is a redundant transmitter). The second die 560 consequently includes a sideband data receiver and a sideband clock receiver, i.e., sideband data receivers 561, 562 (sideband data receiver 562 is a redundant receiver). The second die 560 further includes sideband clock receivers 564, 565 (sideband clock receiver 565 is a redundant receiver).

[0059] 5, there is detection circuitry in the second die 560 that may be used to perform sideband detection that may be part of the sideband initialization to determine which lanes are included in the functional sideband and which lanes may be part of the redundant sideband. As shown, there are multiple detectors 570 0-3Each detector 570 receives an incoming sideband data signal and an incoming sideband clock signal such that each detector 570 receives signals from a different combination of sideband receivers of the second die 560. During sideband initialization, the incoming sideband data signal may be a predetermined sideband initialization packet that includes a predetermined pattern. The detector 570 may be configured to detect the presence of the pattern and generate a first result (e.g., logic 1) in response to a valid detection of the pattern (e.g., for a number of repetitions of the pattern) and generate a second result (e.g., logic 0) in response to the predetermined pattern not being detected. Although embodiments are not limited in this respect, in one implementation, the detector 570 may be configured with a shift register, counter, or the like to perform this detection operation and produce four combinations by sampling the data and redundant data using a clock signal and a redundant clock signal to generate corresponding results.

[0060] It should be noted that if the redundant sideband circuitry is not used for repair purposes, the redundant sideband circuitry may be used to increase the bandwidth of sideband communications, especially for data-intensive transfers. By way of example, a sideband according to an embodiment may be used to communicate large amounts of information to be downloaded, such as firmware and / or downloads. Or, the sideband may be used to communicate management information, for example according to a given management protocol. It should be noted that such communications may occur simultaneously with other sideband information communications in the functional sideband.

[0061] Referring now to FIG. 6A, a timing diagram illustrating sideband signaling according to one embodiment is shown. As shown in FIG. 6A, the timing diagram 600 includes a sideband clock signal 610 and a sideband message signal 620. The sideband message format may be defined as a 64-bit header that includes 32 or 64 bits of data communicated during 64 unit intervals (UIs). The sideband message signal 620 illustrates a 64-bit serial packet. The sideband data may be transmitted with edges aligned with a clock (strobe) signal. The receiver of the sideband interface samples the incoming data with the strobe. For example, the negative edge of the strobe may be used to sample the data when the data uses SDR signaling.

[0062] Referring now to FIG. 6B, a timing diagram illustrating sideband packet serial transmission according to one embodiment is shown. As shown in FIG. 6B, timing diagram 601 illustrates communication of a first sideband packet 622 followed by a second sideband packet 624. As shown, each packet may be a 64-bit serial packet transmitted for 64 UI duration. More specifically, a first sideband packet 622 is transmitted resulting in a logic low on both the clock lane and the data lane for a duration of 32 UI, after which a second sideband packet 624 is communicated. In an embodiment, such signaling may be used for various sideband communications, including sideband messages during sideband initialization.

[0063] Referring now to FIG. 7, a flow diagram illustrating a bring-up flow of an on-package multi-protocol capable interconnect according to one embodiment is shown. As shown in FIG. 7, the bring-up flow 700 begins by independently executing a reset flow for two dies (die 0 and 1) coupled together, for example, via a UCIe interconnect (shown as a D2D channel in FIG. 7). Thus, the first die (die 0) executes an independent reset flow at stage 710, and the second die (die 1) also executes an independent reset flow at stage 710. Note that each die may finish the reset flow at a different time. Next, at stage 720, sideband detection and training may be performed. At stage 720, sidebands may be detected and trained. In the case of an advanced package where lane redundancy is available, available lanes may be detected and used for sideband messages. Note that since each die may finish the reset flow at a different time as described above, this sideband detection and training, including sideband initialization as described herein, may be used to detect the presence of activity in the coupled die. In one or more embodiments, the trigger to exit reset and begin link training is the detection of a sideband message pattern. When training during link bring-up, if the physical layer transitions out of reset, the hardware is allowed to make multiple attempts to train. During this bring-up operation, Both Side Die About Status and Substate In and out of but , these A four-way sideband message handshake between dies ensures they are in lockstep Therefore, synchronization may occur. .

[0064] At stage 730, a training parameter exchange may be performed for the functional sideband, and main band training is performed. At stage 730, the main band is initialized, repaired, and trained. Finally, at stage 740, protocol parameter exchange may be performed for the sideband. At stage 740, the entire link may be initialized by determining the local die capabilities, parameter exchange with the remote die, and bring-up of the FDI that couples the corresponding protocol layer with the D2D adapter of the die. In one embodiment, the main band initializes by default at the lowest allowed data rate in the main band initialization, where repair and reversal detection are performed. The link speed then transitions to the highest common data rate detected through parameter exchange. After link initialization, the physical layer may be enabled to perform protocol flit transfers over the main band.

[0065] In one or more embodiments, different types of packets may be communicated over the sideband interface, including: (1) configuration (CFG) or register accesses, which may be memory-mapped reads or writes and may be 32-bit or 64-bit (b); (2) data-free messages, which may be link management (LM) or vendor-defined packets and do not carry additional data payload; and (3) data-containing messages, which may be parameter exchange (PE), link training related or vendor-defined and may carry 64b of data. Packets may carry a 5-bit opcode, a 3-bit source identifier (srcid) and a 3-bit destination identifier (dstid). The 5-bit opcode indicates the packet type and whether the packet carries 32b of data or 64b of data.

[0066] Flow Control and Data Integrity Sideband packets may be transferred across FDI, RDI or UCIe sideband links. Each of these has independent flow control. For each transmitter associated with an FDI or RDI, design time parameters of the interface may be used to determine the number of credits (up to a maximum of 32 credits) that are signaled by the receiver. Each credit corresponds to a 64-bit header and 64-bits of potentially associated data. Thus, there is only one type of credit for all sideband packets, regardless of how much data they carry. Every transmitter / receiver pair has an independent credit loop. For example, in RDI, for sideband packets sent from the adapter to the physical layer, credits are signaled from the physical layer to the adapter; and for sideband packets sent from the physical layer to the adapter, credits are signaled from the adapter to the physical layer. Transmitters check for available credits before sending register access requests and messages. Transmitters do not check for credits before sending register access completions, and receivers guarantee unconditional sinking for any register access completion packets. Messages carrying requests or responses consume credits in the FDI and RDI but are guaranteed to make forward progress by the receiver and not block behind a register access request. Both the RDI and FDI provide dedicated signals for sideband credit return across their interfaces. All receivers associated with the RDI and FDI check received messages for data or control parity errors and these errors are mapped to an Uncorrectable Internal Error (UIE) and transition the RDI to the LinkError state.

[0067] Referring now to Figure 8, a flow diagram of a link training state machine is shown according to one embodiment. As shown in Figure 8, method 800 is an example of link initialization performed by logical physical layer circuitry, which may include, for example, a link state machine. Table 1 is a high level description of the states of a link training state machine according to one embodiment, with details and actions taken in each state described below. [Table 1] [Table 1]

[0068] Referring to FIG. 8, the method 800 begins with a reset state 810. In one embodiment, the PHY remains in the reset state for a predetermined minimum duration (e.g., 4 ms) to allow various circuits, including the phase-locked loop (PLL), to stabilize. This state may be exited when the power supply is stable, the sideband clock is available and running, the main band and die-to-die adapter clocks are stable and available, the main band clock is set to the slowest IO data rate (e.g., 2 GHz at 4 GT / s), and a link training trigger has occurred. Control then proceeds to a sideband initialization (SBINIT) state 820, where sideband initialization may be performed. In this state, the sideband interface is initialized and repaired (if applicable). During this state, the main band transmitter may be tri-stated and the main band receiver is allowed to be disabled.

[0069] With further reference to FIG. 8, from the sideband initialization state 820, control proceeds to a main band initialization (MBINIT) state 830 where main band initialization is performed. In this state, the main band interface is initialized and repaired or degraded (if applicable). The main band data rate may be set to the lowest supported data rate (e.g., 4 GT / s). For advanced packages, interface interconnect repair may be performed. Sub-states in MBINIT allow for data, clock, trace and active lane detection and repair. For standard package interfaces where lane repair is not required, sub-states are used to check functionality at the lowest data rate and perform width degradation if required.

[0070] Next, at block 840, a Main Band Training (MBTRAIN) state 840 may be entered where main band link training may be performed. In this state, the operating speed is set and clock to data centering is performed. At higher speeds, additional calibrations such as receiver clock correction, transmit and receive deskew may be performed in sub-states to ensure link performance. Each sub-state is entered by the module and exited through a sideband handshake. If no specific action is required in a sub-state, the UCIe module is allowed to exit that state through a sideband handshake without performing the action of that sub-state. In one or more embodiments, this state may be common to advanced and standard package interfaces.

[0071] Control then passes to block 850, where a link initialization (LINKINIT) state occurs, in which link initialization may be performed. In this state, the die-to-die adapter completes initial link management before entering an active state on the RDI. Once the RDI is in the active state, the PHY clears its copy of the "start UCIe link training" bit from the link control register. In an embodiment, the linear feedback shift register (LFSR) is reset upon entering this state. In one or more embodiments, this state may be common to advanced and standard package interfaces.

[0072] Finally, control proceeds to the active state 860, where communication can take place in normal operation. More specifically, packets from higher layers can be exchanged between the two dies. In one or more embodiments, all data in this state can be scrambled using a scrambler LFSR.

[0073] With further reference to FIG. 8, note that during the active state 860, a transition to a retrain (PHYRETRAIN) state 870 may occur or a low power (L2 / L1) link state 880 may occur. As can be seen, depending on the level of the low power link state, one may proceed from exit to either the main band training state 840 or the reset state 810. The low power link state consumes less power than dynamic clock gating in the active state. This state may be entered when the RDI transitions to a power management state. If the local adapter requests RDI active or the remote link partner requests L1 exit, the PHY exits to the MBTRAIN.SPEEDIDLE state. In one or more embodiments, the L1 exit is coordinated with a corresponding L1 state exit transition on the RDI. If the local adapter requests RDI active or the remote link partner requests L2 exit, the PHY exits to the reset state. Note that the L2 exit may be coordinated with a corresponding L2 state exit transition in the RDI.

[0074] 8, if an error occurs during any of the bring-up states, control may proceed to block 890 where a training error state may be generated. This state may be used as a transient state due to some fatal or non-fatal event to return the state machine to a reset state. If sideband is active, a sideband handshake is performed with the link partner to enter the TRAINERROR state from any state other than SBINIT.

[0075] In one embodiment, a die may enter the PHYRETRAIN state for multiple reasons. The trigger may be from an adapter-directed PHY retrain or a PHY-initiated PHY retrain. The local PHY initiates the retrain when it detects a valid framing error. The remote die may request a PHY retrain, which results in the local PHY entering PHY retrain when it receives the request. This retrain state may also be entered if a change in the runtime link test control register during the MBTRAIN.LINKSPEED state is detected. While shown at this high level in the embodiment of FIG. 8, it should be understood that many variations and alternatives are possible.

[0076] 9, a flow diagram of further details of main band initialization according to one embodiment is shown. Method 900 may be implemented by a link state machine to perform main band initialization. As shown, this initialization progresses through a number of states including a parameter exchange state 910, a calibration state 920, a recovery clock state 930, a recovery verify state 940, an invert main band state 950 and finally a main band recovery state 960. After completion of this main band initialization, control proceeds to main band training.

[0077] In parameter exchange state 910, an exchange of parameters may take place to set up maximum negotiated speed and other PHY settings. In one embodiment, the following parameters may be exchanged with the link partner (e.g., per module): voltage swing; maximum data rate; clock mode (e.g., strobe or continuous clock); clock phase; and module ID. In state 920, any calibrations required (e.g., transmit duty cycle correction, receiver offset and Vref calibration) may be performed.

[0078] Next, in block 930, detection and repair (if needed) may be performed on the clock and trace lanes of the advanced package interface and the clock and trace lanes of the standard package interface to check their functionality. In block 940, the module may set the clock phase to the center of the data UI of its main band transmitter. The module partner uses the received transmit clock to sample the received valid. All data lanes may be held low during this state. This state may be used to detect and apply repair (if needed) to the valid lane.

[0079] Still referring to FIG. 9, block 950 is entered only if the clock and enabled lanes are functional. In this state, a data lane inversion is detected. All transmitters and receivers of the module are enabled. The module sets the transmit clock phase to the center of its main band data UI. The module partner uses the incoming transmit clock to sample the incoming data. The 16-bit "per-lane ID" pattern (not scrambled) is a lane-specific pattern with the lane ID of the corresponding lane.

[0080] 9, in block 960, which is entered only after successful lane reversal detection and application, all transmitters and receivers of the module are enabled. The module sets its clock phase to the center of its main band data UI. The module partner uses the main band receiver's incoming forwarding clock to sample the incoming data. In this state, main band lanes are detected and repaired if required for advanced package interfaces and for functionality checks and width reductions for standard package interfaces. In other words, if an error is detected in a lane, a redundant circuit can be enabled over the redundant lane.

[0081] In an exemplary embodiment, during bring-up and operation, several de-rate techniques may be used to enable the link to find an operational setting. First, if an error is detected (either during initial bring-up or function operation) but repair is not required, a rate reduction may be performed. Such a rate reduction mechanism may transition the link to the next lower allowable frequency; this is repeated until a stable link is established. Second, if repair is not possible (as in the case of a standard package link with no repair resources), a width reduction may be performed; as an example, a width reduction to a half-width configuration may be allowed. For example, a 16-lane interface may be configured to operate as an 8-lane interface.

[0082] Referring now to FIG. 10, a flow diagram of main band training according to one embodiment is shown. As shown in FIG. 10, a method 1000 may be implemented by a link state machine to perform main band training. In main band training, the main band data rate is set to the highest common data rate of the two connected devices. Data to Clock training, deskew and Vref training may be performed using multiple sub-states. As shown in FIG. 10, main band training goes through multiple states or sub-states. As shown, main band training begins by performing a valid reference voltage training state 1005. In state 1005, the receiver reference voltage (Vref) for sampling the incoming Valid is optimized. The main band data rate remains at the lowest supported data rate. The module partner sets the forward clock phase to the center of the data UI of its main band transmitter. The receiver module uses the forward clock to sample the pattern of the Valid signal. All data lanes are held low during the Valid lane reference voltage training. Control then proceeds to a data reference voltage state 1010, where the receiver reference voltage (Vref) for sampling the incoming data is optimized while the data rate remains at the lowest supported data rate (e.g., 4 GT / s). The transmitter sets the forwarding clock phase to the center of the data UI. An idle speed state 1015 then occurs, and frequency changes may be allowed in this electrical idle state; more specifically, the data rate may be set to the maximum common data rate determined in the previous state. Circuit parameters may then be updated in the transmitter and receiver calibration states (1020 and 1025).

[0083] With further reference to FIG. 10, various training states 1030, 1035, 1040 and 1045 may proceed to train the valid-to-clock training reference voltage level, the full data-to-clock training and the data receiver reference voltage, respectively. In state 1030, the valid-to-clock training is performed before the data lane training to ensure that the valid signal is functional. The receiver samples the pattern of valid using the forwarded clock. In state 1035, the module may optimize the reference voltage (Vref) to sample the incoming valid at the operating data rate. In state 1040, the module performs the full data-to-clock training (including valid) using the LFSR pattern. In state 1045, the module may optimize its data receiver reference voltage (Vref) to optimize the sampling of the incoming data at the operating data rate.

[0084] With further reference to FIG. 10, a receiver deskew state 1050 may then occur, which is a receiver-initiated training phase in which the receiver performs lane-to-lane deskew to improve timing margins. Another data training state 1055 may then occur, in which the module may re-center the clock and aggregate data if the module partner's receiver has performed per-lane deskew. Control may then proceed to link rate state 1060, where link stability at this operating data rate may be checked after the final sampling point is set in state 1055. If the link performance does not meet this data rate, it is slowed down to the next lower supported data rate and training is performed again. Depending on the outcome of such a state, main band training may be completed and control may then proceed to link initialization. Otherwise, a link speed change in either state 1015 or repair state 1065 may be made. It should also be noted that entry into states 1015 and 1065 may occur from a low power state (e.g., L1 link power state) or a retraining state. Although shown at this high level in the embodiment of FIG. 10, it should be understood that many variations and alternatives are possible.

[0085] In different implementations, a different number of redundant lanes may be provided. In one example, approximately 3 to 5% redundant lanes may be added in the PHY layer to recover the die if some functional lanes are damaged, for example during the package chiplet assembly process. Of course, there may be additional or fewer redundant lanes in a given implementation.

[0086] 11A, a diagram of data lane remapping possibilities is shown according to one embodiment. More specifically, in FIG. 11A, a portion of a die 1100 includes multiple data lanes 11100-1110. 31The data lanes 1110 may be the physical transmit data lanes onto which corresponding logical transmit data lanes are mapped. In a practical application, the data lanes 1110 may terminate in a series of bumps or other conductors on the exterior surface of the die, allowing for direct adaptation to a package substrate or interposer (or directly to another die).

[0087] As further shown in FIG. 11A, the arrowed lines represent intersecting tangents between different data lanes to enable the repair and inversion operations described herein.

[0088] FIG. 11A illustrates the corresponding redundant data lanes 1120. 0、1 This is in the context of logical remapping by a redundant data lane. This remapping may allow a defective data lane to be remapped to an adjacent lane. Such remapping may be done sequentially, such that the data of a given data lane is eventually provided to a corresponding redundant data lane. Thus, as shown in FIG. 11A, data traffic of data lane 11100 may be remapped to redundant data lane 11200, and similarly, data traffic of data lane 1110 may be remapped to redundant data lane 11200. 31 data traffic may be remapped to the redundant data lane 11201.

[0089] In one particular embodiment, the module may support remapping (repair) of up to two data lanes for each group of 32 data lanes (e.g., two redundant data lanes of the first set of physical data lanes (e.g., transmit and receive physical data lanes, TD_P[31:0] (RD_P[31:0])) and two redundant data lanes of the second set of physical data lanes (TD_P[63:32] (RD_P[63:32])). In this way, two separate groups of 32 lanes can each be independently repaired using redundant data lanes (TRD_P[1:0] (RRD_P[1:0]) and TRD_P[3:2] (RRD_P[3:2]). Although in FIG. 11A there are two redundant resources per 32 data lanes for functional data recovery, in other embodiments there may be additional redundant resources. In different embodiments there may be between approximately 1 and 6 redundant data lanes per set of 32 data lanes. It should also be understood that the size of the set of data lanes may vary in different implementations, for example, between 4 and 128 in some cases.

[0090] In one or more embodiments, lane remapping can be a "left shift" or a "right shift" operation This can be achieved by the left shift: operation A right shift occurs when data traffic of a logical lane TD_L[n] associated with a physical data lane TD_P[n] is multiplexed onto a different physical data lane TD_P[n-1]. operationThis occurs when data traffic of logical data lane TD_L[n] is multiplexed onto physical data lane TD_P[n+1]. After the data lanes are remapped, the physical layer may control and disable (e.g., tri-state) the transmitter associated with the corrupted physical lane and control and disable the corresponding receiver. As a result, the transmitter and receiver of the redundant lane used for repair are enabled. Both "left shift" and "right shift" remapping may be performed to optimally repair up to any two lanes in the group. Of course, additional lanes may be repaired even if additional redundant resources exist.

[0091] In addition to redundant data lane resources, there may be dedicated redundant clock lanes for differential clock circuits. In one embodiment, clock lane remapping allows for single lane failure repair for both differential and pseudo-differential implementations of the clock circuit. Similar redundant circuits may also be provided for tracking lanes.

[0092] Referring now to FIG. 11B, a representative remapping configuration for a single lane failure is shown. The numbering scheme in FIG. 11B follows that of FIG. 11A, but a different die 1101 is shown. As can be seen, the data lanes 1110 29An error in TD_P[0] (e.g., this physical lane was damaged during assembly) causes a remapping in the direction of the redundant data lane 11200, which serves as a repair resource. According to this scheme, in the embodiment of FIG. 11B, the redundant data lane 11200 (TRD_P[0] (RRD_P[0])) is used as a redundant lane for remapping any single physical lane failure of TD_P[31:0] (RD_P[31:0]), and TRD[2] (RRD[2]) is used as a redundant lane for remapping any single lane failure of TD_P[63:32] (RD_P[63:32]). Of course, this is equally possible for the redundant data lane 11201, which becomes a repair resource, with the remapping going in the opposite direction. Thus, the bidirectional shifting mechanism allows forward and backward shifting of any lane data to adjacent lanes, so that the device can support sufficient connections or bandwidth.

[0093] Referring now to Tables A and B, pseudocode representations of repairs in the lower and upper lanes according to one embodiment are shown. [Table A] Pseudocode for lane repair in TD_P[31:0](RD_P[31:0])(0<=x<=31): [Table 2] [Table B] Pseudocode for lane repair in TD_P[63:32](RD_P[63:32])(32<=x<=63): [Table 3]

[0094] 11C, a representative remapping configuration for a two lane failure in a module is shown. In this configuration, any two lanes in a group of 32 lanes may be repaired using redundant resources and full functional traffic may be restored. For example, as shown in FIG. 11C, physical data lanes 1110 and 1111 may be repaired using redundant resources and full functional traffic may be restored. 25and 1110 26 Suppose that the data is corrupted. To repair this lane, the data is 1110 25 From 1110 24 and all lanes below it are shifted towards the redundant lane 11200 (TRD_P0). To repair TD26, data is steered to TD27 and all lanes above TD27 are shifted towards the redundant lane 11201 (TRD_P1). Using the above bidirectional shifting scheme, any combination of single or dual lane damage can be fully recovered.

[0095] Thus, in one embodiment, for any two physical lane failures in TD_P[31:0] (RD_P[31:0]), the lower lane is remapped to TRD_P[0] (RRD_P[0]) and the upper lane is remapped to TRD_P[1] (RRD_P[1]). For any two physical lane failures in TD_P[63:32] (RD_P[63:31]), the lower lane is remapped to TRD_P[2] (RRD_P[2]) and the upper lane is remapped to TRD_P[3] (RRD_P[3]).

[0096] Referring now to Tables C and D, a pseudocode representation of two lane repair in the lower and upper lanes is shown according to one embodiment. Note that for all of the above examples, both the transmitter and the corresponding receiver apply the remapping shown. [Table C] Pseudocode for two lane repair in TD_P[31:0](RD_P[31:0])(0<=x,y<=31): [Table 4] [Table D] Pseudocode for two lane repair in TD_P[63:32](RD_P[63:32])(32<=x,y<=63): [Table 5]

[0097] Referring now to Figure 12A, a schematic diagram of a portion of a die-to-die connection according to one embodiment is shown. As shown in Figure 12A, a semiconductor package 1200 includes a first die and a second die. The first die includes a number of transmit logical data lanes TD_L[n-1, n+2]. Of course, each module or cluster may include more than these shown data lanes.

[0098] As can be seen, each data lane is connected to a corresponding lane repair multiplexer 1210. n-1、n+2 As further shown, data from adjacent logical data lanes in the left and right directions are provided via shift lines also coupled to multiplexer 1210. If repair is not required, multiplexer 1210 transmits the data of the corresponding logical data lane to multiple transmitters 1220 associated with the corresponding physical data lane. n-1、n+2 Alternatively, if repair is required, the corresponding left or right shift operation is performed to provide data of adjacent logical data lanes and multiplexer 1210 is controlled accordingly. Thus, multiplexer 1210 may be configured to select the corresponding (true) bit lane data {n} or the previous bit lane data {n-1} or the next bit lane data {n-2}.

[0099] With further reference to FIG. 12A, the data output by the transmitter 1220 is represented by a corresponding bump 1230. n-1、n+2 and interconnection 1240 n-1、n+ 2 to the second die. n-1、n+2 The data is sent to the receiver 1260 n-1、n+2 from which the corresponding lane repair multiplexer 1270 n-1、n+2 As shown, the corresponding left and right shifts operationThis allows the remapping operation to be performed on the receiver side as well. More specifically, the opposite remapping performed on the first die transmitter side can be performed on the second die receiver side (e.g., left shift operation If is done on the first die, then the corresponding right shift operation (This is done in the second die.) Although shown at this high level in FIG. 12A, many variations and alternatives are possible.

[0100] Referring now to FIG. 12B, the left shift operation is done at the sender to repair the defective data lane, and the corresponding right shift operation Single lane remapping in package 1201 is performed at the receiver operation In this case, the logical shift operation Note that only the shift operation It should be understood that this is not shown here in order to illustrate the

[0101] 12C, a two-lane remapping situation is shown. As can be seen, at the transmit side of package 1202, left and right shifts are performed to repair two bad data lanes. operation , and similarly, corresponding opposite right and left shift operations are performed at the receiving end. As mentioned above, FIG. 12C illustrates the logical shift operation Shows.

[0102] For standard package (e.g., ×16) modules where lane repair is not supported, resilience to defective lanes can be provided by configuring the link to a smaller (e.g., ×8) width (e.g., logical lanes 0 through 7 or logical lanes 8 through 15, which eliminates the defective lanes). For example, if one or more defective lanes are in logical lanes 0 through 7, the link is configured to a ×8 width using logical lanes 8 through 15. This configuration is done during link initialization or retraining, and the transmitters of the disabled lanes may be placed into a high impedance state (hi-Z) and the receivers are disabled.

[0103] The device may also be configured to support lane reversal within a module. An example of lane inversion is when physical data lane 0 on the local die is connected to physical data lane (N-1) on the remote die (physical data lane 1 is connected to physical data lane n-2, etc.), e.g., N=16 in the standard package and N=64 in the advanced package. The redundant lanes in the advanced package case may also be inverted. In one or more embodiments, lane inversion is implemented only for the transmitter. The transmitter reverses the logical lane order on the data and redundant data lanes. In one embodiment, lane inversion is discovered and applied during initialization and training. To enable lane inversion discovery, each logical data and redundant lane in a module is assigned a unique lane ID. In some embodiments, the tracking, enable, clock and sideband signals are not inverted.

[0104] Lane reversal according to one embodiment may use a multiplexer structure similar to that described above with respect to Figures 12A-12C, except that the multiplexer selects between the LSB and MSB bits. Also, lane shifting for lane reversal is only done on the transmit or receive side, not on both sides. For example, when lanes are reversed, physical lane 0 of the transmitter (TD_P[0]) is connected to physical lane N-1 of the receiver (RD_P[N-1]), where N is 64 for advanced package modules and 16 for standard package modules. Note that lane repair mapping may be changed when lane reversal is implemented.

[0105] When repairing a single lane with lane reversal, the transmitter side remapping is reversed to maintain the shift order of the receiver side remapping. Referring now to Table E, pseudocode is shown for repairing a single lane failure with a reversal in TD_P[31:0] (RD_P[32:63]) (0<=x<=31). [Table E] [Table 6]

[0106] Referring now to Table F, pseudocode for a one lane failure with inversion in TD_P[63:32] (RD_P[0:31]) (32<=x<=63) is shown. [Table F] [Table 7]

[0107] For two-lane repair with lane reversal, the transmitter side remapping is reversed to maintain the shift order of the receiver side remapping. Referring now to Table G, pseudocode for a two-lane failure with a reversal in TD_P[31:0] (RD_P[32:63]) (0<=x<=31) is shown. [Table G] [Table 8]

[0108] Referring now to Table H, pseudocode for a one lane failure with inversion in TD_P[63:32] (RD_P[0:31]) (32<=x<=63) is shown. [Table H] [Table 9]

[0109] The main band repair process may possibly be performed during main band initialization. This process may be performed in a repair state that is entered only after successful lane reversal detection and application. In this state, all transmitters and receivers of the module are enabled. The module sets the clock phase to the center of the main band data UI. The link partner uses the main band receiver incoming transmit clock to sample the incoming data. In this state, the main band lanes are detected and repaired if required for the advanced package interface and for functionality checks and width degradation for the standard package interface.

[0110] In one embodiment, the following sequence may be used for main band repair of the Advanced Package Interface.

[0111] 1. The module sends a sideband message {MBINIT.REPAIRMB start req} and waits for a response. The link partner responds with {MBINIT.REPAIRMB start resp}.

[0112] 2. The module performs transmitter initiated data-to-clock point training for that transmitter lane (transmit pattern having 128 repetitions of unscrambled "per-lane ID" pattern in continuous mode). The receiver performs per-lane comparison and detection on the receiver lane is deemed successful if at least a predefined number (e.g., 16 consecutive repetitions) of the "per-lane ID" pattern is detected.

[0113] 3. At the end of the transmitter initiated data-to-clock point test, the module receives per-lane pass / fail information via sideband message.

[0114] 4. If lane repair is required and repair resources are available, the module applies the repair to its main band transmitter and sends a {MBINIT.REPAIRMB Apply repair req} sideband message. On receiving this sideband message, the link partner applies the repair to its main band receiver and sends a {MBINIT.REPAIRMB Apply repair resp} sideband message. If the number of lane defects is greater than the repair capability, the main band is not repairable and the module exits to the TRAINERROR state after performing a TRAINERROR handshake.

[0115] 5. If no repair is required, perform step 7.

[0116] 6. If a lane repair has been applied (step 4), the applied repair is checked by the module by repeating steps 2 and 3. If a lane error after the repair was logged in step 5, the module performs a TRAINERROR handshake and then exits to TRAINERROR. If the repair is successful, step 7 is executed.

[0117] 7. The module sends a {MBINIT.REPAIRMB end req} sideband message and the link partner responds with {MBINIT.REPAIRMB end resp}. When the module has sent and received {MBINIT.REPAIRMB end resp} it exits on MBTRAIN.

[0118] Although this embodiment is described using this particular implementation, variations may be made in other embodiments, for example, a similar process may be used to perform repairs after retraining or link speed sub-states.

[0119] In the case of a standard package interface, the main band is checked for function operation at the lowest data rate. Broadly the same steps as described above may be followed. However, if an error is identified in a data lane, it is determined whether a width degradation is possible. If so, the module with the defective transmitter lane applies the degradation (to both its transmitter and receiver) and sends a message {MBINIT.REPAIRMB apply degrade req} containing the logical lane map to the remote link partner. The link partner applies the degradation (to both its transmitter and receiver) and sends a message {MBINIT.REPAIRMB apply degrade resp}.

[0120] In one embodiment, for a standard package interface, if the number of lanes experiencing an error are all contained within lanes 0-7 or lanes 8-15, the width is reduced to a x8 link (lanes 0...lane 7 or lane 8...lane 15).

[0121] Referring now to Figure 13, a flow diagram of a lane repair method according to one embodiment is shown. As shown in Figure 13, method 1300 may be performed during link initialization. Additionally, while method 1300 is in the context of main band data lane repair, the concepts described herein apply equally to other lane repair situations, including clock lanes, sideband lanes, and tracking lanes. Method 1300 may be performed, at least in part, via physical layer circuitry.

[0122] As shown in FIG. 13, the method 1300 begins with the transmission of a predetermined pattern of data-to-clock point training (block 1310). This predetermined pattern, which may be a per-lane pattern such as a per-lane ID pattern, may be transmitted continuously via a transmitter associated with each physical data lane. For example, 128 iterations of this pattern may be transmitted. Next, at block 1320, result information may be received from the second die via the sideband. This result information may include per-lane pass / fail information.

[0123] 13, it is next determined at diamond 1330 whether at least one data lane failure has been identified. If not, control proceeds to block 1340 where the Repair Detection state, which is a sub-state of the Initialization state of the link training state machine, may be exited by sending and receiving a sideband message indicating an end of the Repair state. Control then proceeds to block 1350 where this state is exited to the Training state.

[0124] 13, if instead a fault is identified, control passes to diamond 1350 which determines whether this fault is identified in the post-repair situation. If so, an error is raised by a sideband message and control passes to block 1390 which ends this repair substate into the training error state.

[0125] If the failure is an initial failure, control proceeds to diamond 1370, where it is determined whether repair resources (e.g., including sufficient redundant lanes to accommodate the number of failed data lanes) are available. If so, control proceeds to block 1380 for application of lane repair. More specifically, in block 1380, in the transmit direction, the physical layer circuitry may apply repair to one or more transmitters associated with the data lanes to remap data traffic of at least one logical data lane onto at least one other physical lane. It should be appreciated that similar lane repair may be performed on the second die by appropriate remapping of receivers, with communication of information regarding the defective lane, such that the correct data traffic is provided on the intended logical data lane. Although shown at this high level in the embodiment of FIG. 13, many variations and alternatives are possible.

[0126] It should be noted that in various embodiments, one or more of the features described herein may be configurable to be enabled or disabled, for example under dynamic user control, based on information stored in one or more configuration registers (which may be present, for example, in one or more of the D2D adapter or physical layers). In addition to dynamic (or boot-time) enabling or disabling of various features, it is also possible to provide configurability regarding the operational parameters of certain aspects of UCIe communications.

[0127] Embodiments may support two broad usage models. The first is package-level integration for power-saving and cost-effective performance. For example, board-level attached components such as memory, accelerators, networking devices, modems, etc. may be integrated at the package level, with applicability from handheld to high-end servers. In such use cases, die from potentially multiple sources may be connected through different packaging options, even on the same package.

[0128] The second use is to provide off-package connectivity using different types of media (e.g., optical, electrical cable, mmWave) using UCIe retimers to enable resource pooling, resource sharing and / or message passing to transport the underlying protocols (e.g., PCIe, CXL) at the rack or pod level with load-store semantics beyond the node level to the rack / pod level to provide better power saving and cost effective performance at the edge and data center.

[0129] As mentioned above, embodiments may be implemented in a data center use case, for example in connection with a rack or pod. As an example, multiple compute nodes from different compute chassis may connect to a CXL switch. In turn, the CXL switch may connect to multiple CXL accelerators / Type-3 memory devices, which may be located in one or more separate drawers.

[0130] Referring now to Figure 14, a block diagram of another exemplary system is shown in accordance with one embodiment. In Figure 14, system 1400 may be all or part of a rack-based server having multiple hosts in the form of compute drawers that may be coupled to pooled memory via one or more switches.

[0131] As shown, there are multiple hosts 1430-1-n (also referred to herein as "hosts 1430"). Each host may be implemented as a compute drawer having one or more SoCs, memory, storage and interface circuits, etc. In one or more embodiments, each host 1430 may include one or more virtual hierarchies corresponding to different cache coherence domains. The hosts 1430 may couple to a switch 1420, which may be implemented as a UCIe or CXL switch (e.g., a CXL 2.0 (or later) switch). In one embodiment, each host 1430 may couple to the switch 1420 using an off-package interconnect, e.g., a UCIe interconnect running the CXL protocol through at least one UCIe retimer (which may be present in one or both of the hosts 1430 and the switch 1420).

[0132] The switch 1420 may couple to multiple devices 1410-1-x (also referred to herein as "devices 1410"), each of which may be a memory device (e.g., a type 3 CXL memory expansion device) and / or an accelerator. In the example of FIG. 14, each device 1410 is shown as a type 3 memory device having any number of memory regions (e.g., defined partitions, memory ranges, etc.). Depending on the configuration and use case, a particular device 1410 may include memory regions assigned to a particular host, while others may include at least some memory regions designated as shared memory. Although embodiments are not limited in this respect, the memory included in the device 1410 may be implemented using any type of computer memory (e.g., dynamic random access memory (DRAM), static random access memory (SRAM), non-volatile memory (NVM), a combination of DRAM and NVM, etc.).

[0133] Referring now to Figure 15, a block diagram of a system according to another embodiment, such as an edge platform, is shown. As shown in Figure 15, a multiprocessor system 1500 includes a first processor 1570 and a second processor 1580 coupled via an interconnect 1550, which may be a UCIe interconnect according to one embodiment that implements a coherency protocol. As shown in Figure 15, each of the processors 1570 and 1580 may be a many core processor including representative first and second processor cores (i.e., processor cores 1574a and 1574b and processor cores 1584a and 1584b).

[0134] 15 embodiment, processors 1570 and 1580 further include point-to-point interconnects 1577 and 1587 that couple to switches 1559 and 1560 via interconnects 1542 and 1544 (which may be UCIe links according to one embodiment). In turn, switches 1559, 1560 couple to pooled memories 1555 and 1565 (e.g., via UCIe links).

[0135] With further reference to FIG. 15, the first processor 1570 further includes a memory controller hub (MCH) 1572, point-to-point (PP) interfaces 1576 and 1578. Similarly, the second processor 1580 includes an MCH 1582, PP interfaces 1586 and 1588. As shown in FIG. 15, the MCHs 1572 and 1582 couple the processors to respective memories, i.e., memories 1532 and 1534, which may be portions of system memory (e.g., DRAM) locally attached to the respective processors. The first processor 1570 and the second processor 1580 may be coupled to a chipset 1590 via PP interconnects 1576 and 1586, respectively. As shown in FIG. 15, the chipset 1590 includes PP interfaces 1594 and 1598.

[0136] Additionally, the chipset 1590 includes an interface 1592 for coupling the chipset 1590 to a high performance graphics engine 1538 via a PP interconnect 1539. As shown in FIG. 15, various input / output (I / O) devices 1514 may be coupled to the first bus 1516 along with a bus bridge 1518 that couples the first bus 1516 to a second bus 1520. In one embodiment, various devices may be coupled to the second bus 1520 including, for example, a keyboard / mouse 1522, a communication device 1526, and a data storage unit 1528 such as a disk drive or other mass storage device that may include code 1530. Additionally, an audio I / O 1524 may be coupled to the second bus 1520.

[0137] Referring now to FIG. 16, a block diagram of a system 1600 according to another embodiment is shown. As shown in FIG. 16, the system 1600 may be any type of computing device, and in one embodiment, may be a server system. In the embodiment of FIG. 16, the system 1600 includes multiple CPUs 1610a,b which in turn couple to respective system memories 1620a,b (which in embodiments may be implemented as DIMMs, such as double data rate (DDR) memory, persistent or other types of memory). It should be noted that the CPUs 1610 may be coupled together via an interconnect system 1615, such as UCIe, or other interconnect implementing a coherency protocol.

[0138] There may be multiple interconnects 1630a1-b2 to allow coherent accelerator devices and / or smart adapter devices to couple to the CPU 1610 via potentially multiple communication protocols. Each interconnect 1630 may be a given instance of a UCIe link according to one embodiment.

[0139] In the illustrated embodiment, each CPU 1610 couples to a corresponding field programmable gate array (FPGA) / accelerator device 1650a,b (which may include a GPU in one embodiment). In addition, the CPU 1610 also couples to smart NIC devices 1660a,b, which in turn couple to switches 1680a,b (e.g., CXL switches, in one embodiment). The switches in turn couple to pooled memory 1690a,b, such as persistent memory. In an embodiment, the various components illustrated in FIG. 16 may implement circuitry to perform the techniques described herein.

[0140] The following examples relate to further embodiments.

[0141] In one example, an apparatus includes a first die; The first die includes: a die-to-die adapter for communicating with a protocol layer circuit and a physical layer circuit, wherein the die-to-die adapter receives message information including first information of a first interconnection protocol; and the physical layer circuit coupled to the die-to-die adapter, where the physical layer circuit receives and outputs the first information to a second die via an interconnect; having The physical layer circuit includes: a first plurality of transmitters for transmitting data over a first plurality of data lanes; and at least one redundant transmitter, wherein the physical layer circuitry remaps a first data lane of the first plurality of data lanes to the at least one redundant transmitter. Includes.

[0142] In one example, the device comprises: a first plurality of bumps fitted on the first die, where the first plurality of bumps are associated with the first plurality of data lanes; and at least one redundant bump fitted on the first die, where the physical layer circuitry remaps the first data lane from a first bump of the plurality of bumps to the at least one redundant bump. It further comprises:

[0143] In one example, The at least one redundant transmitter comprises: a first redundant transmitter, where the physical layer circuitry remaps the first data lane to the first redundant transmitter to repair a lane failure in the first data lane; and a second redundant transmitter, wherein the physical layer circuitry remaps a second data lane of the first plurality of data lanes to the second redundant transmitter to repair a lane failure in the second data lane. Includes.

[0144] In one example, the physical layer circuitry repairs two data lanes of a group of 32 data lanes via the first redundant transmitter and the second redundant transmitter.

[0145] In one example, the apparatus further comprises a first plurality of multiplexers coupled to the first plurality of transmitters, where the physical layer circuitry controls the first plurality of multiplexers to pass data from one of a corresponding data lane, a first adjacent data lane, or a second adjacent data lane.

[0146] In one example, The physical layer circuit includes: shifting left using a first portion of the first plurality of multiplexers operation is executed to repair the lane failure in the first data lane; and shifting right using a second portion of the first plurality of multiplexers; operationis executed so that the lane failure in the second data lane is repaired.

[0147] In one example, the device comprises: a first plurality of receivers for receiving second message information over a second plurality of data lanes; and at least one redundant receiver, where in response to a failure in a first data lane of the second plurality of data lanes, the physical layer circuitry remaps the second data lane of the second plurality of data lanes to the at least one redundant receiver. It further comprises:

[0148] In one example, the physical layer circuitry comprises: a first clock transmitter for transmitting a clock signal over a first clock lane; and at least one redundant clock transmitter, wherein the physical layer circuitry remaps the first clock lane to the at least one redundant transmitter. It further has:

[0149] In one example, the physical layer circuitry reverses the logical lane order of at least some of the first plurality of data lanes.

[0150] In one example, when a first data lane associated with a first transmitter of the first plurality of transmitters is coupled to an Nth data lane of the second die, the physical layer circuitry reverses the logical lane order, where N is equal to the number of data lanes in a module.

[0151] In one example, the physical layer circuitry remaps at least one data lane of the first plurality of data lanes in response to a failure in the first data lane and reverses the logical lane order of at least some of the first plurality of data lanes.

[0152] In another example, a method includes identifying, via a physical layer circuit of a first die of a package including a first die and a second die and an interconnect coupling the first die and the second die, a fault in a first physical data lane of a first plurality of physical data lanes of a main band of the interconnect, the interconnect including the main band and a side band; in response to identifying the fault, remapping, via the physical layer circuitry of the first die, a first data traffic of a first logical data lane onto a second physical data lane of the main band; and communicating information regarding the remapping to the second die via the sideband; Equipped with.

[0153] In one example, the method further includes remapping the first data traffic of the first logical data lane from the second physical data lane in the second die to the first logical data lane via physical layer circuitry of the second die.

[0154] In one example, the method comprises: identifying, via the physical layer circuitry of the first die, a fault in a clock physical data lane of the interconnect; remapping a clock signal from the clock physical data lane to a redundant clock physical data lane via the physical layer circuitry of the first die in response to identifying the fault in the clock physical data lane; and communicating information regarding the remapping of the clock signal to the second die via the sideband; It further comprises:

[0155] In one example, the method comprises: providing the first data traffic of the first logical data lane to a first multiplexer of the first die; and controlling the first multiplexer via the physical layer circuitry of the first die to provide the first data traffic to a second transmitter of the first die associated with the second physical data lane. It further comprises:

[0156] In one example, the method further comprises disabling, via the physical layer circuitry, a first transmitter of the first die associated with the first physical data lane.

[0157] In another example, a computer readable medium comprising instructions performs the method of any of the above examples.

[0158] In a further example, a computer readable medium comprising data is used by at least one machine to manufacture at least one integrated circuit that performs the method of any one of the above examples.

[0159] In another further example, an apparatus comprises means for performing the method of any one of the above examples.

[0160] In another example, a package may include a first die having a CPU and a protocol stack, and a second die coupled to the first die via an interconnect, the first die may include a die-to-die adapter for communicating with a protocol layer circuit via an FDI and a physical layer circuit via an RDI, where the die-to-die adapter communicates message information including first information of a first interconnect protocol; a physical layer circuit coupled to the die-to-die adapter via the RDI, where the physical layer circuit receives and outputs first information to the second die via the interconnect, where the physical layer circuit includes a first plurality of receivers for receiving data via a first plurality of physical data lanes; and at least one redundant receiver, where the physical layer circuit shifts data traffic of a first plurality of physical data lanes to adjacent ones of the first plurality of physical data lanes and at least one redundant lane associated with the at least one redundant receiver in response to a lane failure in a first physical lane of the first plurality of physical data lanes.

[0161] In one example, the physical layer circuitry enables at least one redundant transmitter and tri-states a first transmitter of the first plurality of transmitters associated with another physical lane in response to another lane failure.

[0162] In one example, the interconnect includes a main band including the first plurality of physical data lanes, and a sideband, and the physical layer circuitry receives information regarding the shifting of the data traffic from the second die via the sideband.

[0163] In one example, the second die includes an accelerator, where the first die communicates with the second die according to at least one of a flit mode of a PCIe protocol or a flit mode of a CXL protocol.

[0164] In another example, the apparatus comprises: an interconnect means including a main band and a side band, the side band coupling a first die means and a second die means, means for identifying a fault in a first physical data lane means of a first plurality of physical data lane means of the main band of the interconnect means; means for remapping first data traffic of a first logical data lane means onto a second physical data lane means of said main band; means for communicating information regarding the remapping to the second die means via the sideband; Equipped with.

[0165] In one example, the apparatus further comprises means for remapping the first data traffic of the first logical data lane means from the second physical data lane means in the second die means to the first logical data lane means.

[0166] In one example, the device comprises: means for identifying faults in a clock physical data lane means of said interconnect means; means for remapping a clock signal from said clock physical data lane means to a redundant clock physical data lane means; and means for communicating information regarding the remapping of the clock signal to the second die means via the sideband; It further comprises:

[0167] In one example, the device comprises: means for providing the first data traffic of the first logical data lane means to a first multiplexer means of the first die means; and means for controlling the first multiplexer means to provide the first data traffic to a second transmitter means of the first die means associated with the second physical data lane means; It further comprises:

[0168] In one example, the apparatus further comprises means for disabling a first transmitter means of the first die means associated with the first physical data lane means.

[0169] It should be understood that various combinations of the above examples are possible.

[0170] It should be noted that the terms "circuit" and "circuitry" are used interchangeably herein. These terms and the term "logic" are used herein to refer to analog circuits, digital circuits, hardwired circuits, programmable circuits, processor circuits, microcontroller circuits, hardware logic circuits, state machine circuits, and / or any other type of physical hardware components, alone or in any combination. Each embodiment may be used in many different types of systems. For example, in one embodiment, a communications device may be arranged to perform the various methods and techniques described herein. Of course, the scope of the invention is not limited to communications devices. Instead, other embodiments may be directed to other types of apparatus for processing instructions, or one or more machine-readable media containing instructions that, in response to being executed on a computing device, cause the device to perform one or more of the methods and techniques described herein.

[0171] The embodiments may be implemented in code and may be stored on a non-transitory storage medium that includes instructions that may be used to program a system to execute the instructions. The embodiments may be implemented in data and may be stored on a non-transitory storage medium. The non-transitory storage medium, when used by at least one machine, causes the at least one machine to manufacture at least one integrated circuit that performs one or more operations. Another further embodiment may be implemented in a computer-readable storage medium that includes information that, when manufactured as an SoC or other processor, configures the SoC or other processor to perform one or more operations. The storage medium may include, but is not limited to, a floppy disk, an optical disk, a solid state drive (SSD), a compact disk read only memory (CD-ROM), a compact disk rewriteable (CD-RW), any type of disk including a magneto-optical disk, a read only memory (ROM), a random access memory (RAM) such as a dynamic random access memory (DRAM) or a static random access memory (SRAM), a semiconductor device such as an erasable programmable read only memory (EPROM), a flash memory, an electrically erasable programmable read only memory (EEPROM), a magnetic or optical card, or any other type of medium suitable for storing electronic instructions.

[0172] While the present disclosure has been described with respect to a limited number of implementations, those skilled in the art having the benefit of this disclosure will appreciate numerous modifications and variations therefrom, and it is intended that the appended claims cover all such modifications and variations. [Item 1] An apparatus comprising a first die, The first die includes: a die-to-die adapter for communicating with a protocol layer circuit and a physical layer circuit, wherein the die-to-die adapter receives message information including first information of a first interconnection protocol; and the physical layer circuit coupled to the die-to-die adapter, where the physical layer circuit receives and outputs the first information to a second die via an interconnect; having The physical layer circuit includes: a first plurality of transmitters for transmitting data over a first plurality of data lanes; and at least one redundant transmitter, wherein the physical layer circuitry remaps a first data lane of the first plurality of data lanes to the at least one redundant transmitter. Including, Device. [Item 2] a first plurality of bumps fitted on the first die, where the first plurality of bumps are associated with the first plurality of data lanes; and at least one redundant bump fitted on the first die, where the physical layer circuitry remaps the first data lane from a first bump of the plurality of bumps to the at least one redundant bump. Item 1. The apparatus of item 1, further comprising: [Item 3] The at least one redundant transmitter comprises: a first redundant transmitter, where the physical layer circuitry remaps the first data lane to the first redundant transmitter to repair a lane failure in the first data lane; and a second redundant transmitter, wherein the physical layer circuitry remaps a second data lane of the first plurality of data lanes to the second redundant transmitter to repair a lane failure in the second data lane. Including, Item 1. The device according to item 1. [Item 4] 4. The apparatus of claim 3, wherein the physical layer circuitry repairs two data lanes of a group of 32 data lanes via the first redundant transmitter and the second redundant transmitter. [Item 5] 2. The apparatus of claim 1, further comprising a first plurality of multiplexers coupled to the first plurality of transmitters, wherein the physical layer circuitry controls the first plurality of multiplexers to pass data from one of a corresponding data lane, a first adjacent data lane, or a second adjacent data lane. [Item 6] The physical layer circuit includes: shifting left using a first portion of the first plurality of multiplexers operation is executed to repair the lane failure in the first data lane; and shifting right using a second portion of the first plurality of multiplexers; operation is executed so that the lane failure on the second data lane is repaired. Item 5. The device according to item 5. [Item 7] a first plurality of receivers for receiving second message information over a second plurality of data lanes; and at least one redundant receiver, where in response to a failure in a first data lane of the second plurality of data lanes, the physical layer circuitry remaps the second data lane of the second plurality of data lanes to the at least one redundant receiver. Item 1. The apparatus of item 1, further comprising: [Item 8] The physical layer circuit includes: a first clock transmitter for transmitting a clock signal over a first clock lane; and at least one redundant clock transmitter, wherein the physical layer circuitry remaps the first clock lane to the at least one redundant transmitter. Further comprising 8. The device according to any one of items 1 to 7. [Item 9] 2. The apparatus of claim 1, wherein the physical layer circuitry reverses the logical lane order of at least some of the first plurality of data lanes. [Item 10] 10. The apparatus of claim 9, wherein when a first data lane associated with a first transmitter of the first plurality of transmitters is coupled to an Nth data lane of the second die, where N is equal to a number of data lanes in a module, the physical layer circuitry reverses the logical lane order. [Item 11] 11. The apparatus of claim 1, wherein the physical layer circuitry remaps at least one data lane of the first plurality of data lanes in response to a failure in the first data lane and reverses a logical lane order of at least some of the first plurality of data lanes. [Item 12] identifying, via a physical layer circuit of a first die of a package including a first die and a second die and an interconnect coupling the first die and the second die, a fault in a first physical data lane of a first plurality of physical data lanes of a main band of the interconnect, the interconnect including the main band and a side band; in response to identifying the fault, remapping, via the physical layer circuitry of the first die, a first data traffic of a first logical data lane onto a second physical data lane of the main band; and communicating information regarding the remapping to the second die via the sideband; A method for providing the above. [Item 13] 13. The method of claim 12, further comprising remapping the first data traffic of the first logical data lane from the second physical data lane in the second die to the first logical data lane via a physical layer circuit of the second die. [Item 14] identifying, via the physical layer circuitry of the first die, a fault in a clock physical data lane of the interconnect; remapping a clock signal from the clock physical data lane to a redundant clock physical data lane via the physical layer circuitry of the first die in response to identifying the fault in the clock physical data lane; and communicating information regarding the remapping of the clock signal to the second die via the sideband; Item 13. The method of item 12, further comprising: [Item 15] providing the first data traffic of the first logical data lane to a first multiplexer of the first die; and controlling the first multiplexer via the physical layer circuitry of the first die to provide the first data traffic to a second transmitter of the first die associated with the second physical data lane. Item 13. The method of item 12, further comprising: [Item 16] Item 16. The method of item 15, further comprising disabling, via the physical layer circuitry, a first transmitter of the first die associated with the first physical data lane. [Item 17] 17. A computer-readable storage medium comprising computer-readable instructions for implementing the method of any one of items 12 to 16 when executed. [Item 18] 17. An apparatus comprising means for carrying out the method according to any one of items 12 to 16. [Item 19] 1. A package comprising a first die having a central processing unit (CPU) and a protocol stack and a second die coupled to the first die via an interconnect, The first die includes: a die-to-die adapter for communicating with a protocol layer circuit via a flit-aware die-to-die interface (FDI) and with a physical layer circuit via a raw die-to-die interface (RDI), wherein the die-to-die adapter communicates message information including first information of a first interconnection protocol; and the physical layer circuit coupled to the die-to-die adapter via the RDI, where the physical layer circuit receives and outputs the first information to the second die via the interconnect; having The physical layer circuit includes: a first plurality of receivers for receiving data over a first plurality of physical data lanes; and at least one redundant receiver, where the physical layer circuitry is configured to shift data traffic of a first plurality of physical data lanes to adjacent ones of the first plurality of physical data lanes and to at least one redundant lane associated with the at least one redundant receiver in response to a lane failure in a first physical lane of the first plurality of physical data lanes; Including, package. [Item 20] 20. The package of claim 19, wherein the physical layer circuitry enables at least one redundant transmitter and tri-states a first transmitter of a first plurality of transmitters associated with another physical lane in response to another lane failure. [Item 21] 20. The package of claim 19, wherein the interconnect includes a main band including the first plurality of physical data lanes, and a side band, and the physical layer circuitry receives information regarding the shifting of the data traffic from the second die via the side band. [Item 22] 20. The package of claim 19, wherein the second die has an accelerator, and wherein the first die communicates with the second die according to at least one of a Peripheral Component Interconnect Express (PCIe) protocol flit mode or a Compute Express Link (CXL) protocol flit mode. [Item 23] an interconnect means including a main band and a side band, the side band coupling a first die means and a second die means, means for identifying a fault in a first physical data lane means of a first plurality of physical data lane means of the main band of the interconnect means; means for remapping first data traffic of a first logical data lane means onto a second physical data lane means of said main band; means for communicating information regarding the remapping to the second die means via the sideband; An apparatus comprising: [Item 24] 24. The apparatus of claim 23, further comprising: means for remapping the first data traffic of the first logical data lane means from the second physical data lane means in the second die means to the first logical data lane means. [Item 25] means for identifying faults in a clock physical data lane means of said interconnect means; means for remapping a clock signal from said clock physical data lane means to a redundant clock physical data lane means; and means for communicating information regarding the remapping of the clock signal to the second die means via the sideband; 24. The apparatus of item 23, further comprising: [Item 26] means for providing the first data traffic of the first logical data lane means to a first multiplexer means of the first die means; and means for controlling the first multiplexer means to provide the first data traffic to a second transmitter means of the first die means associated with the second physical data lane means; 24. The apparatus of item 23, further comprising:

Claims

1. 1. An apparatus comprising a first die, The first die includes: a die-to-die adapter in communication with the protocol layer circuit and the physical layer circuit, the die-to-die adapter receiving message information including first information of a first interconnection protocol; the physical layer circuit coupled to the die-to-die adapter, the physical layer circuit receiving and outputting the first information to a second die via an interconnect; The physical layer circuit includes: a first plurality of transmitters for transmitting data over a first plurality of data lanes; and redundant transmitters provided at start and end positions of the first plurality of data lanes, respectively, wherein the physical layer circuitry remaps first data traffic of a first data lane of the first plurality of data lanes to an adjacent data lane in a first direction of the first data lane or a second direction opposite to the first direction, and shifts at least one data lane arranged in the first direction or the second direction relative to the adjacent data lane toward a direction of one of the redundant transmitters. Device.

2. a first plurality of bumps fitted on the first die, the first plurality of bumps being associated with the first plurality of data lanes; and at least one redundant bump fitted on the first die, the physical layer circuitry remapping the first data lane from a first bump of the first plurality of bumps to the at least one redundant bump; The apparatus of claim 1 further comprising:

3. The redundant transmitter comprises: a first redundant transmitter, the physical layer circuitry re-mapping the first data traffic of the first data lane to the first redundant transmitter to repair a lane failure in the first data lane; a second redundant transmitter, the physical layer circuitry re-mapping second data traffic of a second data lane of the first plurality of data lanes to the second redundant transmitter to repair a lane failure in the second data lane. Including, 2. The apparatus of claim 1.

4. 4. The apparatus of claim 3, wherein the physical layer circuitry repairs two data lanes in a group of 32 data lanes via the first redundant transmitter and the second redundant transmitter.

5. 2. The apparatus of claim 1, further comprising: a first plurality of multiplexers coupled to the first plurality of transmitters, the physical layer circuitry controlling the first plurality of multiplexers to pass data from one of a data lane of the first plurality of data lanes corresponding to each multiplexer, an adjacent data lane to the left of each multiplexer, or an adjacent data lane to the right of each multiplexer.

6. The physical layer circuit includes: performing a left shift operation using a first portion of the first plurality of multiplexers such that the lane failure in the first data lane is repaired; performing a right shift operation using a second portion of the first plurality of multiplexers such that the lane failure in the second data lane is repaired.

4. The apparatus of claim 3.

7. a first plurality of receivers for receiving second message information over a second plurality of data lanes; at least one redundant receiver, the physical layer circuitry being responsive to a failure in a first data lane of the second plurality of data lanes to remap the first data lane of the second plurality of data lanes toward the at least one redundant receiver; The apparatus of claim 1 further comprising:

8. The physical layer circuit includes: a first clock transmitter for transmitting a clock signal over a first clock lane; at least one redundant clock transmitter, the physical layer circuitry remapping the first clock lane to the at least one redundant transmitter; Further comprising 8. Apparatus according to any one of claims 1 to 7.

9. 2. The apparatus of claim 1, wherein the physical layer circuitry reverses a logical lane order of at least some of the first plurality of data lanes.

10. 10. The apparatus of claim 9, wherein the physical layer circuitry reverses the logical lane order when the first data lane associated with a first transmitter of the first plurality of transmitters is coupled to an Nth data lane of the second die, where N is equal to a number of data lanes in a module.

11. The physical layer circuit includes: remapping at least one data lane of the first plurality of data lanes in response to a fault in the first data lane; 11. The apparatus of claim 10, further comprising: reversing the logical lane order of the at least some data lanes of the first plurality of data lanes.

12. identifying, via a physical layer circuit of a first die of a package including a first die, a second die, and an interconnect coupling the first die and the second die, a fault in a first physical data lane of a first plurality of physical data lanes of a main band of the interconnect, the interconnect including the main band and a side band, the first plurality of physical data lanes including redundant physical data lanes disposed at a start position and an end position, respectively; in response to identifying the fault, remapping, via the physical layer circuitry of the first die, a first data traffic of a first logical data lane onto a second physical data lane adjacent in a first direction of the first physical data lane or a second direction opposite to the first direction, and shifting at least one physical data lane aligned in the first direction or the second direction relative to the adjacent second physical data lane toward one of the redundant physical data lanes; and communicating information regarding the remapping to the second die via the sideband; A method for providing the above.

13. 13. The method of claim 12, further comprising: remapping the first data traffic of the first logical data lane from a second physical data lane in the second die to the first logical data lane via physical layer circuitry of the second die.

14. identifying, via the physical layer circuitry of the first die, a fault in a clock physical data lane of the interconnect; remapping a clock signal from the clock physical data lane to a redundant clock physical data lane via the physical layer circuitry of the first die in response to identifying the fault in the clock physical data lane; and communicating information regarding the remapping of the clock signal to the second die via the sideband; The method of claim 12 further comprising:

15. providing the first data traffic of the first logical data lane to a first multiplexer of the first die; and controlling, via the physical layer circuitry of the first die, the first multiplexer to provide the first data traffic to a second transmitter of the first die associated with the second physical data lane; The method of claim 12 further comprising:

16. 16. The method of claim 15, further comprising disabling, via the physical layer circuitry, a first transmitter of the first die associated with the first physical data lane.

17. A computer program product for causing a processor to implement the method according to any one of claims 12 to 16.

18. 20. A computer readable storage medium storing a computer program according to claim 17.

19. Apparatus comprising means for carrying out the method according to any one of claims 12 to 16.

20. a first die having a central processing unit (CPU) and a protocol stack; a second die coupled to the first die via an interconnect, The first die includes: a die-to-die adapter communicating with a protocol layer circuit via a flit-aware die-to-die interface (FDI) and communicating with a physical layer circuit via a load die-to-die interface (RDI), the die-to-die adapter communicating message information including first information of a first interconnection protocol; the physical layer circuit coupled to the die-to-die adapter via the RDI, the physical layer circuit receiving the first information via the interconnect and outputting the first information to the second die; having The physical layer circuit includes: a first plurality of receivers for receiving data over a first plurality of physical data lanes; redundant receivers respectively disposed at a beginning and an end of the first plurality of physical lanes, wherein the physical layer circuitry, in response to a lane failure in a first physical data lane of the first plurality of physical data lanes, remaps data traffic of the first physical data lane to an adjacent physical data lane in a first direction of the first physical data lane or in a second direction opposite to the first direction, and shifts at least one physical data lane aligned in the first direction or the second direction relative to the adjacent physical data lane toward a redundant lane associated with one of the redundant receivers; Including, package.

21. 21. The package of claim 20, wherein the physical layer circuitry is responsive to another lane failure to enable at least one redundant transmitter and tri-state a first transmitter of the first plurality of transmitters associated with another physical lane.

22. 21. The package of claim 20, wherein the interconnect includes a main band and a side band, the main band including the first plurality of physical lanes, and the physical layer circuitry receives information regarding the shifting of the data traffic from the second die via the side band.

23. 21. The package of claim 20, wherein the second die has an accelerator, and the first die communicates with the second die according to at least one of a Peripheral Component Interconnect Express (PCIe) protocol flit mode or a Compute Express Link (CXL) protocol flit mode.

24. A method for identifying a fault in a first physical data lane of a first plurality of physical data lanes of a main band included in an interconnect coupling a first die and a second die, the interconnect including the main band and a side band, the first plurality of physical data lanes including redundant physical data lanes at respective start and end positions; means for remapping a first data traffic of a first logical data lane onto a second physical data lane adjacent in a first direction of the first physical data lane or a second direction opposite to the first direction, and shifting at least one physical data lane aligned in the first direction or the second direction relative to the adjacent second physical data lane toward one of the redundant physical data lanes; means for communicating information regarding the remapping to the second die via the sideband; An apparatus comprising:

25. 25. The apparatus of claim 24, further comprising: means for remapping the first data traffic of the first logical data lane from a second physical data lane in the second die to the first logical data lane.

26. means for identifying faults in a clock physical data lane of the interconnect; means for remapping a clock signal from the clock physical data lane to a redundant clock physical data lane; means for communicating information regarding the remapping of the clock signal to the second die via the sideband; 25. The apparatus of claim 24, further comprising: