Lane Repair and Lane Reversal Implementation for Device-to-Device (D2D) Interconnection
By integrating redundant hardware circuits and advanced interconnect technology, the semiconductor package can recover from single bump connection failures, enhancing yield and reducing waste and costs associated with advanced packaging technology.
Patent Information
- Application Number
- JP2023567003
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-06-20
- Filing Date
- 2022-11-22
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-11-22
AI Technical Summary
The high yield and low design cost of chiplet integration using advanced packaging technology are compromised by the failure of a single bump connection in the manufactured die, leading to waste and increased costs due to the necessity of discarding the entire packaged part.
Implementing a recovery scheme within the semiconductor package using advanced interconnect technology, including redundant hardware circuits and carefully designed distribution of redundant bumps and circuits, to improve yield and enable lane remapping and lane inversion between transmit and receive lanes.
The solution effectively recovers from yield loss due to lane connection problems, enables efficient lane repair, and eliminates the need for multiple tape-ins of the same die by allowing die rotation and mirroring, thereby reducing waste and costs.
Smart Images

Figure 0007693833000010 
Figure 0007693833000011 
Figure 0007693833000012
Abstract
Description
Background Art
[0001] Since the chiplets are integrated using advanced packaging technology, the yield is high and a low design cost can be maintained. However, if there is a single bump connection failure in the manufactured die, the entire packaged part will be discarded, which can lead to waste and cost problems.
Brief Description of the Drawings
[0002]
Figure 1
[0003]
Figure 2A
Figure 2B
Figure 2C
Figure 2D
[0004]
Figure 3A
Figure 3B
[0005]
Figure 4A
Figure 4B
[0006]
Figure 5
[0007]
Figure 6A
Figure 6B
[0008]
Figure 7
[0009]
Figure 8
[0010]
Figure 9
[0011]
Figure 10
[0012]
Figure 11A
[0013]
Figure 11B
[0014]
Figure 11C
[0015]
Figure 12A
[0016]
Figure 12B
[0017]
Figure 12C
[0018]
Figure 13
[0019]
Figure 14
[0020]
Figure 15
[0021]
Figure 16
[0022] In various embodiments, within a semiconductor package implemented in one or more die implementations may provide a recovery scheme as , using advanced interconnect technology to improve the yield of the packaged portion re and may include redundant hardware circuits. In an implementation, the distribution of redundant bumps and other redundant circuits and corresponding recovery mechanisms within the physical layer can be carefully designed to achieve maximum coverage with minimal overhead. Also, embodiments provide the ability to perform lane remapping and / or lane inversion between transmit and receive lanes, for example when these lanes are not connected in a bit-lane matching fashion due to floorplan constraints.
[0023] Embodiments implement very general means for solving all types of lane remapping problems, such as repair, remapping, etc., with minimal overhead to the bump area and performance. For this purpose, redundant lanes can be provided at the start and end of each module with sufficient flexibility to independently shift lane-by-lane data forward and backward on both the transmitting and receiving sides. Sufficient utilization of redundant lanes can be used to repair and recover any type of bump connection problem.
[0024] In various embodiments, a multi-protocol on-package interconnect can be used to communicate between the separate dies of a package. By initializing and training this interconnect with an ordered bring-up flow, independent reset of different dies, detection of "reset complete" of the partner die, and ordered initialization and training of the sideband and mainband interfaces of the interconnect can be enabled (in that order). More specifically, sideband initialization can be performed to detect that the link partner die has completed reset and to initialize and train the sideband. Thereafter, the mainband can be initialized and trained, which may include any lane inversion and / or repair operations further described herein. Such mainband operations can utilize the already brought-up sideband to communicate synchronization and status information.
[0025] In embodiments that perform lane inversion and / or repair, the yield loss due to lane connection problems in an advanced package multi-chip package (MCP) can be recovered. Further, by a lane repair technique according to one embodiment, both the left shift technique and the right shift technique can cover the entire bump map for efficient lane repair. Still further, lane inversion detection enables die rotation and die mirroring, which can enable multiple on-package instantiations on the same die. Thus, lane inversion can eliminate multiple tape-ins of the same die.
[0026] Embodiments may be implemented in connection with a multi - protocol - compliant on - package interconnect protocol that can be used to connect multiple chiplets or dies on a single package. In this interconnect protocol, the powerful ecosystems of separate die architectures can be interconnected together. This on - package interconnect protocol may be referred to as the "Universal Chiplet Interconnect Express (UCIe) interconnect protocol" and may be in accordance with the UCIe specification that may be issued by a Special Interest Group (SIG) or other promoter or other entity. Although referred to herein as "UCIe", it should be understood that the multi - protocol - compliant on - package interconnect protocol may adopt another term.
[0027] This UCIe interconnect protocol can support a plurality of underlying interconnect protocols including a flit - based mode of a particular communication protocol. In one or more embodiments, the UCIe interconnect protocol is, for example, the flit mode of the CXL protocol in accordance with a given version of the Compute Express Link (CXL) specification (CXL specification version 2.0 (released in November 2020), any future updates, versions or variations thereof, etc.); the flit mode of the PCIe protocol in accordance with a given version of, for example, the Peripheral Component Interconnect Express (PCIe) specification (PCIe - based specification version 6.0 (released in 2022) or any future updates, versions or variations thereof, etc.); and a low (or streaming) mode used to map any protocol supported by a link partner. In one or more embodiments, note that the UCIe interconnect protocol may not be backward - compatible and, instead, may be compatible with current and future versions of the above - mentioned protocols, or other protocols that support the flit mode of communication.
[0028] Embodiments can be used to provide compute, memory, storage, and connectivity across the compute continuum, spanning cloud, edge, enterprise, 5G, automotive, high performance computing, and handheld segments. Embodiments can be used to package or otherwise combine dies from different sources, including different fabs, different designs, and different packaging technologies.
[0029] Also, chiplet integration on package enables different trade - offs for different market segments by allowing customers to choose different numbers and types of dies. For example, depending on the segment, different numbers of compute, memory, and I / O dies can be selected. Thus, lower product stock - keeping unit (SKU) costs are incurred since different die designs for different segments are not required.
[0030] Referring now to FIG. 1, a block diagram of a package according to one embodiment is shown. As shown in FIG. 1, package 100 can be any type of integrated circuit package. In the particular figure shown, package 100 includes a central processing unit (CPU) die 110 0-n , an accelerator die 120, an input / output (I / O) tile 130, and a memory 140 1-4 comprising a plurality of chiplets or dies. At least certain of these dies can be coupled together via on - package interconnects according to one embodiment. As shown, interconnect 150 1-3 can be implemented as a UCIe interconnect. CPU 110 can be coupled via another on - package interconnect 155 that may optionally provide a CPU - CPU connection on - package using a UCIe interconnect that executes a coherence protocol. As one such example, this coherence protocol can be Intel® Ultra Path Interconnect (UPI); of course, other examples are possible.
[0031] The protocols mapped to the UCIe protocol described herein include PCIe and CXL, but it should be understood that the embodiments are not limited in this regard. In an exemplary embodiment, the mapping of any underlying protocol can be performed using a flit format that includes a low mode. In one implementation, these protocol mappings enable more on-package integration by replacing certain physical layer circuits (e.g., PCIe SERDES PHY and PCIe / CXL LogPHY with link-level retry) with a UCIe die-to-die adapter and PHY according to an embodiment, which can improve power and performance characteristics. Additionally, the low mode may be protocol-independent to enable usage such as the integration of a stand-alone SERDES / transceiver tile (e.g., Ethernet (R)) on-package while allowing other protocols to be mapped. As further shown in FIG. 1, the off-package interconnect can conform to various protocols including CXL / PCIe protocols and double data rate (DDR) memory interconnect protocols, etc.
[0032] In an exemplary implementation, the accelerator 120 and / or the I / O tile 130 can be connected to the CPU 110 using CXL transactions operating on the UCIe interconnect 150 by leveraging the I / O, coherence, and memory protocols of CXL. In the embodiment of FIG. 1, the I / O tile 130 can provide an interface to CXL, PCIe, and DDR pins external to the package. Static or dynamic, the accelerator 120 can also be connected to the CPU 110 using PCIe transactions operating on the UCIe interconnect 150.
[0033] A package according to one embodiment can be implemented in many different types of computing devices, from small portable devices such as smartphones to larger devices including client computing devices and servers or other data center computing devices. Thus, UCIe interconnects can enable local and long-distance connections at the rack / pod level. Although not shown in FIG. 1, it should be understood that at least one UCIe retimer may be used to extend the UCIe connection beyond this package using off-package interconnects. Examples of off-package interconnects include electrical cables, optical cables, or any other technology for connecting packages at the rack / pod level.
[0034] Embodiments can also be used to support rack / pod level disaggregation using the CXL 2.0 (or later) protocol. In such a configuration, multiple compute nodes (e.g., virtual tiers) from different compute chassis can be coupled to a CXL switch that can be coupled to multiple CXL accelerators / Type-3 memory devices that can be placed within one or more separate drawers. Each compute drawer can be coupled to this switch using an off-package interconnect that executes the CXL protocol through a UCIe retimer.
[0035] Referring now to FIGS. 2A through 2D, cross-sectional views of different packaging options incorporating embodiments are shown. As shown in FIG. 2A, package 200 may be an advanced package that provides advanced packaging technology. In one or more embodiments, advanced package implementations can be used for performance optimization applications, including power saving performance applications. In some such exemplary use cases, the channel reach can be short (e.g., less than 2 mm), and the interconnect can be optimized for high bandwidth and low latency with the best performance and power efficiency characteristics.
[0036] As shown in FIG. 2A, package 200 includes a plurality of dies 210 0-2 Although three specific dies are shown in FIG. 2A, it should be understood that more dies may be present in other implementations. Dies 210 are adapted onto package substrate 220. In one or more embodiments, dies 210 may be adapted to substrate 220 via bumps. As shown, package substrate 220 includes a plurality of silicon bridges 225 1-2 including on-package interconnects 226 1-2 In one embodiment, interconnect 226 may be implemented as a UCIe interconnect, and silicon bridge 225 may be implemented as an Intel® EMIB bridge.
[0037] Referring now to FIG. 2B, another embodiment of an advanced package in which the package configuration is implemented as chip-on-wafer-on-substrate (CoWoS) is shown. In this figure, package 201 includes dies 210 adapted onto an interposer 230, which includes corresponding on-package interconnects 236. As a result, interposer 230 is adapted to package substrate 220 via bumps.
[0038] Referring now to FIG. 2C, another embodiment of an advanced package in which the package configuration is implemented using a fan-out organic interposer 230 is shown. In this figure, package 202 includes dies 210 adapted onto an interposer 230 that includes corresponding on-package interconnects 236. As a result, interposer 230 is adapted to package substrate 220 via bumps.
[0039] Referring now to FIG. 2D, another package diagram is shown. Package 203 may be a standard package that provides standard packaging technology. In one or more embodiments, a standard package implementation may be used for low-cost and long-distance (e.g., from 10 mm to 25 mm) interconnections using traces on an organic package / substrate while still providing significantly better BER characteristics compared to off-package SERDES. In this implementation, package 203 includes die 210 adapted to a package substrate 220 in which on-package interconnect 226 is directly adapted within the package substrate 220 without including a silicon bridge or the like.
[0040] Referring now to FIGS. 3A / 3B, a block diagram of a layered protocol in which one or more embodiments may be implemented is shown. As shown at a high level in FIG. 3A, a plurality of layers of the layered protocol implemented within circuit 300 may implement an interconnect protocol. Protocol layer 310 may communicate information of one or more application-specific protocols. In one or more implementations, protocol layer 310 may operate according to one or more of PCIe or CXL flit mode and / or streaming protocols to provide a general-purpose mode for user-defined protocols to be transmitted. For each protocol, different organizations and associated flit transfers are available.
[0041] As a result, protocol layer 310 couples to die-to-die adapter (D2D) adapter 320 via interface 315. In one embodiment, interface 315 may be implemented as a flit-aware D2D interface (FDI). In one embodiment, D2D adapter 320 may be configured to ensure the success of data transfer across UCIe link 340 by way of protocol layer 310 and physical layer 330 cooperating with and may be configured to provide a low-latency and optimized data path for protocol flits by minimizing the logic on the main data path as much as possible.
[0042] Figure 3A shows various functions executed within the D2D adapter 320. The D2D adapter 320 may provide link state management and parameter negotiation for the connected die (also referred to as a "chiplet"). Still further, for example, if the low bit error rate (BER) is less than 1e-27, the D2D adapter 320 may optionally guarantee reliable delivery of data through cyclic redundancy check (CRC) and link level retry mechanisms. If multiple protocols are supported, the D2D adapter 320 may define the underlying arbitration mechanism. For example, when transporting CXL protocol communication, the adapter 320 may provide an arbiter / multiplexer (ARB / MUX) function that supports communication of multiple simultaneous protocols. In one or more embodiments, if the D2D adapter 320 is responsible for reliable transfer, a flow control unit (flit) of a given size, such as 256 bytes, may define the underlying transfer mechanism.
[0043] If the operation is in flit mode, the die-to-die adapter 320 may insert and check CRC information. In contrast, if the operation is in low mode, all information of the flit (e.g., bytes) is populated by the protocol layer 310. If applicable, the adapter 320 may also perform retries. The adapter 320 may further be configured to coordinate higher level link state machine management and bring up protocol option related parameter exchange with a remote link partner and, if supported, power management coordination with the remote link partner. Different underlying protocols may be used depending on the usage model. For example, in one embodiment, data transfer using direct memory access, software discovery, and / or error handling, etc. may be processed using PCIe / CXL.io; memory use cases may be processed through CXL.Mem; and cache requirements for applications such as accelerators may be processed using CXL.cache.
[0044] As a result, the D2D adapter 320 couples to the physical layer 330 via the interface 325. In one embodiment, the interface 325 may be a Low D2D Interface (RDI). As shown in Figure 3B, the physical layer 330 includes circuitry for interfacing with the die-to-die interconnect 340 (which may be a UCIe interconnect or another multi-protocol compliant on-package interconnect in one embodiment). In one or more embodiments, the physical layer 330 may be responsible for electrical signaling, clocking, link training, sideband, etc.
[0045] The interconnect 340 may include sideband and mainband links, which may be in the form of so-called "lanes", which are physical circuits for carrying signaling. In one embodiment, a lane may comprise circuitry for carrying a pair of signals (one for transmission and one for reception) mapped to physical bumps or other conductive elements. In one embodiment, xN The UCIe link is composed of N lanes.
[0046] As shown in Figure 3B, the physical layer 330 includes three sub-components, namely, physical (PHY) logic 332, electrical analog front end (AFE) 334, and sideband circuitry 336. In one embodiment, the interconnect 340 includes a mainband interface that provides a main data path on a plurality of physical bumps that may be organized as a group of multiple lanes called a module or cluster.
[0047] The unit of construction of the interconnect 340 is equally referred to herein as a "cluster" or a "module". In one embodiment, a cluster may include N single-ended unidirectional full-duplex data lanes, one single-ended lane for Valid, one lane for tracking, a differential transfer clock for each direction, and two lanes for each direction for sideband (single-ended clock and data). Thus, a module (or cluster) forms the atomic granularity of the structural design implementation of the AFE 334. There may be a different number of lanes provided for each module of standard and advanced packages. For example, for a standard package, 16 lanes constitute a single module, while for an advanced package, 64 lanes constitute a single module. Although the embodiments are not limited in this regard, the interconnect 340 is a physical interconnect that can be implemented using one or more of conductive traces, conductive pads, bumps, etc., that provide an interconnect between PHY circuits present on the link partner die.
[0048] A given instance of the protocol layer 310 or the D2D adapter 320 can transmit data through a plurality of modules in which bandwidth scaling is implemented. The physical link of the interconnect 340 between the dies may include two separate connections: (1) a sideband connection; and (2) a mainband connection. In an embodiment, the sideband connection is used for parameter exchange, register access for debugging / link training, and compliance to and coordination with a remote partner for management.
[0049] In one or more embodiments, the sideband interface is formed by at least one data lane and at least one clock lane in each direction. In other words, the sideband interface is a two-signal interface for transmission and reception directions. In the use of an advanced package, additional data and clock pairs in each direction can be provided with redundancy for repair or bandwidth increase. The sideband interface may include transfer clock pins and data pins in each direction. In one or more embodiments, the sideband clock signal can be generated by an auxiliary clock source configured to operate at 800 MHz regardless of the main data path speed. The sideband circuit 336 of the physical layer 330 may be provided with an auxiliary power supply and may be included in a domain that is always on. In one embodiment, the sideband data can be communicated as a single data rate signal (SDR) of 800 megatransfers per second (MT / s). The sideband can be configured to operate on a power supply and an auxiliary clock source that are always on. Each module has its own set of sideband pins.
[0050] The mainband interface that constitutes the main data path may include a transfer clock, a data valid pin, and N data lanes per module. In the advanced package option, N = 64 (also referred to as ×64), and all four additional pins for lane repair are provided in the bump map. In the standard package option, N = 16 (also referred to as ×16), and no additional pins for repair are provided. The physical layer 330 can be configured to coordinate different functions and their relative sequencing for proper link bring-up and management (e.g., sideband transfer, mainband training and repair, etc.).
[0051] In one or more embodiments, an advanced package implementation may support redundant lanes (also referred to herein as "spare lanes") to handle defective lanes (including clock, valid, sideband, etc.). In one or more embodiments, a standard package implementation may support lane width reduction to handle failures. In some embodiments, multiple clusters may be aggregated to achieve higher performance per link.
[0052] Referring now to FIG. 4A, a block diagram of a multi-die package according to one embodiment is shown. As shown in FIG. 4A, package 400 includes at least a first die 410 and a second die 450. It should be understood that dies 410 and 450 may be various types of dies, including CPUs, accelerators, or I / O devices, etc. In the high-level diagram shown in FIG. 4A, the interconnect 440 that couples the dies together is shown as a dashed line. Interconnect 440 may be an instantiation of an on-package multi-protocol compliant interconnect, e.g., the UCIe interconnect described herein. Although not shown in detail in FIG. 4A, it should be understood that interconnect 440 may be implemented using conductive bumps adapted to each die that may be coupled together to provide an interconnect between the dies. Additionally, interconnect 440 may further include in-package circuitry, such as conductive lines on or within one or more substrates. As used herein, the term "lane" should be understood to refer to any interconnect circuit that couples one die to another die.
[0053] In certain embodiments, the interconnect 440 may be a UCIe interconnect having one or more modules, each module including a sideband interface and a mainband interface. In this high-level diagram, the mainband interface couples to the mainband receiver and transmitter circuits within each die. Specifically, die 410 includes mainband receiver circuit 420 and mainband transmitter circuit 425, and as a result, die 450 includes mainband receiver circuit 465 and mainband transmitter circuit 460.
[0054] FIG. 4A further shows the connection of the sideband interface. Generally, the sideband includes data lanes and clock lanes in each direction, and in the use of advanced packages, redundancy may be provided for additional data and clock pairs in each direction. Thus, FIG. 4A shows a first possible connection implementation between the sideband circuits of these two dies. Die 410 includes sideband circuit 430, which includes a first sideband circuit 432 that couples to the corresponding sideband clock and data receivers (R_C and R_D) and sideband clock and data transmitters (T_C and T_D) of the sideband circuit 470 of the second die 450, respectively. The sideband circuit 430 also includes a second sideband circuit 434 having similar circuits for redundant sideband clock and data transmitters and receivers (abbreviated as transmitters and receivers above shown in Abbreviations for transmitters and receivers with "R" appended at the end of ).
[0055] In FIG. 4A, a first sideband connection instantiation is shown. The sideband circuits 432 and 472 function as the functional sideband, and the sideband circuits 434 and 474 function as the redundant sideband.
[0056] In response to sideband detection performed during sideband initialization, it may be determined that there is a defect in one or more of the sideband lanes and / or associated sideband circuits, and thus at least a portion of the redundant sideband circuit can be used as part of the functional sideband. More specifically, FIG. 4B shows a second possible connection implementation between the sideband circuits of these two dies. In this example, redundant sideband data transmitters and receivers are present within sideband circuit 472 and function as part of the functional sideband.
[0057] In different implementations, any connection can be made by the initialization and bring-up flow as long as data-to-data and clock-to-clock connections are maintained. If redundancy is not required based on such initialization, faster message exchange can be made possible by using both sideband circuit pairs to expand the sideband bandwidth. FIGS. 4A and 4B are shown in the context of an advanced package configuration, but note that similar sideband circuits may exist in the dies used within a standard package. However, in certain implementations, the standard package may not provide redundancy and lane repair support, so redundant sideband circuits and redundant sideband lanes may not be present within the standard package.
[0058] Referring now to FIG. 5, a schematic diagram showing a die-to-die connection according to one embodiment is shown. As shown in FIG. 5, package 500 includes a first die 510 and a second die 560. An interconnect 540, e.g., a UCIe interconnect, includes a plurality of sideband lanes, namely sideband lanes 541-544. Although a single direction of the sideband lanes is shown, it should be understood that a corresponding set of sideband lanes may also be provided in the other direction. The first die 510 includes sideband data transmitters and sideband clock transmitters, namely sideband data transmitters 511, 512 (sideband data transmitter 512 is a redundant transmitter). The first die 510 further includes sideband clock transmitters 514, 515 (sideband clock transmitter 515 is a redundant transmitter). The second die 560, as a result, includes sideband data receivers and sideband clock receivers, namely sideband data receivers 561, 562 (sideband data receiver 562 is a redundant receiver). The second die 560 further includes sideband clock receivers 564, 565 (sideband clock receiver 565 is a redundant receiver).
[0059] Referring still to FIG. 5, a detection circuit that may be used to perform sideband detection, which may be part of sideband initialization to determine which lanes are included in the functional sideband and which lanes may be part of the redundant sideband, is present within the second die 560. As shown, a plurality of detectors 570 0-3is provided. Each detector 570 receives incoming sideband data signals and incoming sideband clock signals such that each detector 570 receives signals from different combinations of the sideband receivers of the second die 560. During sideband initialization, the incoming sideband data signal may be a predetermined sideband initialization packet that includes a predetermined pattern. The detector 570 is configured to detect the presence of this pattern, generate a first result (e.g., logic 1) in response to a valid detection of this pattern (e.g., with respect to the number of repetitions of this pattern), and generate a second result (e.g., logic 0) in response to the predetermined pattern not being detected. Although embodiments are not limited in this regard, in one implementation, the detector 570 may be configured using a shift register, a counter, etc. to perform this detection operation and sample data and redundant data using the clock signal and redundant clock signal to produce four combinations to generate corresponding results.
[0060] Note that when the redundant sideband circuit is not used for repair purposes, the redundant sideband circuit can be used to increase the bandwidth of sideband communication, particularly for data-intensive transfers. As an example, a sideband according to one embodiment can be used to communicate large amounts of information being downloaded, such as firmware and / or downloads. Or, the sideband can be used to communicate management information, for example, according to a given management protocol. Note that such communication can be performed concurrently with other sideband information communication on the functional sideband.
[0061] Referring now to FIG. 6A, a timing diagram showing sideband signaling according to one embodiment is shown. As shown in FIG. 6A, timing diagram 600 includes a sideband clock signal 610 and a sideband message signal 620. The sideband message format can be defined as a 64-bit header that includes 32-bit or 64-bit data communicated during 64 unit intervals (UI). The sideband message signal 620 represents a 64-bit serial packet. The sideband data can be transmitted with edges aligned with the clock (strobe) signal. The receiver of the sideband interface samples the incoming data with the strobe. For example, the negative edge of the strobe can be used to sample the data when the data uses SDR signaling.
[0062] Referring now to FIG. 6B, a timing diagram showing sideband packet continuous transmission according to one embodiment is shown. As shown in FIG. 6B, timing diagram 601 shows the communication of a first sideband packet 622 followed by a second sideband packet 624. As shown, each packet can be a 64-bit serial packet transmitted during a 64 UI duration. More specifically, the first sideband packet 622 is transmitted, and as a result, logic low on both the clock lane and the data lane lasts for 32 UI duration, after which the second sideband packet 624 is communicated. In an embodiment, such signaling can be used for various sideband communications including sideband messages during sideband initialization.
[0063] Referring now to FIG. 7, there is shown a flow diagram depicting the bring-up flow of an on-package multi-protocol compliant interconnect according to one embodiment. As shown in FIG. 7, the bring-up flow 700 begins by independently executing a reset flow for two dies (die 0 and 1) coupled together via, for example, a UCIe interconnect (shown as a D2D channel in FIG. 7). Thus, the first die (die 0) executes an independent reset flow at stage 710, and the second die (die 1) also executes an independent reset flow at stage 710. Note that each die may finish the reset flow at different times. Next, at stage 720, sideband detection and training may be performed. At stage 720, the sideband may be detected and trained. In the case of an advanced package where lane redundancy is available, the available lanes are detected and may be used for sideband messages. Note that since each die may finish the reset flow at different times as described above, this sideband detection and training, including sideband initialization described herein, may be used to detect the presence of activity in the coupled dies. In one or more embodiments, the trigger to end the reset and start link training is the detection of a sideband message pattern. When the physical layer transitions from a reset state, when training during link bring-up, the hardware is permitted to attempt training multiple times. During this bring-up operation, both both dies for all status and sub - states in and out are , these are guaranteed to be in lockstep by a four-mode sideband message handshake between the dies so synchronization may occur .
[0064] At stage 730, training parameter exchange may be performed for the functional sideband, and main band training is carried out. At stage 730, the main band is initialized, repaired, and trained. Finally, at stage 740, protocol parameter exchange may be performed for the sideband. At stage 740, the entire link may be initialized by determining the local die function, parameter exchange with the remote die, and bringing up the FDI that couples the corresponding protocol layer to the die's D2D adapter. In one embodiment, the main band is initialized, by default, at the lowest acceptable data rate in main band initialization where repair and inversion detection are performed. Next, the link speed transitions to the highest common data rate detected through parameter exchange. After link initialization, the physical layer may be enabled to perform protocol flit transfer via the main band.
[0065] In one or more embodiments, different types of packets may be communicated via the sideband interface and may include (1) configuration (CFG), or register access that may be memory mapped read or write and may be 32 bits or 64 bits (b); (2) link management (LM) or vendor-defined packets, which are messages that do not contain data and do not carry an additional data payload; (3) parameter exchange (PE), link training related or vendor-defined messages that may carry 64b of data. The packet may carry a 5-bit opcode, a 3-bit source identifier (srcid), and a 3-bit destination identifier (dstid). The 5-bit opcode indicates the packet type and whether the packet carries 32b or 64b of data.
[0066] Flow control and data integrity sideband packets can be transferred over the FDI, RDI, or UCIe sideband links. Each of these has independent flow control. For each transmitter associated with FDI or RDI, interface design-time parameters can be used to determine the number of credits (up to 32 credits) notified by the receiver. Each credit corresponds to a 64-bit header and 64 bits of potentially associated data. Thus, there is only one type of credit for all sideband packets regardless of how much data they carry. All transmitter / receiver pairs have independent credit loops. For example, in RDI, credits are notified from the physical layer to the adapter for sideband packets transmitted from the adapter to the physical layer; also, credits are notified from the adapter to the physical layer for sideband packets transmitted from the physical layer to the adapter. The transmitter checks for available credits before sending register access requests and messages. The transmitter does not check for credits before sending a register access completion, and the receiver guarantees unconditional sinking for any register access completion packet. Messages carrying requests or responses consume credits in FDI and RDI, but it is guaranteed that the receiver makes forward progress and is not blocked behind a register access request. Both RDI and FDI provide dedicated signals for sideband credit return across their interfaces. All receivers associated with RDI and FDI check received messages for data or control parity errors, and these errors are mapped to uncorrectable internal errors (UIE), transitioning RDI to the LinkError state.
[0067] Referring to FIG. 8, a flow diagram of a link training state machine according to an embodiment is shown. As shown in FIG. 8, method 800 is an example of link initialization that may be performed by a logical physical layer circuit that includes, for example, a link state machine. Table 1 is a high-level description of the states of a link training state machine according to an embodiment, and the details and actions performed in each state are described below. [Table 1]
Table 1
[0068] Referring to FIG. 8, method 800 begins in a reset state 810. In one embodiment, the PHY remains in the reset state for a predetermined minimum duration (e.g., 4 ms) to allow various circuits including a phase-locked loop (PLL) to stabilize. This state can end when the power supply is stable, the sideband clock is available and operating, the mainband and die-to-die adapter clocks are stable and available, the mainband clock is set to the slowest IO data rate (e.g., 4 GT / s at 2 GHz), and a link training trigger is performed. Next, control proceeds to a sideband initialization (SBINIT) state 820 where sideband initialization may be performed. In this state, the sideband interface is initialized and repaired (if applicable). During this state, the mainband transmitter may be in a tri-state, and the mainband receiver is permitted to be disabled.
[0069] Referring further to FIG. 8, from the sideband initialization state 820, control proceeds to the mainband initialization (MBINIT) state 830 where mainband initialization is performed. In this state, the mainband interface is initialized and repaired, or degraded (if applicable). The data rate of the mainband can be set to the lowest supported data rate (e.g., 4 GT / s). For advanced packages, interface interconnect repair can be performed. In the sub-state in MBINIT, detection and repair of data, clock, tracking, and valid lanes are enabled. For standard package interfaces that do not require lane repair, the sub-state is used to check functionality at the lowest data rate and perform width reduction if required.
[0070] Next, in block 840, it enters the mainband training (MBTRAIN) state 840 where mainband link training can be performed. In this state, the operating speed is set and centering from the clock to the data is performed. At higher speeds, additional calibrations such as receiver clock correction, transmit and receive deskew are performed in the sub-state to ensure link performance. The module enters each sub-state and each state ends through a sideband handshake. If no specific action in the sub-state is required, the UCIe module is allowed to end the state through a sideband handshake without performing the operation of that sub-state. In one or more embodiments, this state may be common to advanced and standard package interfaces.
[0071] Next, the control proceeds to block 850, where a LINKINIT state occurs and link initialization can be performed. In this state, the die-to-die adapter completes the initial link management before entering the active state on the RDI. Once the RDI becomes active, the PHY clears a copy of the bit "start UCIe link training" from the link control register. In an embodiment, the linear feedback shift register (LFSR) is reset when entering this state. In one or more embodiments, this state may be common to advanced and standard package interfaces.
[0072] Finally, the control proceeds to the active state 860, where communication can occur in normal operation. More specifically, packets from the upper layer can be exchanged between these two dies. In one or more embodiments, all data in this state can be scrambled using the scrambler LFSR.
[0073] Referring further to FIG. 8, it should be noted that during the active state 860, a transition to a re-training (PHYRETRAIN) state 870 may occur, or a low-power (L2 / L1) link state 880 may occur. As can be seen, depending on the level of the low-power link state, it can proceed to either the mainband training state 840 or the reset state 810 from the end. In the low-power link state, less power is consumed than in the dynamic clock gating of the die in the active state. This state can be entered when the RDI transitions to a power management state. When the local adapter requests the activation of the RDI or the remote link partner requests the end of L1, the PHY exits to the MBTRAIN.SPEEDIDLE state. In one or more embodiments, the end of L1 is coordinated with the corresponding L1 state end transition on the RDI. When the local adapter requests the active state of the RDI or the remote link partner requests the end of L2, the PHY exits to the reset state. Note that the end of L2 can be coordinated with the corresponding L2 state end transition on the RDI.
[0074] As further shown in FIG. 8, if an error occurs between any of the multiple bring-up states, the control proceeds to block 890, and a training error state may occur. This state is used as a transient state resulting from some fatal or non-fatal event to reset the state machine to the reset state. When the sideband is active, a sideband handshake is performed for the link partner to enter the TRAINERROR state from any state other than SBINIT.
[0075] In one embodiment, the die may enter the PHYRETRAIN state for multiple reasons. The trigger may be due to PHY re-training targeted at the adapter or PHY-initiated PHY re-training. The local PHY starts re-training when it detects a Valid framing error. The remote die may request PHY re-training, and as a result, the local PHY enters PHY re-training when it receives this request. It may also enter this re-training state when a change is detected in the runtime link test control register during the MBTRAIN.LINKSPEED state. Although shown at this high level in the embodiment of FIG. 8, it should be understood that many variations and alternatives are possible.
[0076] Referring now to FIG. 9, a more detailed flowchart of mainband initialization according to one embodiment is shown. Method 900 may be implemented by a link state machine to perform mainband initialization. As shown, this initialization passes through multiple states including a parameter exchange state 910, a calibration state 920, a repair clock state 930, a repair verification state 940, an inverted mainband state 950, and finally a mainband repair state 960. After the completion of this mainband initialization, the control proceeds to mainband training.
[0077] In the parameter exchange state 910, parameter exchange can be performed to set up the maximum negotiation speed and other PHY settings. In one embodiment, the following parameters, namely, voltage amplitude; maximum data rate; clock mode (e.g., strobe or continuous clock); clock phase; and module ID can be exchanged with the link partner (e.g., for each module). In state 920, any required calibration (e.g., transmit duty cycle correction, receiver offset, and Vref calibration) can be performed.
[0078] Next, in block 930, detection and repair (if required) for the functional check of the clock and trace lanes of the advanced package interface and the clock and trace lanes of the standard package interface can be performed. In block 940, the module can set the clock phase to the center of the data UI of its main band transmitter. The module partner samples the received Valid using the received transfer clock. All data lanes can be held low during this state. This state can be used to detect and apply repairs (if required) to the Valid lanes.
[0079] Referring further to FIG. 9, entry into block 950 occurs only if the clock and valid lanes are functioning. In this state, data lane inversion is detected. All transmitters and receivers of the module are enabled. The module sets the transfer clock phase to the center of the data UI of its main band. The module partner samples the incoming data using the incoming transfer clock. The 16-bit "lane-by-lane ID" pattern (not scrambled) is a lane-specific pattern using the lane ID of the corresponding lane.
[0080] Referring further to FIG. 9, in block 960 which is entered only after lane reversal detection and application have been successful, all transmitters and receivers of the module are enabled. The module sets the clock phase to the center of the data UI of its main band. The module partner samples the incoming data using the incoming transfer clock of the main band receiver. In this state, the main band lanes are detected and repaired if required for the advanced package interface and for the function check and width reduction of the standard package interface. Put another way, if an error is detected within a lane, the redundant circuitry can be enabled via the redundant lane.
[0081] In an exemplary embodiment, several fallback techniques may be used to find the operational settings with the link enabled during bring-up and operation. First, if an error is detected (during initial bring-up or function operation) but no repair is required, a speed reduction may be performed. Such a speed reduction mechanism may then shift the link to the next lower acceptable frequency; this is repeated until a stable link is established. Second, if repair is not possible (for a standard package link without repair resources), a width reduction may be performed, and as an example, a reduction in width to a half-value width configuration may be allowed. For example, a 16-lane interface may be configured to operate as an 8-lane interface.
[0082] Referring now to FIG. 10, a flow diagram of main band training according to one embodiment is shown. As shown in FIG. 10, method 1000 may be implemented by a link state machine to perform main band training. In main band training, the main band data rate is set to the highest common data rate of the two connected devices. Data-to-clock training, deskew, and Vref training may be performed using a plurality of sub-states. As shown in FIG. 10, main band training goes through a plurality of states or sub-states. As shown, main band training begins by performing valid reference voltage training state 1005. In state 1005, the receiver reference voltage (Vref) for sampling incoming Valid is optimized. The data rate of the main band continues to be the lowest supported data rate. The module partner sets the transfer clock phase to the center of the data UI of its main band transmitter. The receiver module samples the pattern of the Valid signal using the transfer clock. All data lanes are held low during valid lane reference voltage training. Next, control proceeds to data reference voltage state 1010, where the receiver reference voltage (Vref) for sampling incoming data is optimized, while the data rate continues to be the lowest supported data rate (e.g., 4 GT / s). The transmitter sets the transfer clock phase to the center of the data UI. Thereafter, an idle speed state 1015 occurs, and frequency changes may be allowed in this electrical idle state; more specifically, the data rate may be set to the maximum common data rate determined in the previous state. Thereafter, the circuit parameters may be updated in transmitter and receiver calibration states (1020 and 1025).
[0083] Referring further to FIG. 10, various training states 1030, 1035, 1040, and 1045 proceed to train a valid-to-clock training reference voltage level, a full data-to-clock training, and a data receiver reference voltage, respectively. In state 1030, valid-to-clock training is performed prior to data lane training to ensure that the valid signal is functional. The receiver samples the valid pattern using the transfer clock. In state 1035, the module can optimize the reference voltage (Vref) to sample the incoming valid at this operating data rate. In state 1040, the module performs full data-to-clock training (including valid) using an LFSR pattern. In state 1045, the module can optimize the reference voltage (Vref) of its data receiver to optimize the sampling of the incoming data at this operating data rate.
[0084] Referring further to FIG. 10, a receiver deskew state 1050 may then occur, which is a training stage initiated by the receiver to perform lane-to-lane deskewing to improve the timing margin. Next, another data training state 1055 occurs, in which the receiver of the module partner performs per-lane deskewing, the module can re-center the clock and aggregate the data. Next, the control proceeds to the link speed state 1060, and after the final sampling point is set in state 1055, the link stability at this operating data rate can be checked. If the link performance does not meet this data rate, the speed is then reduced to the next lower supported data rate and training is executed again. Depending on the result of such a state, the main band training may be completed, and then the control proceeds to link initialization. Otherwise, a link speed change can be made either in state 1015 or the repair state 1065. Note that entering states 1015 and 1065 can occur from a low power state (e.g., L1 link power state) or a re-training state. Although shown at this high level in the embodiment of FIG. 10, it should be understood that many variations and alternatives are possible.
[0085] In different implementations, different numbers of redundant lanes can be provided. In one example, if some functional lanes are damaged during, for example, package chiplet assembly, approximately 3 to 5% redundant lanes can be added in the PHY layer to recover the die. Of course, in a given implementation, additional or fewer redundant lanes may be present.
[0086] Referring now to FIG. 11A, a diagram of data lane remapping possibility according to an embodiment is shown. More specifically, in FIG. 11A, a portion of die 1100 is shown with a plurality of data lanes 11100-1110 31is shown together. Data lane 1110 may be a physical transmission data lane to which a corresponding logical transmission data lane is mapped upward. In an actual application, data lane 1110 may terminate at a series of bumps or other conductors on the external surface of the die to enable direct (or direct to another die) adaptation to a package substrate or an interposer.
[0087] As further shown in FIG. 11A, the arrow lines represent cross-tangent lines between different data lanes to enable the repair and inversion operations described herein.
[0088] FIG. 11A is in the context of the logical remapping possibility by the corresponding redundant data lane 1120 0、1 This remapping may enable defective data lanes to be remapped to adjacent lanes. Such remapping may be performed sequentially so that ultimately the data of a given data lane is provided to the corresponding redundant data lane. Thus, as shown in FIG. 11A, the data traffic of data lane 11100 may be remapped to redundant data lane 11200, and similarly, the data traffic of data lane 1110 31 may be remapped to redundant data lane 11201.
[0089] In one particular embodiment, the module can support remapping (repair) of up to two data lanes per group of 32 data lanes (e.g., two redundant data lanes of a first set of physical data lanes (e.g., TD_P[31:0] (RD_P[31:0]) which are transmit and receive physical data lanes) and two redundant data lanes of a second set of physical data lanes (TD_P[63:32] (RD_P[63:32]))). Thus, two separate groups of 32 lanes can each be repaired independently using redundant data lanes (TRD_P[1:0] (RRD_P[1:0]) and TRD_P[3:2] (RRD_P[3:2])). In FIG. 11A, there are two redundant resources per 32 data lanes for functional data recovery, but in other embodiments, additional redundant resources may exist. In different embodiments, there may be between approximately one and six redundant data lanes per set of 32 data lanes. Also, it should be understood that the size of the set of data lanes can vary in different implementations, for example, in some cases between four and 128.
[0090] In one or more embodiments, lane remapping can be achieved by a "left shift" or a "right shift". operation The left shift operation is performed when the data traffic of the logical lane TD_L[n] associated with the physical data lane TD_P[n] is multiplexed onto a different physical data lane TD_P[n - 1]. The right shift operation This is performed when the data traffic of the logical data lane TD_L[n] is multiplexed onto the physical data lane TD_P[n+1]. After the data lanes are remapped, the physical layer may control and invalidate (e.g., tri-state) the transmitter associated with the damaged physical lane and control and invalidate the corresponding receiver. As a result, the transmitters and receivers of the redundant lanes used for repair are enabled. To optimally repair any two lanes at most within the group, both "left shift" and "right shift" remappings may be performed. Of course, additional lanes may be repaired if additional redundant resources are available.
[0091] In addition to the redundant data lane resources, dedicated redundant clock lanes for differential clock circuits may exist. In one embodiment, clock lane remapping enables the repair of single lane faults for both differential and pseudo-differential implementations of the clock circuit. Similar redundant circuits may also be provided for the trace lanes.
[0092] Referring now to FIG. 11B, a representative remapping configuration for a single lane fault is shown. The numbering scheme in FIG. 11B follows that of FIG. 11A, but a different die 1101 is shown. As can be seen, data lane 1110 29An error in (e.g., this physical lane is damaged during assembly) causes remapping in the direction of redundant data lane 11200 that functions as a repair resource. According to this method, in the embodiment of FIG. 11B, redundant data lane 11200 (TRD_P[0](RRD_P[0])) is used as a redundant lane for remapping any single physical lane failure of TD_P[31:0](RD_P[31:0]), and TRD[2](RRD[2]) is used as a redundant lane for remapping any single lane failure of TD_P[63:32](RD_P[63:32]). Of course, this is equally possible for redundant data lane 11201 that becomes a repair resource, and the remapping proceeds in the opposite direction. Therefore, the bidirectional shift mechanism enables forward and backward shifting of any lane data to adjacent lanes, and as a result, the device can support sufficient connection or bandwidth.
[0093] Referring here to Table A and Table B, a pseudo-code representation of repair in the lower and upper lanes according to an embodiment is shown. [Table A] Pseudo-code for lane repair in TD_P[31:0](RD_P[31:0])(0<=x<=31):
Table 2
Table 3
[0094] Referring here to FIG. 11C, a representative remapping configuration of two lane failures within a module is shown. In this configuration, any two lanes within a group of 32 lanes are repaired using redundant resources, and full functional traffic can be restored. For example, as shown in FIG. 11C, physical data lane 1110 25and 1110 26 is damaged. For this lane repair, data is shifted from 1110 25 to 1110 24 and all lanes below it are shifted towards redundant lane 11200 (TRD_P0). For the repair of TD26, data is diverted to TD27 and all lanes above TD27 are shifted towards redundant lane 11201 (TRD_P1). By using the above bidirectional shift scheme, any combination of single or dual lane damage can be fully recovered.
[0095] Thus, in one embodiment, for any two physical lane failures in TD_P[31:0] (RD_P[31:0]), the lower lane is remapped to TRD_P[0] (RRD_P[0]) and the upper lane is remapped to TRD_P[1] (RRD_P[1]). For any two physical lane failures in TD_P[63:32] (RD_P[63:31]), the lower lane is remapped to TRD_P[2] (RRD_P[2]) and the upper lane is remapped to TRD_P[3] (RRD_P[3]).
[0096] Referring now to Tables C and D, a pseudo-code representation of two-lane repair in the lower and upper lanes according to one embodiment is shown. Note that for all of the above examples, both the transmitter and the corresponding receiver apply the indicated remapping. [Table C] Pseudo-code for two-lane repair in TD_P[31:0] (RD_P[31:0]) (0 <= x, y <= 31): [Table 4] [Table D] Pseudo-code for two-lane repair in TD_P[63:32] (RD_P[63:32]) (32 <= x, y <= 63): [Table 5]
[0097] Referring to FIG. 12A here, a schematic diagram of a part of a die-to-die connection according to an embodiment is shown. As shown in FIG. 12A, the semiconductor package 1200 includes a first die and a second die. The first die includes a plurality of transmission logic data lanes TD_L[n - 1, n + 2]. Of course, each module or cluster may include more than these shown data lanes.
[0098] As can be seen, each data lane provides data to the corresponding lane repair multiplexer 1210 n-1、n+2 Furthermore, as shown, data from adjacent logic data lanes in the left and right directions is also provided via shift lines coupled to the multiplexer 1210. When no repair is needed, the multiplexer 1210 is controlled to provide the data of the corresponding logic data lane to the corresponding one of the plurality of transmitters 1220 n-1、n+2 associated with the corresponding physical data lane. Instead, when repair is needed, the corresponding left or right shift operation is performed, so that data from adjacent logic data lanes is provided, and the multiplexer 1210 is controlled accordingly. Thus, the multiplexer 1210 may be configured to select the corresponding (true) bit lane data {n} or the previous bit lane data {n - 1} or the next bit lane data {n - 2}.
[0099] Referring further to FIG. 12A, the data output by the transmitter 1220 is coupled to the second die through the corresponding bump 1230 n-1、n+2 and through the interconnect 1240 n-1、n+ 2. Through the corresponding bump 1250 n-1、n+2 the data is received at the receiver 1260 n-1、n+2 and from there provided to the corresponding lane repair multiplexer 1270 n-1、n+2 As shown, the corresponding left and right shifts operation By providing this, it becomes possible to perform the remapping operation also on the receiving side. More specifically, the opposite remapping performed on the first die transmitter side can be performed on the second die receiver side (for example, a left shift operation is performed in the first die, the corresponding right shift operation is performed in the second die). Although this is shown at a high level in FIG. 12A, many variations and alternatives are possible.
[0100] Referring here to FIG. 12B, a left shift operation is performed on the transmitting side to repair the defective data lane, and the corresponding right shift operation is performed on the receiving side for the single-lane remapping in package 1201. Along with this, the configuration of FIG. 12A is shown. In this case, note that only the logical shift operation is shown; it should be understood that the physical circuit (shown in FIG. 12A) still exists but is not shown here for the sake of showing the shift operation . operation
[0101] operation Referring here to FIG. 12C, a two-lane remapping situation is shown. As can be seen, on the transmitting side of package 1202, both a left and a right shift operation are performed to repair two defective data lanes, and similarly, on the receiving side, the corresponding opposite right and left shift operations are performed. As described above, FIG. 12C shows the logical shift .
[0102] For standard packages (e.g., x16) modules that do not support lane repair, resilience against defective lanes can be provided by configuring the link to a smaller (e.g., x8) width (e.g., logical lanes 0 to 7 or logical lanes 8 to 15 excluding the defective lane). For example, if one or more defective lanes are in logical lanes 0 to 7, the link is configured to x8 width using logical lanes 8 to 15. This configuration is done during link initialization or re-training, and the transmitters of the disabled lanes may be placed in a high impedance state (hi-Z), and the receivers are disabled.
[0103] Also, the device may be configured to support lane inversion within the module. An example of lane inversion is when the physical data lane 0 on the local die is connected to the physical data lane (N - 1) on the remote die (physical data lane 1 is connected to physical data lane N - 2, etc.), for example, N = 16 for the standard package and N = 64 for the advanced package. The redundant lanes in the case of the advanced package can also be inverted. In one or more embodiments, lane inversion is implemented only for the transmitter. The transmitter inverts the logical lane order on the data and redundant data lanes. In one embodiment, lane inversion is discovered and applied during initialization and training. To enable discovery of lane inversion, a unique lane ID is assigned to each logical data and redundant lane within the module. In some embodiments, the tracking, enable, clock, and sideband signals are not inverted.
[0104] Except for the multiplexer selecting between the LSB and MSB bits, the lane reversal according to one embodiment may use a multiplexer structure similar to that described above with respect to FIGS. 12A through 12C. Also, the lane shift for lane reversal is performed only on the transmitting or receiving side, not on both sides. For example, when a lane is reversed, physical lane 0 of the transmitter (TD_P[0]) is connected to physical lane N - 1 of the receiver (RD_P[N - 1]) (N is 64 in the case of an advanced package module and 16 in the case of a standard package module). Note that the lane repair mapping may be changed when lane reversal is implemented.
[0105] When repairing a single lane with lane reversal, the remapping on the transmitter side is reversed to maintain the shift order of the remapping on the receiver side. Referring here to Table E, pseudo - code is shown for repairing one lane failure with inversion at TD_P[31:0](RD_P[32:63]) (0 <= x <= 31). [Table E]
Table 6
[0106] Referring here to Table F, pseudo - code is shown for one lane failure with inversion at TD_P[63:32](RD_P[0:31]) (32 <= x <= 63). [Table F]
Table 7
[0107] In the case of two - lane repair with lane reversal, the remapping on the transmitter side is reversed to maintain the shift order of the remapping on the receiver side. Referring here to Table G, pseudo - code is shown for two - lane failure with inversion at TD_P[31:0](RD_P[32:63]) (0 <= x <= 31). [Table G]
Table 8
[0108] Referring to Table H here, the pseudo code for the 1-lane failure with inversion in TD_P[63:32](RD_P[0:31]) (32 <= x <= 63) is shown. [Table H]
Table 9
[0109] The main band repair process may, in some cases, be executed during main band initialization. This process can be executed only in the repair state that enters after the lane inversion detection and application are successful. In this state, all transmitters and receivers of the module are enabled. The module sets the clock phase to the center of the data UI of the main band. The link partner samples the incoming data using the incoming transfer clock of the main band receiver. In this state, the main band lanes are detected and repaired if required for the advanced package interface and for the function check and width reduction of the standard package interface.
[0110] In one embodiment, the following sequence may be used for the main band repair of the advanced package interface.
[0111] 1. The module sends the sideband message {MBINIT.REPAIRMB start req} and waits for a response. The link partner responds with {MBINIT.REPAIRMB start resp}.
[0112] 2. The module performs data-to-clock point training initiated by the transmitter on its transmitter lane (with a transmission pattern having 128 repetitions in a continuous mode of an "ID per lane" pattern that is not scrambled). The receiver performs a per-lane comparison, and detection on the receiver lane is considered successful if at least a predetermined number (e.g., 16 consecutive repetitions) of "ID per lane" patterns are detected.
[0113] 3. At the end of the data-to-clock point test initiated by the transmitter, the module receives per-lane pass / fail information via a sideband message.
[0114] 4. If lane repair is required and repair resources are available, the module applies repair to its mainband transmitter and sends a {MBINIT.REPAIRMB Apply repair req} sideband message. Upon receiving this sideband message, the link partner applies repair to its mainband receiver and sends a {MBINIT.REPAIRMB Apply repair resp} sideband message. If the number of lane failures is greater than the repair capabilities, the mainband is non-repairable, and the module enters the TRAINERROR state after performing a TRAINERROR handshake.
[0115] 5. If no repair is required, step 7 is executed.
[0116] 6. If lane repair is applied (step 4), the applied repair is checked by the module repeating steps 2 and 3. If post-repair lane errors are logged in step 5, the module enters TRAINERROR after performing a TRAINERROR handshake. If the repair is successful, step 7 is executed.
[0117] 7. The module sends a {MBINIT.REPAIRMB end req} sideband message, and the link partner responds with {MBINIT.REPAIRMB end resp}. When the module sends and receives {MBINIT.REPAIRMB end resp}, it exits to MBTRAIN.
[0118] Although this particular implementation is used in this embodiment, variations may be made in other embodiments. For example, similar processing may be used to perform repair after retraining or a link speed sub-state.
[0119] In the case of a standard package interface, for function operations at the lowest data rate, the main band is checked. Roughly the same steps as described above may be performed. However, if an error is identified in a data lane, it is determined whether width reduction is possible. If so, the module with the defective transmitter lane applies the reduction (to both its transmitter and receiver) and sends a message {MBINIT.REPAIRMB apply degrade req} including the logical lane map to the remote link partner. The link partner applies the reduction (to both its transmitter and receiver) and sends a message {MBINIT.REPAIRMB apply degrade resp}.
[0120] In one embodiment, for a standard package interface, if the number of lanes facing an error is all included within lanes 0 - 7 or lanes 8 - 15, the width is reduced to an ×8 link (lanes 0…lanes 7 or lanes 8…lanes 15).
[0121] Referring now to FIG. 13, a flow diagram of a lane repair method according to one embodiment is shown. As shown in FIG. 13, method 1300 can be executed during link initialization. Further, although method 1300 is in the context of main band data lane repair, the concepts described herein are equally applicable to other lane repair situations including clock lanes, sideband lanes, and tracking lanes. Method 1300 can be executed at least in part via a physical layer circuit.
[0122] As shown in FIG. 13, method 1300 begins with the transmission of a predetermined pattern of data-to-clock point training (block 1310). This predetermined pattern, which may be a per-lane pattern such as a per-lane ID pattern, can be continuously transmitted via a transmitter associated with each physical data lane. For example, 128 repetitions of this pattern can be transmitted. Next, in block 1320, result information can be received from a second die via the sideband. This result information can include per-lane pass / fail information.
[0123] Referring further to FIG. 13, next, in diamond 1330, it is determined whether at least one data lane failure has been identified. If not, control proceeds to block 1340, and the repair detection state, which is a sub-state of the initialization state of the link training state machine, can be terminated by transmitting and receiving a sideband message indicating the end of the repair state. Thereafter, control proceeds to block 1350, and this state ends and the training state is entered.
[0124] Referring further to FIG. 13, alternatively, if a failure is identified, control proceeds to diamond 1350, and it is determined whether this failure is identified in the post-repair situation. If so, an error is raised via a sideband message, and control proceeds to block 1390, and this repair sub-state ends and the training error state is entered.
[0125] If the failure is an initial failure, control proceeds to diamond 1370, and it is determined whether repair resources (e.g., including sufficient redundant lanes to accommodate the number of failed data lanes) are available. If so, control proceeds to block 1380 for application of lane repair. More specifically, at block 1380, in the transmit direction, the physical layer circuitry may apply repair to one or more transmitters associated with the data lane to remap the data traffic of at least one logical data lane onto at least one other physical lane. It should be understood that similar lane repair can be performed on the second die by appropriate remapping at the receiver by communication of information regarding the defective lane so that accurate data traffic is provided to the intended logical data lane. Although shown at this high level in the embodiment of FIG. 13, many variations and alternatives are possible.
[0126] In various embodiments, note that one or more of the features described herein may be configured to be enabled or disabled based on information stored in one or more configuration registers (which may be present, for example, in one or more of a D2D adapter or the physical layer) under, for example, dynamic user control. In addition to dynamic (or boot-time) enabling or disabling of various features, it is also possible to provide configurability regarding the operational parameters of certain aspects of UCIe communication.
[0127] Embodiments may support two broad usage models. The first is package-level integration for power-saving and cost-effective performance. For example, components mounted at the board level, such as memory, accelerators, networking devices, modems, etc., can be integrated at the package level while remaining applicable from handheld servers to high-end servers. In such use cases, dies from potentially multiple sources can be connected through different packaging options, even on the same package.
[0128] A second use is to provide off-package connections using different types of media (e.g., optical, electrical cables, millimeter wave) with UCIe retimers to enable resource pooling, resource sharing, and / or message passing for transporting underlying protocols (e.g., PCIe, CXL) at the rack or pod level with load-store semantics beyond the node level to achieve better power savings and cost-effective performance in edge and data centers.
[0129] As described above, embodiments may be implemented in a data center use case, for example, in relation to a rack or pod. As an example, multiple compute nodes from different compute chassis may be connected to a CXL switch. As a result, the CXL switch may be connected to multiple CXL accelerators / Type-3 memory devices that may be disposed in one or more separate drawers.
[0130] Referring now to FIG. 14, a block diagram of another exemplary system according to an embodiment is shown. In FIG. 14, system 1400 may be all or part of a rack-based server having a plurality of hosts in the form of compute drawers that may be coupled to pooled memory via one or more switches.
[0131] As shown, there are a plurality of hosts 1430-1-n (also referred to herein as "host 1430"). Each host can be implemented as a compute drawer having one or more SoCs, memory, storage, and interface circuits, etc. In one or more embodiments, each host 1430 may include one or more virtual tiers corresponding to different cache coherence domains. The host 1430 can be coupled to a switch 1420 that can be implemented as a UCIe or CXL switch (e.g., a CXL2.0 (or later) switch). In one embodiment, each host 1430 can be coupled to the switch 1420 using an off-package interconnect, e.g., a UCIe interconnect that executes the CXL protocol through at least one UCIe retimer (which may be present in one or both of the host 1430 and the switch 1420).
[0132] The switch 1420 may be coupled to a plurality of devices 1410-1-x (also referred to herein as "device 1410"), each of which may be a memory device (e.g., a type 3 CXL memory expansion device) and / or an accelerator. In the illustration of FIG. 14, each device 1410 is shown as a type 3 memory device having any number of memory regions (e.g., defined partitions, memory ranges, etc.). Depending on the configuration and use case, a particular device 1410 may include a memory region assigned to a particular host, or otherwise may include at least some memory regions designated as shared memory. Although embodiments are not limited in this regard, the memory included in the device 1410 can be implemented using any type of computer memory (e.g., dynamic random access memory (DRAM), static random access memory (SRAM), non-volatile memory (NVM), a combination of DRAM and NVM, etc.).
[0133] Referring now to FIG. 15, a block diagram of a system according to another embodiment, such as an edge platform, is shown. As shown in FIG. 15, the multiprocessor system 1500 includes a first processor 1570 and a second processor 1580 coupled via an interconnect 1550 that can be a UCIe interconnect according to one embodiment that executes a coherence protocol. As shown in FIG. 15, each of the processors 1570 and 1580 can be a multi-core processor that includes a number of representative first and second processor cores (i.e., processor cores 1574a and 1574b and processor cores 1584a and 1584b).
[0134] In the embodiment of FIG. 15, the processors 1570 and 1580 further include point-to-point interconnects 1577 and 1587 that couple to switches 1559 and 1560 via interconnects 1542 and 1544, which can be UCIe links according to one embodiment. As a result, the switches 1559, 1560 couple to pooled memories 1555 and 1565 (e.g., via UCIe links).
[0135] Referring further to FIG. 15, the first processor 1570 further includes a memory controller hub (MCH) 1572, point-to-point (P-P) interfaces 1576 and 1578. Similarly, the second processor 1580 includes an MCH 1582, P-P interfaces 1586 and 1588. As shown in FIG. 15, the MCHs 1572 and 1582 couple the processors to memories 1532 and 1534, respectively, which can be portions of system memory (e.g., DRAM) locally attached to their respective processors. The first processor 1570 and the second processor 1580 can each be coupled to a chipset 1590 via P-P interconnects 1576 and 1586. As shown in FIG. 15, the chipset 1590 includes P-P interfaces 1594 and 1598.
[0136] Furthermore, the chipset 1590 includes an interface 1592 for coupling the chipset 1590 to the high-performance graphics engine 1538 via the P-P interconnect 1539. As shown in FIG. 15, various input / output (I / O) devices 1514 can be coupled to the first bus 1516, along with a bus bridge 1518 that couples the first bus 1516 to the second bus 1520. In one embodiment, for example, various devices can be coupled to the second bus 1520, including a keyboard / mouse 1522, a communication device 1526, and a data storage unit 1528 such as a disk drive or other mass storage device that can include code 1530. Additionally, audio I / O 1524 can be coupled to the second bus 1520.
[0137] Referring now to FIG. 16, a block diagram of a system 1600 according to another embodiment is shown. As shown in FIG. 16, the system 1600 can be any type of computing device and, in one embodiment, can be a server system. In the embodiment of FIG. 16, the system 1600 includes a plurality of CPUs 1610a, b that are coupled to respective system memories 1620a, b (which can be implemented as DIMMs, such as double data rate (DDR) memory, persistent or other types of memory, in the embodiment). Note that the CPUs 1610 can be coupled together via an interconnect system 1615 such as UCIe or other interconnect that implements a coherence protocol.
[0138] A plurality of interconnects 1630a1-b2 can be present to enable coherent accelerator devices and / or smart adapter devices to be potentially coupled to the CPUs 1610 via multiple communication protocols. Each interconnect 1630 can be a given instance of a UCIe link according to one embodiment.
[0139] In the illustrated embodiment, each CPU 1610 is coupled to a corresponding field programmable gate array (FPGA) / accelerator device 1650a, b (which may include a GPU in one embodiment). Additionally, the CPU 1610 is also coupled to smart NIC devices 1660a, b. As a result, the smart NIC devices 1660a, b are coupled to switches 1680a, b (e.g., a CXL switch according to one embodiment). The switches are in turn coupled to pooled memories 1690a, b such as persistent memory. In an embodiment, the various components shown in FIG. 16 may implement circuits for performing the techniques described herein.
[0140] The following examples relate to further embodiments.
[0141] In one example, the apparatus comprises a first die, wherein the first die, a die-to-die adapter for communicating with a protocol layer circuit and a physical layer circuit, wherein the die-to-die adapter receives message information including first information of a first interconnect protocol; and the physical layer circuit coupled to the die-to-die adapter, wherein the physical layer circuit receives the first information via an interconnect and outputs it to a second die and has, wherein the physical layer circuit, a first plurality of transmitters for transmitting data via a first plurality of data lanes; and at least one redundant transmitter, wherein the physical layer circuit remaps a first data lane of the first plurality of data lanes to the at least one redundant transmitter and includes.
[0142] In one example, the apparatus, a first plurality of bumps adapted on the first die, wherein the first plurality of bumps are associated with the first plurality of data lanes; and At least one redundant bump adapted on the first die, wherein the physical layer circuit remaps the first data lane from a first bump of the plurality of bumps to the at least one redundant bump further comprises.
[0143] In one example, the at least one redundant transmitter a first redundant transmitter, wherein the physical layer circuit remaps the first data lane to the first redundant transmitter to repair a lane failure in the first data lane; and a second redundant transmitter, wherein the physical layer circuit remaps a second data lane of the first plurality of data lanes to the second redundant transmitter to repair a lane failure in the second data lane includes.
[0144] In one example, the physical layer circuit repairs two data lanes out of a group of 32 data lanes via the first redundant transmitter and the second redundant transmitter.
[0145] In one example, the apparatus further comprises a first plurality of multiplexers coupled to the first plurality of transmitters, wherein the physical layer circuit controls the first plurality of multiplexers to pass data from one of a corresponding data lane, a first adjacent data lane, or a second adjacent data lane.
[0146] In one example, the physical layer circuit uses a first portion of the first plurality of multiplexers to perform a left shift operation so that a lane failure in the first data lane is repaired; and uses a second portion of the first plurality of multiplexers to perform a right shift operation By being executed, a lane failure in the second data lane is repaired.
[0147] In one example, the apparatus includes a first plurality of receivers for receiving second message information via a second plurality of data lanes; and at least one redundant receiver, wherein, in response to a failure in a first data lane of the second plurality of data lanes, the physical layer circuit remaps the second data lane of the second plurality of data lanes to the at least one redundant receiver and further includes.
[0148] In one example, the physical layer circuit includes a first clock transmitter for transmitting a clock signal via a first clock lane; and at least one redundant clock transmitter, wherein the physical layer circuit remaps the first clock lane to the at least one redundant transmitter and further has.
[0149] In one example, the physical layer circuit reverses the logical lane order of at least some of the first plurality of data lanes.
[0150] In one example, when a first data lane associated with a first transmitter of the first plurality of transmitters is coupled to an Nth data lane of a second die, the physical layer circuit reverses the logical lane order, where N is equal to the number of data lanes in the module.
[0151] In one example, the physical layer circuit remaps at least one data lane of the first plurality of data lanes in response to a failure in the first data lane and reverses the logical lane order of at least some of the first plurality of data lanes.
[0152] In another example, the method includes identifying, via a physical layer circuit of the first die of a package that includes a first die, a second die, and an interconnect that couples the first die and the second die, a failure in a first physical data lane of a plurality of physical data lanes of a main band of the interconnect, the interconnect including the main band and a side band; remapping, in response to the identification of the failure, first data traffic of a first logical data lane via the physical layer circuit of the first die onto a second physical data lane of the main band; and communicating, via the side band, information regarding the remapping to the second die. The method includes.
[0153] In one example, the method further includes remapping, via a physical layer circuit of the second die, the first data traffic of the first logical data lane from the second physical data lane within the second die back to the first logical data lane.
[0154] In one example, the method includes identifying, via the physical layer circuit of the first die, a failure in a clock physical data lane of the interconnect; remapping, in response to the identification of the failure in the clock physical data lane, a clock signal via the physical layer circuit of the first die from the clock physical data lane to a redundant clock physical data lane; and communicating, via the side band, information regarding the remapping of the clock signal to the second die. The method further includes.
[0155] In one example, the method includes providing the first data traffic of the first logical data lane to a first multiplexer of the first die; and Controlling the first multiplexer via the physical layer circuit of the first die to provide the first data traffic to a second transmitter of the first die associated with the second physical data lane further comprises.
[0156] In one example, the method further comprises disabling a first transmitter of the first die associated with the first physical data lane via the physical layer circuit.
[0157] In another example, a computer-readable medium comprising instructions executes any of the methods of the above examples.
[0158] In a further example, a computer-readable medium comprising data is used by at least one machine to manufacture at least one integrated circuit that executes any one of the methods of the above examples.
[0159] In yet another further example, an apparatus comprises means for executing any one of the methods of the above examples.
[0160] In another example, the package includes a first die having a CPU and a protocol stack, and a second die coupled to the first die via an interconnect. The first die includes a die-to-die adapter for communicating with a protocol layer circuit via FDI and a physical layer circuit via RDI, where the die-to-die adapter communicates message information including first information of a first interconnect protocol; the physical layer circuit coupled to the die-to-die adapter via the RDI, where the physical layer circuit receives the first information via the interconnect and outputs it to the second die, and where the physical layer circuit includes a first plurality of receivers for receiving data via a first plurality of physical data lanes; at least one redundant receiver, where the physical layer circuit shifts data traffic of the first plurality of physical data lanes to adjacent ones of the first plurality of physical data lanes and at least one redundant lane associated with the at least one redundant receiver in response to a lane failure in a first physical lane of the first plurality of physical data lanes.
[0161] In one example, the physical layer circuit activates at least one redundant transmitter and tri-states a first transmitter of a first plurality of transmitters associated with another physical lane in response to another lane failure.
[0162] In one example, the interconnect includes a main band including the first plurality of physical data lanes and a side band, and the physical layer circuit receives information regarding the shift of the data traffic from the second die via the side band.
[0163] In one example, the second die has an accelerator, and the first die communicates with the second die according to at least one of a flit mode of a PCIe protocol or a flit mode of a CXL protocol.
[0164] In another example, the apparatus Interconnection means including a main band and a side band, for identifying a failure in the first physical data lane means of a plurality of first physical data lane means of the main band of the interconnection means that couples first die means and second die means; Means for remapping first data traffic of the first logical data lane means onto a second physical data lane means of the main band; Means for communicating information regarding the remapping to the second die means via the side band Comprising.
[0165] In one example, the apparatus further comprises means for remapping the first data traffic of the first logical data lane means from the second physical data lane means within the second die means back to the first logical data lane means.
[0166] In one example, the apparatus Means for identifying a failure in the clock physical data lane means of the interconnection means; Means for remapping a clock signal from the clock physical data lane means to a redundant clock physical data lane means; and Means for communicating information regarding the remapping of the clock signal to the second die means via the side band Further comprising.
[0167] In one example, the apparatus Means for providing the first data traffic of the first logical data lane means to first multiplexer means of the first die means; and Means for controlling the first multiplexer means to provide the first data traffic to second transmitter means of the first die means associated with the second physical data lane means Further comprising.
[0168] In one example, the apparatus further comprises means for disabling the first transmitter means of the first die means associated with the first physical data lane means.
[0169] It should be understood that various combinations of the above examples are possible.
[0170] Note that the terms "circuit" and "circuitry" are used herein with the same meaning. In this specification, these terms and the term "logic" are used to refer to analog circuits, digital circuits, hardwired circuits, programmable circuits, processor circuits, microcontroller circuits, hardware logic circuits, state machine circuits, and / or any other type of physical hardware component, either alone or in any combination. Each embodiment may be used in many different types of systems. For example, in one embodiment, a communication device may be arranged to perform the various methods and techniques described herein. Of course, the scope of the present invention is not limited to communication devices. Instead, other embodiments may be directed to one or more machine-readable media containing instructions that, in response to being executed by another type of device for processing instructions, or a computing device, cause the device to perform one or more of the methods and techniques described herein.
[0171] Embodiments may be implemented in code and stored in a non-transitory storage medium storing instructions that can be used to program a system to execute the instructions. Each embodiment may also be implemented in data and stored in a non-transitory storage medium. The non-transitory storage medium causes at least one integrated circuit that executes one or more operations to be fabricated in at least one machine when used by the at least one machine. Another further embodiment may be implemented in a computer-readable storage medium containing information that configures a SoC or other processor to execute one or more operations when fabricated as the SoC or other processor. The storage medium may include, but is not limited to, any type of disk including floppy disks, optical disks, solid state drives (SSDs), compact disc read only memories (CD-ROMs), compact disc rewritables (CD-RWs), magneto-optical disks, read only memories (ROMs), random access memories (RAM) such as dynamic random access memories (DRAM) and static random access memories (SRAM), semiconductor devices such as erasable programmable read only memories (EPROMs), flash memories, electrically erasable programmable read only memories (EEPROMs), magnetic or optical cards, or any other type of medium suitable for storing electronic instructions.
[0172] Although the present disclosure has been described with respect to a limited number of implementations, those of ordinary skill in the art having the benefit of the present disclosure will appreciate a number of modifications and variations therefrom. The appended claims are intended to cover all such modifications and variations. [Item 1] An apparatus comprising a first die, wherein the first die a die-to-die adapter for communicating with a protocol layer circuit and a physical layer circuit, where the die-to-die adapter receives message information including first information of a first interconnect protocol; and The physical layer circuit coupled to the die-to-die adapter, where the physical layer circuit receives the first information via an interconnect and outputs it to a second die and has the physical layer circuit a first plurality of transmitters for transmitting data via a first plurality of data lanes; and at least one redundant transmitter, where the physical layer circuit remaps a first data lane of the first plurality of data lanes to the at least one redundant transmitter and includes a device. [Item 2] a first plurality of bumps adapted on the first die, where the first plurality of bumps are associated with the first plurality of data lanes; and at least one redundant bump adapted on the first die, where the physical layer circuit remaps the first data lane from a first bump of the plurality of bumps to the at least one redundant bump The device according to item 1, further comprising. [Item 3] The at least one redundant transmitter includes a first redundant transmitter, where the physical layer circuit remaps the first data lane to the first redundant transmitter to repair a lane failure in the first data lane; and a second redundant transmitter, where the physical layer circuit remaps a second data lane of the first plurality of data lanes to the second redundant transmitter to repair a lane failure in the second data lane and includes the device according to item 1. [Item 4] The device according to item 3, where the physical layer circuit repairs two data lanes out of a group of 32 data lanes via the first redundant transmitter and the second redundant transmitter. [Item 5] The apparatus according to item 1, further comprising a first plurality of multiplexers coupled to the first plurality of transmitters, wherein the physical layer circuit controls the first plurality of multiplexers to pass data from one of a corresponding data lane, a first adjacent data lane, or a second adjacent data lane. [Item 6] The physical layer circuit uses a first portion of the first plurality of multiplexers to perform a left shift operation so as to repair a lane failure in the first data lane; and uses a second portion of the first plurality of multiplexers to perform a right shift operation so as to repair a lane failure in the second data lane The apparatus according to item 5. [Item 7] A first plurality of receivers for receiving second message information via a second plurality of data lanes; and At least one redundant receiver, wherein in response to a failure in the first data lane of the second plurality of data lanes, the physical layer circuit remaps the second data lane of the second plurality of data lanes to the at least one redundant receiver The apparatus according to item 1, further comprising. [Item 8] The physical layer circuit A first clock transmitter for transmitting a clock signal via a first clock lane; and At least one redundant clock transmitter, wherein the physical layer circuit remaps the first clock lane to the at least one redundant transmitter The apparatus according to any one of items 1 to 7, further comprising The apparatus according to any one of items 1 to 7. [Item 9] The apparatus according to item 1, wherein the physical layer circuit reverses the logical lane order of at least some of the first plurality of data lanes. [Item 10] If the first data lane associated with the first transmitter of the first plurality of transmitters is coupled to the Nth data lane of the second die, the physical layer circuit reverses the logical lane order, where N is equal to the number of data lanes in the module, the apparatus of item 9. [Item 11] The physical layer circuit remaps at least one data lane of the first plurality of data lanes in response to a failure in the first data lane and reverses the logical lane order of at least some of the first plurality of data lanes, the apparatus according to any one of items 1 to 10. [Item 12] Identifying a failure in the first physical data lane of the first plurality of physical data lanes of the main band of the interconnect via the physical layer circuit of the first die of a package including a first die and a second die and an interconnect coupling the first die and the second die, the interconnect including the main band and a side band; Remapping, in response to identifying the failure, the first data traffic of the first logical data lane onto the second physical data lane of the main band via the physical layer circuit of the first die; and Communicating information regarding the remapping to the second die via the side band A method comprising. [Item 13] The method of item 12, further comprising remapping, via the physical layer circuit of the second die, the first data traffic of the first logical data lane from the second physical data lane in the second die to the first logical data lane. [Item 14] Identifying a failure in the clock physical data lane of the interconnect via the physical layer circuit of the first die; In response to identifying the fault in the clock physical data lane, remapping a clock signal from the clock physical data lane to a redundant clock physical data lane via the physical layer circuit of the first die; and communicating information regarding the remapping of the clock signal to the second die via the sideband The method according to item 12, further comprising. [Item 15] Providing the first data traffic of the first logical data lane to a first multiplexer of the first die; and Controlling the first multiplexer via the physical layer circuit of the first die to provide the first data traffic to a second transmitter of the first die associated with the second physical data lane The method according to item 12, further comprising. [Item 16] The method according to item 15, further comprising disabling a first transmitter of the first die associated with the first physical data lane via the physical layer circuit. [Item 17] A computer-readable storage medium comprising computer-readable instructions for implementing the method according to any one of items 12 to 16 when executed. [Item 18] An apparatus comprising means for executing the method according to any one of items 12 to 16. [Item 19] A package comprising a first die having a central processing unit (CPU) and a protocol stack and a second die coupled to the first die via an interconnect, wherein The first die is A die-to-die adapter for communicating with a protocol layer circuit via a framer aware die-to-die interface (FDI) and a physical layer circuit via a low die-to-die interface (RDI), wherein the die-to-die adapter communicates message information including first information of a first interconnect protocol; and The physical layer circuit coupled to the die-to-die adapter via the RDI, wherein the physical layer circuit receives the first information via the interconnect and outputs it to the second die and has the physical layer circuit a first plurality of receivers for receiving data via a first plurality of physical data lanes; and at least one redundant receiver, wherein the physical layer circuit shifts the data traffic of the first plurality of physical data lanes to adjacent ones of the first plurality of physical data lanes and at least one redundant lane associated with the at least one redundant receiver in response to a lane failure in the first physical lane of the first plurality of physical data lanes including package [Item 20] The package according to item 19, wherein the physical layer circuit activates at least one redundant transmitter in response to another lane failure and tri-states a first transmitter of a first plurality of transmitters associated with another physical lane [Item 21] The interconnect includes a main band including the first plurality of physical data lanes and a side band, and the physical layer circuit receives information regarding the shift of the data traffic from the second die via the side band. The package according to item 19 [Item 22] The second die has an accelerator, wherein the first die communicates with the second die according to at least one of a flit mode of a peripheral component interconnect express (PCIe) protocol or a flit mode of a compute express link (CXL) protocol. The package according to item 19 [Item 23] Interconnection means including a main band and a side band, for identifying a failure in the first physical data lane means of a plurality of first physical data lane means of the main band of the interconnection means that couples first die means and second die means; Means for remapping first data traffic of the first logical data lane means onto a second physical data lane means of the main band; Means for communicating information regarding the remapping to the second die means via the side band An apparatus comprising the same. [Item 24] The apparatus according to item 23, further comprising means for remapping the first data traffic of the first logical data lane means from the second physical data lane means in the second die means back to the first logical data lane means. [Item 25] Means for identifying a failure in the clock physical data lane means of the interconnection means; Means for remapping a clock signal from the clock physical data lane means to a redundant clock physical data lane means; and Means for communicating information regarding the remapping of the clock signal to the second die means via the side band The apparatus according to item 23, further comprising the same. [Item 26] Means for providing the first data traffic of the first logical data lane means to first multiplexer means of the first die means; and Means for controlling the first multiplexer means to provide the first data traffic to second transmitter means of the first die means associated with the second physical data lane means The apparatus according to item 23, further comprising the same.
Claims
1. An apparatus comprising a first die, wherein the first die is a die-to-die adapter that communicates with a protocol layer circuit and a physical layer circuit, the die-to-die adapter receiving message information including first information of a first interconnect protocol, the die-to-die adapter; the physical layer circuit coupled to the die-to-die adapter, the physical layer circuit having the physical layer circuit that receives the first information via an interconnect and outputs it to a second die, the physical layer circuit a first plurality of transmitters that transmit data via a first plurality of data lanes, and redundant transmitters respectively provided at start and end positions of the first plurality of data lanes, the physical layer circuit remapping first data traffic of a first data lane among the first plurality of data lanes to an adjacent data lane in a first direction or a second direction opposite to the first direction of the first data lane, and shifting at least one data lane arranged in the first direction or the second direction with respect to the adjacent data lane in a direction of one of the redundant transmitters of the redundant transmitters, the apparatus.
2. a first plurality of bumps adapted on the first die, the first plurality of bumps being associated with the first plurality of data lanes, the first plurality of bumps; at least one redundant bump adapted on the first die, the physical layer circuit remapping the first data lane from a first bump among the first plurality of bumps to the at least one redundant bump, at least one redundant bump The apparatus according to claim 1, further comprising.
3. The redundant transmitter is A first redundant transmitter, wherein the physical layer circuit remaps the first data traffic of the first data lane to the first redundant transmitter to repair a lane failure in the first data lane, the first redundant transmitter; A second redundant transmitter, wherein the physical layer circuit remaps the second data traffic of the second data lane among the plurality of first data lanes to the second redundant transmitter to repair a lane failure in the second data lane comprising The apparatus according to claim 1.
4. The apparatus according to claim 3, wherein the physical layer circuit repairs two data lanes in a group composed of 32 data lanes via the first redundant transmitter and the second redundant transmitter.
5. The apparatus according to claim 1, further comprising a first plurality of multiplexers coupled to the first plurality of transmitters, wherein the physical layer circuit controls the first plurality of multiplexers to control the first plurality of multiplexers. Among the data lanes, the data from one of the data lane corresponding to each multiplexer, the adjacent data lane in the left direction of each multiplexer, or the adjacent data lane in the right direction of each multiplexer is passed.
6. The physical layer circuit is using a first portion of the first plurality of multiplexers to perform a left shift operation so that the lane failure in the first data lane is repaired, using a second portion of the first plurality of multiplexers to perform a right shift operation so that the lane failure in the second data lane is repaired The apparatus according to claim 3.
7. A first plurality of receivers for receiving second message information via a second plurality of data lanes, At least one redundant receiver, wherein the physical layer circuit remaps the first data lane among the second plurality of data lanes in the direction of the at least one redundant receiver in response to a failure in the first data lane among the second plurality of data lanes The apparatus according to claim 1, further comprising **Claim 8** The physical layer circuit A first clock transmitter that transmits a clock signal via a first clock lane, At least one redundant clock transmitter, wherein the physical layer circuit remaps the first clock lane to the at least one redundant transmitter The apparatus according to any one of claims 1 to 7, further comprising The apparatus according to any one of claims 1 to 7. **Claim 9** The apparatus according to claim 1, wherein the physical layer circuit reverses the logical lane order of at least some of the first plurality of data lanes. **Claim 10** When the first data lane associated with the first transmitter among the first plurality of transmitters is coupled to the Nth data lane of the second die, the physical layer circuit reverses the logical lane order, where N is equal to the number of data lanes in the module. The apparatus according to claim 9. **Claim 11** The physical layer circuit In response to a failure in the first data lane, remaps at least one of the first plurality of data lanes, The apparatus according to claim 10, wherein the physical layer circuit reverses the logical lane order of at least some of the first plurality of data lanes. **Claim 12** Identifying a failure in a first physical data lane of a plurality of first physical data lanes in a main band of the interconnect via a physical layer circuit of the first die of a package including the first die, the second die, and an interconnect coupling the first die and the second die, wherein the interconnect includes the main band and a side band, and the plurality of first physical data lanes includes redundant physical data lanes provided at a start position and an end position, respectively; In response to identifying the failure, remapping first data traffic of a first logical data lane via the physical layer circuit of the first die onto a second physical data lane adjacent to the first physical data lane in a first direction or a second direction opposite to the first direction of the first physical data lane, and shifting at least one physical data lane arranged in the first direction or the second direction with respect to the adjacent second physical data lane in a direction of one of the redundant physical data lanes of the redundant physical data lanes; and Communicating information regarding the remapping to the second die via the side band A method comprising.
13. The method according to claim 12, further comprising remapping the first data traffic of the first logical data lane from a second physical data lane in the second die to the first logical data lane via a physical layer circuit of the second die.
14. Identifying a failure in a clock physical data lane of the interconnect via the physical layer circuit of the first die; In response to identifying the failure in the clock physical data lane, remapping a clock signal from the clock physical data lane to a redundant clock physical data lane via the physical layer circuit of the first die; and Communicating information regarding the remapping of the clock signal to the second die via the side band The method according to claim 12, further comprising.
15. providing the first traffic of the first logical data lane to a first multiplexer of the first die; and controlling, via the physical layer circuit of the first die, the first multiplexer to provide the first traffic to a second transmitter of the first die associated with the second physical data lane The method according to claim 12, further comprising.
16. The method according to claim 15, further comprising invalidating a first transmitter of the first die associated with the first physical data lane via the physical layer circuit.
17. A computer program for causing a processor to implement the method according to any one of claims 12 to 16.
18. A computer-readable storage medium storing the computer program according to claim 17.
19. An apparatus comprising means for executing the method according to any one of claims 12 to 16.
20. A package comprising a first die having a central processing unit (CPU) and a protocol stack, and a second die coupled to the first die via an interconnect, wherein the first die is a die-to-die adapter that communicates with a protocol layer circuit via a framer die-to-die interface (FDI) and with a physical layer circuit via a router die-to-die interface (RDI), the die-to-die adapter communicating message information including first information of a first interconnect protocol; and the physical layer circuit coupled to the die-to-die adapter via the RDI, the physical layer circuit receiving the first information via the interconnect and outputting the first information to the second die having, the physical layer circuit a first plurality of receivers that receive data via a first plurality of physical data lanes, redundant receivers respectively provided at start and end positions of the first plurality of physical lanes, wherein the physical layer circuit, in response to a lane failure in a first physical data lane among the first plurality of physical data lanes, remaps data traffic of the first physical data lane to an adjacent physical data lane in a first direction or a second direction opposite to the first direction of the first physical data lane, and shifts at least one physical data lane arranged in the first direction or the second direction with respect to the adjacent physical data lane in a direction of a redundant lane associated with one of the redundant receivers, the redundant receiver including, package.
21. The physical layer circuit activates at least one redundant transmitter in response to another lane failure and tri-states a first transmitter of a first plurality of transmitters associated with another physical lane, the package according to claim 20.
22. The interconnection includes a main band and a side band, the main band includes the first plurality of physical lanes, and the physical layer circuit receives information regarding the shift of the data traffic from the second die via the side band, the package according to claim 20.
23. The second die has an accelerator, and the first die communicates with the second die according to at least one of a flit mode of a Peripheral Component Interconnect Express (PCIe) protocol or a flit mode of a Compute Express Link (CXL) protocol, the package according to claim 20.
24. Means for identifying a failure in a first physical data lane among a plurality of first physical data lanes of a main band included in an interconnection that couples a first die and a second die, wherein the interconnection includes the main band and a side band, and the plurality of first physical data lanes includes redundant physical data lanes respectively provided at a start position and an end position. Means for remapping first data traffic of a first logical data lane onto a second physical data lane adjacent to the first physical data lane in a first direction or a second direction opposite to the first direction of the first physical data lane, and shifting at least one physical data lane arranged in the first direction or the second direction with respect to the adjacent second physical data lane in a direction of one of the redundant physical data lanes of the redundant physical data lanes. Means for communicating information regarding the remapping to the second die via the side band. An apparatus comprising the above.
25. The apparatus according to claim 24, further comprising means for remapping the first data traffic of the first logical data lane from a second physical data lane in the second die to the first logical data lane.
26. Means for identifying a failure in a clock physical data lane of the interconnection. Means for remapping a clock signal from the clock physical data lane to a redundant clock physical data lane. Means for communicating information regarding the remapping of the clock signal to the second die via the side band. The apparatus according to claim 24, further comprising the above.
Citation Information
Patent Citations
Reception circuit, information processor, and control method
JP2013211687A
Distributed processor configuration for use in infusion pumps
US20120157920A1
Reception circuit, information processing apparatus, and control method
US20130262948A1
Shared error detection and correction memory
US20180025789A1
Multichip package link error detection
US20200356436A1