Switching to multiple endpoints across root complex of inter-die data interface

Multi-endpoint switching is realized in PCIe 5.0 system through multi-die integrated chip module and retimer logic, solving channel loss problem, expanding channel range, reducing noise and power consumption, and achieving efficient bandwidth sharing and fault management.

CN120457421APending Publication Date: 2025-08-08KANDOU LABS SA
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202380090703.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-11-09
Filing Date
2023-11-09
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

In PCIe 5.0 systems, channel loss exceeds the target requirements, resulting in signal attenuation and noise increase, and existing retimers cannot effectively expand the channel range and maintain low power consumption.

Method used

Multi-die integrated chip module (ICM) is used to establish data flow between circuit dies through inter-die data interface and retimer logic, realize multi-endpoint switching and bandwidth sharing, use BMC for monitoring and management, and configure channel routing logic to optimize signal transmission.

Benefits of technology

It effectively expands the channel range of PCIe links, reduces signal attenuation and noise, maintains low power consumption of the system, and realizes efficient bandwidth sharing and failover management of multiple endpoints.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120457421A_ABST
    Figure CN120457421A_ABST
Patent Text Reader

Abstract

A plurality of upstream dummy ports (PPs) of the first circuit die, each upstream PP having a connection to a respective one of the at least two root complex devices; a plurality of downstream PPs of the second circuit die, each set of downstream PPs having a connection to a respective one of the at least two endpoints; an inter-die data interface between the first and second circuit dies, the inter-die data interface having an adaptation layer port at each circuit die that operates according to an adaptation layer protocol; channel routing logic disposed within the first and second circuit dies to map at least one of the sets of upstream PPs and a respective set of downstream PPs to respective adaptation layer ports of the first and second circuit dies; and a processor disposed within one of the first and second circuit dies to configure the channel routing logic within the two circuit dies.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 382,901, entitled “Switching of a Root Complex to Multiple Endpoints Across an Inter-Die Data Interface,” filed on November 9, 2022, which is incorporated herein by reference in its entirety for all purposes. Background Art

[0003] With the data rate of PCIe 5.0 (32Gbps) increasing from previous generations (for example, PCIe 4.0 MAX has a data rate of 16Gbps), the channel reach is becoming shorter than before, making the need for retimers more apparent. Common channels include system boards, backplanes, cables, riser cards, and add-in cards. The loss of connections across these channels (usually a combination of such channels and their slots) often exceeds the target loss requirement of -36dB at 16GHz. Retimers can extend the channel reach beyond the reach limit when a retimer is not used.

[0004] The retimer splits the link between the host (root complex, abbreviated as RC) and the device (endpoint) into two independent segments. Therefore, the retimer can re-establish a new forward PCIe link, and this re-establishment process includes retraining and appropriate equalization for the physical layer and link layer.

[0005] Redrivers are purely analog amplifiers that boost signals to compensate for attenuation. However, while boosting the signal, redrivers also increase noise, which often exacerbates jitter. In contrast, retimers incorporate both analog and digital logic to equalize the signal, extract its clock timing information, and output a signal with high amplitude and low noise and jitter. Furthermore, retimers maintain power states to keep system power consumption low.

[0006] The retimer specification debuted in PCIe 4.0 and is expected to continue in PCIe 5.0. Figure 1 and Figure 2 Shown are common application scenarios of retimers in some embodiments. Figure 1 A retimer is used and placed on the motherboard so that it is logically located between the PCIe root complex (RC) and the PCIe endpoints.

[0007] Figure 2 The scenario shown uses two retimers, where the first retimer is also located on the motherboard, while the second retimer is located on a riser card that serves as a connection between the motherboard and an add-in card, where the PCIe endpoints are contained.

[0008] In complex PCIe systems, there may be far more PCIe endpoints than available PCIe ports. In such cases, a switch can be used to increase the number of PCIe ports. A switch connects multiple endpoints to a single root node and routes data packets to specific destinations, rather than mirroring data across all ports. A key feature of switches is bandwidth sharing, allowing all endpoints to share the root node's bandwidth. Summary of the Invention

[0009] The methods and systems described herein include an apparatus comprising: a plurality of sets of upstream pseudo ports (PPs) on a first circuit die, each upstream PP having a connection to a respective one of at least two root complex devices; a plurality of sets of downstream PPs on a second circuit die, each downstream PP having a connection to a respective one of at least two endpoints; an inter-die data interface between the first and second circuit dies, the inter-die data interface configured to establish a retimer physical coding sublayer (RPCS) data flow between the upstream and downstream PPs of the first and second circuit dies through an adaptation layer port within each circuit die in accordance with an adaptation layer protocol; channel routing logic within the first and second circuit dies configured to map at least one of the sets of upstream PPs and a respective set of downstream PPs to respective adaptation layer ports in the first and second circuit dies in accordance with the adaptation layer protocol; and a processor configured to configure the channel routing logic of the first and second circuit dies.

[0010] This "Summary" section provides a brief summary of the concepts described in detail in the "Detailed Description" section below. This "Summary" section is not intended to provide key or primary features of the claimed technical solution, nor is it intended to assist in determining the scope of the claimed technical solution. Other objectives and / or advantages of the embodiments of the present invention will become readily apparent to those skilled in the art by referring to the "Detailed Description" section and the accompanying figures. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 and Figure 2 Shown are two uses of a retimer in some embodiments.

[0012] Figure 3 A block diagram of a chip structure of a multi-die integrated chip module (ICM) that performs multi-endpoint switching between multiple root complexes using high-speed die-to-die (D2D) interconnects in some embodiments.

[0013] Figure 4FIG. 1 is a data flow diagram for a multi-die ICM operating in retimer mode with data path routing within the same die in some embodiments.

[0014] Figure 5 FIG. 1 is a data flow diagram for a multi-die ICM operating in retimer mode with D2D interconnects for routing data paths between circuit dies in some embodiments.

[0015] Figure 6 FIG. 1 is a block diagram of a multiplexing crossbar switch used for data path routing in some implementations.

[0016] Figure 7 Patterning of D2D interconnection in some embodiments.

[0017] Figure 8 FIG. 4 is a block diagram of an adaptation layer for D2D interconnection in some implementation modes.

[0018] Figure 9 FIG. 4 is a block diagram of a chiplet-to-chiplet (T2T) serial peripheral interface (SPI) bus structure in a four-chiplet implementation.

[0019] Figure 10 FIG. 1 is a block diagram of the complete signal path between the central processing unit (CPU) core 900 and the PHY in each chip in a multi-chip module.

[0020] Figure 11 is a flow chart of a method in some embodiments. DETAILED DESCRIPTION

[0021] While the ability to fully integrate multiple systems into the same integrated circuit is increasing, maintaining multiple chip systems and subsystems separately still has significant advantages. For non-limiting descriptive purposes, at least some aspects of the present invention described herein are illustrated in a system environment comprising at least one point-to-point communication interface connecting two integrated circuit chips, representing a root complex (i.e., a host) and an endpoint, respectively, wherein the communication interface is supported by a number of data channels, each of which consists of four high-speed transmission line signal conductors.

[0022] A retimer typically consists of a physical physical layer (PHY) and retimer core logic. The PHY consists of a receiver and a transmitter. The PHY receiver performs data recovery, deserialization, and clock recovery, while the PHY transmitter serializes data and provides amplification before output transmission. The retimer core logic is responsible for deskew (for multi-channel links) and rate adaptation to address frequency differences between ports on both sides.

[0023] Retimers offer additional value because they sit in the path between the root complex (e.g., CPU) and endpoints (e.g., cache blocks). Retimers can also incorporate integrated processing units (e.g., accelerators) to handle data processing in the path from the root complex to the endpoint.

[0024] The PCIe retimer circuit is a chiplet (bare die) with a four-channel retimer and can be connected to a DPU chiplet or another retimer chiplet through a high-speed die-to-die interconnect. Among them, one, two or four channels can form a multi-channel link, and data can be transmitted across all links. Alternatively, each channel can be constructed as a single-channel link. Each channel in the PCIe retimer has two PHYs, one at each end (i.e., the upstream port and the downstream port). In the case of four channels, a PCIe retimer die has eight PHYs. In addition, the PCIe retimer die also has communication lines to enable the exchange of control information between two or more PCIe retimer dies.

[0025] With one (or more) PCIe retimer cores, the following structures can be constructed (these structures are described in further detail below):

[0026] -Quad-channel retimer;

[0027] - Single die with full-flexibility 4×4 static channel routing;

[0028] - Four-channel retimer with accelerator (DPU);

[0029] Two dies in the same package: a retimer die and a DPU die.

[0030] - Eight-channel retimer;

[0031] Two dies within the same package structure, but with limited static channel routing – Flexible 4×4 routing within the same die, but without the ability to cross die boundaries;

[0032] - Eight-channel retimer with full-flexibility channel routing;

[0033] Two dies within the same package structure use high-speed inter-die interconnects to route data across different die, but this incurs additional latency.

[0034] - Eight-channel retimer with accelerator (DPU);

[0035] Three dies in the same package: two retimer dies and one DPU die.

[0036] - Sixteen-channel retimer;

[0037] Four dies within the same package structure, but with limited static channel routing – Flexible 4×4 routing within the same die, but without the ability to cross die boundaries.

[0038] Multi-die ICM with multi-endpoint switching via D2D interface

[0039] Figure 3 FIG3 is a block diagram of a multi-die ICM 300 according to an embodiment of the present invention. As shown, ICM 300 includes a set of serial data transceivers (SerDes, PHYs) for multiple upstream pseudo ports (PPs) of a first circuit die 305, each of which has a connection to a respective one of at least two root complex devices 302 and 304. The apparatus also includes a second circuit die 310, each of which has a respective set of PHYs for respective downstream PPs, each of which has a connection to a respective one of at least two endpoints 315 and 320.

[0040] Figure 3 Also included is an inter-die data interface (D2D) provided between the first circuit die 305 and the second circuit die 310. The D2D interface is used to establish a retimer data flow between the upstream PP and the downstream PP of the first and second circuit dies via the adaptation layer port within each circuit die according to the adaptation layer protocol. The adaptation layer protocol can be used to format the raw data received by the PHY of one type of (upstream / downstream) pseudo port for transmission to the PHY of the opposite type of (downstream / upstream) pseudo port via the D2D interface. As shown below in combination Figure 7 and Figure 8In further detail, the D2D interface uses multiple orthogonal differential vector signaling code (ODVS) streams. It should be noted that in addition to this, other D2D interfaces may also be used, such as the Universal Chiplet Interconnect Express (UCIe) interface. Each of the first and second circuit dies also includes a channel routing logic 600, which is used to map at least one of each group of upstream PPs and a corresponding group of downstream PPs to the corresponding adaptation layer ports of the first and second circuit dies according to the adaptation layer protocol. The device also includes a processor (such as a CPU core) provided in one of the first and second circuit dies, which is used to configure the channel routing logic in the first and second circuit dies. In the multi-die ICM 300, one of the first and second circuit dies is the main circuit die. Although both circuit dies may include a processor provided therein, only the processor of the main circuit die is in an active state. The processor of the main circuit die can configure the channel routing logic in the slave circuit die through the chip-to-chip serial peripheral interface (SPI), as described below. Figure 9 and Figure 10 This is described in further detail in the description of .

[0041] Figure 3 A board management controller (BMC) 325 is included. The BMC may be included in a motherboard, for example, and uses sensors to monitor the status of motherboard components and hardware devices, and transmits these statuses to, for example, the root complex. The BMC may be used, for example, in server room / data center applications and can be remotely managed by an administrator to access information related to the entire system. Some of the BMC's monitoring functions include temperature monitoring, humidity monitoring, power supply voltage monitoring, fan speed monitoring, communication parameter monitoring, and operating system monitoring. When any parameter exceeds a threshold, the BMC can notify the administrator, allowing the administrator to take action. In some embodiments, the BMC may be preconfigured to take certain actions when a parameter exceeds a threshold, such as (but not limited to) executing a failover sequence to a redundant endpoint in the event of a primary endpoint failure. In some embodiments, the BMC 325 monitors the status of the PCIe link between the root complex 302 / 304 and the endpoints 315 / 320. In such embodiments, monitoring the PCIe link status includes measuring the bit error rate of the upstream and downstream data paths. These measurements can be used to determine the overall status of the PCIe link and initiate a link retraining sequence.

[0042] In addition to the above monitoring, BMC 325 can be used to Figure 3In some embodiments, the BMC 325 may manage multiple root complexes in a single device. Specifically, endpoints 315 and 320 may correspond to sharable resources for high-cost functions such as artificial intelligence (AI), sharable computer-readable media (such as hard disk drives (HDDs) or solid-state drives (SSDs)), and network interface cards (NICs) between other endpoints. In such embodiments, the BMC 325 may coordinate the use of endpoints by root complex devices, that is, so that no two root complex devices establish connections to the same endpoint at the same time. In some embodiments, the BMC may utilize credit-based technology to enable sharing of multiple endpoints among multiple root complex devices.

[0043] The BMC may be used to provide instructions to the CPU core of the master core of the ICM 300. Such instructions may be provided, for example, via an SMBus connection or various other point-to-point connections, and may be associated with a root complex to endpoint mapping. The CPU of the master core may configure the channel routing logic within the master core and the slave core to map the upstream pseudo port to the downstream pseudo port associated with the mapping instruction issued by the BMC. In some embodiments, the configuration of the channel routing logic includes modifying a configuration register space within the two circuit dies, wherein the configuration register space includes control signal values that are provided as selection signals to a multiplexing device within the channel routing logic. In some embodiments, as described below, the logic channel of the upstream pseudo port has a static mapping configuration that maps to an adaptation layer port. For example, Figure 6 The upstream pseudo port PHY1 in the master core can be statically mapped to the adaptation layer port 1 of the master core's adaptation layer transmit (TX) portion, and the downstream pseudo port of the slave core can selectively connect to any of the adaptation layer ports 0 to 7 based on the root complex-endpoint mapping. In some embodiments, the downstream pseudo port can be statically mapped to an adaptation layer port, while the upstream pseudo port is used to connect to any adaptation layer port, and vice versa.

[0044] Figure 4 1 is a data flow diagram for a PCIe data link to a first endpoint 315 in some embodiments. As shown, a PHY within an upstream pseudo port receives serial data and includes a deserializer for converting the serial data stream into, for example, 32-bit channel-specific data codewords after deserialization. The data codewords are routed to the retimer core logic via channel routing logic. In some embodiments, the core logic includes a PCS decoding module for performing, for example, 8b / 10b or 128b / 130b decoding before storing in the retimer FIFO. The retimer FIFO includes channel de-skew and rate adaptation functionality across multiple channels within a given circuit die, as well as channel de-skew and rate adaptation functionality between channels across multiple different circuit dies. After reading the channel-specific data codewords from the retimer FIFO, the downstream serial data transceiver transmits them via a transmitter within the PHY of the downstream pseudo port.

[0045] Figure 5 3 is a data flow diagram for a PCIe data link to be used via a D2D interface to the second endpoint 320. Specifically, Figure 5 As shown, data is first received by a set of serial data transceivers and then provided to the second circuit die via an inter-die data interface (such as an adaptation layer and an inter-die transmitter). Figure 5 In the example, data is received by the PHY of the first circuit die. After deserialization, the data is routed through the channel routing logic to the adaptation layer within the first circuit die, which formats the raw data for transmission over the D2D interface. The adaptation layer within the second circuit die receives the data and reverse-formats it to provide it to the target channel in the second circuit die. Before the data is output to the endpoint via the serial data transceiver PHY of the second circuit die, it is also provided to the RPCS logic for rate adaptation and inter-channel de-skew. A similar data path exists in the opposite direction from the endpoint to the root complex.

[0046] like Figure 4 and Figure 5 As shown, the RPCS logic may, for example, include 8b / 10b encoding / decoding functionality for PCIe Gen 1 and Gen 2, and 128b / 130b encoding / decoding functionality for PCIe Gen 3 through Gen 5. The embodiments described herein further contemplate PCIe Gen 6, which utilizes a flow control unit (FLIT) scheme, eliminating the need for 8b / 10b or 128b / 130b. In such embodiments, encoding / decoding functionality may be omitted, but at the same time, PCIe Gen 6-specific functionality, such as FEC decoding (both partial and full decoding), is further included in the datapath in the form of logic. Some functionality of the retimer core logic is shared, such as FIFO inter-channel deskew and rate adaptation.

[0047] Figure 6 FIG. 6 is a block diagram of channel routing logic 600 of an ICM retimer circuit die in some embodiments. Figure 6Includes the block diagram on the left and various channel routing structure diagrams on the right. In the top channel routing structure 605, data is fed into the PHY after being fed through the deserializer. After passing through the core routing logic and the same PHY, it is output to the bottom through the serializer. In the middle structure diagram 610, data is fed into a port, processed by the core routing logic, and then fed out at the opposite PHY at the bottom. In the bottom structure diagram 615, all data is fed into each PHY on the top side of a PCIe retimer circuit and forwarded directly from these PHYs to the high-speed die-to-die interconnect. After passing through the core channel routing logic, the data is further fed into each PHY within another PCIe retimer die. All of these scenarios further have data paths in the opposite direction.

[0048] Figure 6 The left side shows a schematic diagram of the channel routing logic. Each serial data transceiver PHY is numbered 0 to 7 and includes a receive deserializer (DES) and a transmit serializer (SER). The top channel (PHY0 and PHY4) shows three different data paths corresponding to the data paths shown on the right. Figure 6 The data path 605 on the right corresponds to the path shown on the left for data to be input to and output from PHY0 of the PCIe retimer circuit. Figure 6 Path 610 in the figure corresponds to the feedthrough path for data received by PHY0 and passed to PHY4 on the left. Path 615 corresponds to the following path: all received data is forwarded directly to the adaptation layer for transmission via the inter-die data interface; data from the inter-die data interface is forwarded by the second PCIe retimer to the core routing logic; and data is processed by the core logic and output by the included PHY.

[0049] The second channel (PHY1 and PHY5) is used for multiplexing. Each core logic / transmit path can receive data from any of the eight channels and can also obtain data from the adaptation layer port of the D2D data interface. The other channels (PHY0 and PHY4, PHY2 and PHY6, PHY3 and PHY7) have the same switching function. The bottom shows the multiplexing of the adaptation layer port. As shown in the figure, for a given adaptation layer port, any input PHY can be selected as the input source. As shown in the following combination Figure 7 As described above, there may be a fixed mapping between the adaptation layer ports and the D2D data streams. Therefore, in some embodiments, data mirroring may be achieved by selecting the same receiving PHY data for multiple adaptation layer physical ports.

[0050] Data path switching in the routing logic involves: using a 32-bit receive data bus to carry the deserialized channel-specific data codewords; enabling the corresponding data lines; clock recovery; and the corresponding reset operation. After routing the channels to the adaptation layer, the deserialized data is read into the FIFO using the recovered clock signal, causing clock domain crossing. The adaptation layer handles clock timing, for example, so that each D2D data stream outputs a 150-bit D2D codeword to the transmitter; and converts to the 25Gbd clock domain used by the 5b6w interface. It is important to note that only the raw data is multiplexed; no processing is performed on the received data. The raw data multiplexing logic is statically configured using configuration bits, making the switching itself asynchronous. Changing the raw data multiplexing settings during task mode may result in invalid data and may cause glitches on the clock lines. Therefore, the multiplexing logic can be configured during reset.

[0051] In some embodiments, each circuit die includes channel routing logic, such as raw data multiplexing logic for performing channel routing between different circuit dies or within the same circuit die. In one such embodiment, the raw data multiplexing function within each circuit die can be configured by a master circuit die (also referred to as a "master device"), for example, by writing to a configuration register associated with the raw data multiplexing. Figure 9 and Figure 10 Such a chiplet-to-chiplet communication is shown. Figure 9 Figure 2 shows the T2T SPI bus architecture for a four-chiplet scenario. This specific number of chips is not limiting, and the principles described here can be extended to an N-chiplet retimer with one master chiplet and N-1 slave chipslets, where N ≥ 2.

[0052] The master T2T SPI 985 includes a serial clock line SCK, which carries the serial clock signal generated by the master T2T SPI 985. The SCK signal is received by all slave T2T SPIs and is used to coordinate the reading and writing of data over the T2T SPI bus.

[0053] The master T2T SPI 985 also includes a Leader Out Follower In (MOSI) line and a Leader In Follower Out (MISO) line. The MOSI line is used to send data from the master device to the slave device as part of a write operation. The MISO line is used to send data from the slave device to the master device as part of a read operation.

[0054] The master T2T SPI 985 further includes a slave chip select Line, which is used to indicate which slave device will participate in the current operation of the bus, that is, which slave device the data or command in the bus is aimed at. Figure 9 Only one wire is shown for selecting a line from the core particle, but in actual situations, there may be one wire for each line, that is, Figure 9 In the case shown, there may be three different slave die selection wires.

[0055] In addition, each slave T2T SPI 975a, 975b, 975c is connected to all of the above lines to enable bidirectional communication between the master and slave T2Ts, thereby enabling communication between chiplets.

[0056] Figure 10 Shown is the complete signal path between the CPU core 900 and each PHY in each chiplet of the multi-chiplet module.

[0057] The CPU core 900 is connected to the PHY 970 within the master chiplet via the master chiplet APB interconnect 925, thereby enabling communication with the PHY 970 via the APB interconnect 925. The CPU core 900 is also connected to the master T2T SPI 985 via the master chiplet APB interconnect 925. The master T2T SPI 985 is part of the T2T SPI bus that enables the CPU core 900 to communicate with other chiplets.

[0058] like Figure 10 As shown, each slave chip comprises a corresponding slave T2T SPI 975a, 975b, 975c. Each slave SPI is connected to the master T2T SPI 985 to achieve signal transmission between the chiplets.

[0059] Each slave SPI 975a, 975b, 975c is connected to a corresponding PHY 970a, 970b, 970c via the corresponding slave chip APB interconnect 926, 927, 928. Each slave SPI 975a, 975b, 975c acts as a master device for the corresponding APB interconnect 926, 927, 928. In this way, each slave SPI can access all registers within the chip where the slave SPI is located.

[0060] As can be seen, communication between chiplets uses two different buses and protocols. The SPI protocol does not support addressing, while the APB protocol does. Part of the data placed on the T2T SPI bus by the CPU core 900 is APB address information, enabling the local APB interconnect of each slave chiplet to route the message to the target receiving PHY.

[0061] Each PHY is assigned a unique APB address or APB address range, so that the CPU core 900 can read and write to a specific PHY in any chip. From the perspective of the CPU core 900, the entire multi-chip module has a single address space that covers the different address regions of each PHY.

[0062] For ease of explanation, assuming that the APB address is 24 bits and the data codeword size is 32 bits, the control information provided to the SPI bus may take the following format, which is referred to herein as "control data packet".

[0063] r r r r r s s s a a a a a a a a a a a a a a a a a a a a a a a a

[0064] Among them, bits 0 to 23 are address bits ("a"), bits 24, 25, and 26 are slave core selection bits, and bits 27 to 31 are reserved bits ("r"). In this specific case, since there are three slave cores (i.e., three slave T2T SPIs), the number of slave core selection bits is three. The reserved bits leave room for other slave core selection bits—in this case, up to eight slave core selection bits can be provided, thus supporting a maximum of eight slave cores. This principle can be further extended to any number of slave cores by increasing the codeword size.

[0065] The address bits constitute the APB address. Each slave T2T-SPI is configured as the master bus for its corresponding local APB interconnect, enabling each slave T2T-SPI to instruct its corresponding APB interconnect to perform a write or read operation on the corresponding PHY connected to each APB bus. In some cases, address data may not be used, and instead a T2T-SPI bus may be employed that automatically increments the address to obtain the address to write or read data. The address data may be provided to the local APB interconnect after the corresponding slave T2T-SPI receives the control packet, enabling the local APB interconnect to route commands and data to the correct local PHY.

[0066] The slave chip selection bit enables the control packet to indicate the slave chip selection line to be started, that is, the chip to be read or written. The T2T SPI bus controls the slave chip selection line through the slave chip selection bit. For example, 0 indicates that the corresponding slave chip selection circuit should be at a low potential, and 1 indicates that the corresponding slave chip selection circuit should be at a high potential.

[0067] Alternatively, the slave chip select control information can be sent separately from the APB address data. The slave chip select information can be sent in-band as described above, or using another channel such as the System Management Bus (SMBus). The address data can be sent separately before the PHY configuration packet. In some cases, address data can be omitted because the T2T-SPI bus can automatically increment the address to determine the address of the data to be written.

[0068] In each of these scenarios, data can be sent after the slave chip select information and address information are provided (the address information may or may not be provided as needed). The master T2T SPI 985 can keep the slave chip select line set until it receives new instructions related to the configuration of the slave chip select line. Similarly, the associated APB interconnect can continue to write to the specified address (or auto-increment address) until new addressing information is provided. In this way, data and commands can be sent or received to any PHY in any chip.

[0069] The APB address space is a global address space across all corelets. That is, any register in any corelet can be addressed through this global address space. In a specific structure, each corelet is provided with a base address in the form of a corelet identifier multiplied by a constant. The corelet identifier can be the corelet number, and the constant can be the base address of the main corelet. In addition to this, other storage space structures can also be used. Each register in each corelet is assigned a unique address or address range within the global address space. In this way, each PHY in PHY 970, 970a, 970b, 970c is assigned a unique address or address range.

[0070] The CPU core in the master core can coordinate the channel switching circuits in the two cores. The CPU core in the slave core can be in a low-power state. As shown in the figure, the SPI communication bus between the two cores can be used to configure the switching circuit in the slave core to select between the first and second groups of downstream serial data transceiver ports. In some embodiments, a die-to-die (D2D) interface can be set up and the channel routing between the master core and the slave core can be configured using the D2D interface, that is, the serial data stream received by the upstream port of the master core is routed to the downstream port of the slave core, and vice versa. Such a D2D interface can also be used to transmit configuration information as sideband information from the master core to the slave core, for example, to configure the configuration registers of the slave core. In another embodiment, the raw data cross-multiplexer can be configured via a system management bus, which can also be connected to the root complex. In some embodiments, the configuration purpose can be achieved through a virtual channel between the root complex and the retimer chip. In such an embodiment, a vendor-defined message (VDM) may be included in a specific vendor-defined packet field of a PCIe data transmission. Such a VDM may be detected and extracted, for example, via an interrupt protocol and provided to the CPU of the master die. Figure 7 There is only a single slave chiplet, but it should be noted that other slave chips may be included, and in some embodiments, up to three slave chips may be included. In such a case, each slave chiplet may have a specific chiplet ID, and configuration register write commands may be assigned to specific chiplet IDs.

[0071] In some embodiments, the master chiplet may initialize the configuration registers of the slave chiplet's raw data multiplexer so that the receive (RX) adaptation layer port is statically mapped to the downstream port leading to the redundant endpoint. In one such embodiment, the master chiplet may switch the routing of the deserialized channel-specific data codewords between: (i) the downstream port leading to the master endpoint within the same die and (2) the adaptation layer to be routed via the D2D interface.

[0072] Figure 7 This figure is a block diagram of an inter-die data interface (also referred to herein as a "high-speed die-to-die (D2D) interconnect," "D2D link," etc.) in some embodiments. As shown, the D2D link has eight high-speed die-to-die (D2D) data streams, four in each direction. Each D2D data stream operates at 25 GBd, sending 5 bits over 6 wires, for a total throughput of 125 Gbps. Furthermore, the interface includes two differential clock channels operating at 6.25 GHz. Note that interconnects with other sizes, throughputs, and / or encoding schemes may also be employed.

[0073] The original bandwidth of each data stream is within 125Gbps without forward error correction (FEC), and with FEC, the bandwidth is 125Gbps×150 / 160=117.1875Gbps. In some embodiments, the PCIe retimer uses low-latency FEC and scrambling schemes for high-speed die-to-die interconnect operations. In this structure, 150 bits of data are sent in each clock cycle of each data stream. The clock frequency may depend on the link speed. At 125Gbps, the core clock frequency is 125Gbps / (5×32)=781.25MHz. Among them, the 150 bits of data sent at one end of the link are aligned at the receiving end, that is, TX bit 0 is received as RX bit 0. The 150 bits of data in one clock cycle are called a "codeword".

[0074] The inter-die data interface operates using the same 100 MHz reference clock as the PHY. In some embodiments, this interface is configured using an APB interface with an 8-bit wide data bus. In some embodiments, this interface can be used to reduce power consumption by reducing speed. Furthermore, the number of enabled TX / RX data streams can be adjusted based on the amount of bandwidth required for communication.

[0075] Figure 8 FIG. 1 is a block diagram of an adaptation layer (AL) for an inter-die data interface in some embodiments. The adaptation layer formats the payload sent and received via the high-speed die-to-die interconnect. Figure 8 As shown, the Adaptation Layer supports the following types of payloads:

[0076] 1) Raw SERDES RX data (up to eight SERDES);

[0077] 2) Frames / packets from a link controller that supports flow control (up to eight interfaces are active);

[0078] 3) Indirect register read and write commands executed through the APB bus.

[0079] exist Figure 3 In one embodiment, the retimers 305 and 310 may be used Figure 5 The retimer data path is shown in Figure 5 In [1], data is routed via the D2D interface by the adaptation layer. In one such embodiment, to minimize latency, the original encoded data is sent via the original data interface via the D2D interface. When the input traffic is terminated by the link controller, a data frame mode may be used, which is described in further detail below.

[0080] Original data format

[0081] The eight raw SERDES RX data interfaces operate in parallel. Depending on the protocol, the eight data frame interfaces can operate in round-robin fashion or in parallel. The high-speed link is statically configured to transmit both raw SERDES RX data and data frames. Indirect register accesses can be interleaved within both traffic types.

[0082] like Figure 8 As shown, the original SERDES RX data stream receives two 32-bit data codewords from the SERDES in two consecutive receive clock cycles and writes the combined 64-bit data into the asynchronous FIFO. Within the asynchronous FIFO, the clock domain is converted from the recovered clock assigned to the data by the channel routing logic to the adaptation layer clock domain. The data read from the asynchronous FIFO is sent via a dedicated data stream on the high-speed die-to-die link. The original data from the two RX SERDES asynchronous FIFOs is combined and sent via the same dedicated data stream on the high-speed link. The adaptation layer not only provides the clock for reading from the asynchronous FIFO but also provides the clock for merging the 150-bit D2D codeword (160 bits when FEC is present) into the 25Gbd clock used by the transmitter in the D2D data stream.

[0083] The raw data format (i.e., non-frame-based protocol) is a format used to transmit raw 32-bit SERDES data in each data stream clock cycle. The non-frame-based protocol can be used to transmit raw data over the D2D interface in the retimer operating mode of the multi-chip module. The codewords of the non-frame protocol are as follows:

[0084] 149 148:0 protocol Payload

[0085] The protocol bit of the non-frame protocol is set to 1'b1. As shown in Table 1 below, the format of bits 148:0 of the payload field is as follows:

[0086]

[0087]

[0088] Table 1

[0089] The SERDES payload has a high priority, the register command has a medium priority, and messages reserved for future use have a low priority. The SERDES payload is always inserted into the user data cycle, starting with PAYLOAD0, followed by PAYLOAD1, and so on. Register commands can only be inserted when fewer than four SERDES payloads have been prepared in the data stream cycle. Register commands are inserted only into the PAYLOAD3 field. Before a new register write address command is issued, the register write address command precedes the register write data command. A register read address command or a register read data command can be inserted between a register write address command and a register write data command.

[0090] Although the above describes in detail Figure 7 and Figure 8 The specific D2D interface shown is shown, but it should be noted that other D2D interfaces can also be used to transmit PCIe traffic between circuit dies. In some embodiments, the D2D interface can be a Universal Chip Interconnect (UCIe) interface. UCIe has multiple operating modes, one of which is a FLIT-aware operating mode that executes protocols such as CXL / PCIe with a die-to-die adapter. In addition, UCIe also has a streaming transport protocol that provides a common mode for sending raw data for user-defined protocols. In combination Figure 3 In the multi-endpoint switching implementation described above, such a streaming protocol can be used to transfer data between circuit dies in the retimer operating mode. Figure 7 The specific D2D flow shown statically maps the PHY of each upstream pseudo port to the PHY of a corresponding downstream pseudo port. A similar adaptation layer is used to split the traffic sent over two independent PCIe links of the UCIe connection.

[0091] Load distribution: non-load balancing mode

[0092] Depending on the protocol, payloads can be configured to be sent over the D2D link in either load-balanced or non-load-balanced mode. All data flows utilize one or the other operating mode. The D2D link uses non-load-balanced mode when sending PCS payload data (raw SERDES data) and frame-based mode when sending non-PCS payload data. When sending non-PCS payload data in frame-based mode, load-balanced mode is used. Load-balanced mode is described in further detail below.

[0093] In raw SERDES mode, payload data from a fixed set of channels is statically configured to be sent via dedicated D2D data streams. In this case, a "logical channel" corresponds to an adaptation layer physical port, which is the PHY mapping target of the raw data multiplexing crossbar switch. To minimize multiplexing logic, a fixed logical channel-to-data stream mapping is used. In one embodiment, the mapping relationship between eight channel flows from the adaptation layer physical port and four die-to-die data streams is as follows:

[0094] Logical channels 0 to 1 are mapped to data stream 0

[0095] Logical channels 2 to 3 are mapped to data stream 1

[0096] Logical channels 4 to 5 are mapped to data stream 2

[0097] Logical channels 6 to 7 are mapped to data stream 3

[0098] This mapping also applies to non-SERDES payload data. Register commands and message payloads are statically configured to use dedicated data streams, minimizing the logic required by processing only one command per cycle. Message payloads can be configured to use a different data stream than register commands.

[0099] In custom frame class mode, similar to native SERDES mode, each channel can be statically configured to a dedicated D2D data stream in the same manner as described above. For each D2D data stream's two data frame interfaces, the load from these two interfaces is distributed across the D2D link codewords in a round-robin manner. In some embodiments, for the same data frame interface / port within the same data stream, a minimum interval can be applied between D2D link codewords. In some embodiments, this minimum interval can be four cycles.

[0100] In some embodiments, fixed time division multiplexing (TDM) time slots can be programmed. In fixed TDM mode, the transmitter sends codewords to the four ports it supports in a constant manner, such as port 0, port 1, port 2, port 3, port 0, port 1, etc. If a port has no payload to send in a certain time slot, it sends an "idle" period. In some embodiments, the number of ports in the TDM time mechanism can be adjusted programmatically. In addition, similar to the original SERDES mode, register commands and message payloads can also be statically set to use a dedicated D2D data stream to ensure that only one command needs to be processed in a cycle, thereby minimizing the required logic. Message payloads can be configured to use a different data stream than register commands.

[0101] APB master / slave interface

[0102] The D2D interface includes an APB slave interface and an APB master interface. The APB slave interface is an interface to all configuration registers of the adaptation layer, including configuration registers for setting core-to-core (T2T) read and write transactions. T2T transactions refer to indirect register read and write commands sent via the D2D link. The source of the T2T transaction is the adaptation layer in the master core, and the destination is the adaptation layer in the slave core that converts the received T2T read and write commands into APB read and write transactions. Both the APB slave interface and the master interface have a command FIFO, but only the APB master interface has a read return FIFO. The number of entries in these two types of FIFOs can be independent of each other, but in at least one embodiment, these two types of FIFOs are set to have the same size.

[0103] The APB master interface executes the received T2T read and write commands in the APB within the slave chip. For read commands, the corresponding read return data is sent back to the master chip via the D2D link. The command FIFO in the APB master interface allows for some "significant" write operations that may take time for the slave chip to complete. The firmware ensures that the command FIFO does not time out. The FIFO fill level can be read from a register, and the firmware ensures that timeouts do not occur by adding a delay between T2T write transactions or by performing a read after issuing a maximum number of consecutive T2T write transactions and then waiting for read data. The maximum allowed number is determined by the number of command FIFO entries minus one. Because commands do not overwrite each other, T2T read transactions are used to flush the command FIFO. The master chip's APB master interface is idle, meaning it does not receive T2T transactions from the slave chip. The slave chip's APB slave interface is used to access the adaptation registers, but the slave chip does not initiate any T2T transactions.

[0104] Figure 111 is a flow chart of a method 1100 in some embodiments. As shown, method 1100 includes receiving 1105, by a plurality of upstream serial data transceivers of a first circuit die of a multi-die integrated circuit module (ICM), a plurality of serial data lanes associated with a PCIe data link, and generating corresponding deserialized channel-specific data codewords. The method further includes providing 1110 the deserialized channel-specific data codewords for transmission via a set of downstream serial data transceivers within the first circuit die of the multi-die ICM, the set of downstream serial data transceivers having a PCIe data link to a first endpoint. The method further includes rerouting 1115 the deserialized channel-specific data codewords to a second circuit die of the multi-die ICM via an inter-die data interface using an inter-die adaptation layer protocol upon a failure of the PCIe data link to the first endpoint. The method further includes restoring 1120, by the second circuit die, the deserialized channel-specific data codewords from the inter-die data interface. The method further includes sending 1225 the deserialized channel-specific data codewords to the second endpoint via the second PCIe data link by the second set of downstream serial data transceivers.

Claims

1. A device, characterized in that include: a plurality of upstream pseudo ports of the first circuit die, wherein each upstream pseudo port has a connection to a respective one of the at least two root complex devices; a plurality of downstream dummy ports of the second circuit die, each downstream dummy port having a connection to a respective one of the at least two endpoints; an inter-die data interface provided between the first circuit die and the second circuit die, wherein the inter-die data interface is configured to exchange retimer data streams between the upstream pseudo port and the downstream pseudo port of the first circuit die and the second circuit die through an adaptation layer port of each circuit die according to an adaptation layer protocol; channel routing logic disposed on the first circuit die and the second circuit die, configured to map the retimer data flows between at least one set of the upstream pseudo ports and a corresponding set of downstream pseudo ports to corresponding adaptation layer ports of the first circuit die and the second circuit die according to the adaptation layer protocol; and A processor disposed in one of the first circuit die and the second circuit die is configured to configure the channel routing logic of the first circuit die and the second circuit die.

2. The device according to claim 1, wherein The inter-die data interface is a universal chip interconnection interface.

3. The device according to claim 1, wherein Each retimer data stream includes one or more lanes, wherein, for each of the one or more lanes of each retimer data stream, the lane routing logic is used to route (i) deserialized data, (ii) a receive clock signal, (iii) a reset signal, and (iv) a data enable signal to a corresponding adaptation layer port.

4. The device according to claim 1, wherein The inter-die data interface includes a plurality of die-to-die data streams, wherein the retimer data stream is provided to the plurality of die-to-die data streams according to corresponding adaptation layer ports.

5. The device according to claim 4, characterized in that Each die-to-die data flow includes a corresponding clock domain crossing buffer connected to at most one or more of the adaptation layer ports.

6. The device according to claim 4, characterized in that Each die-to-die data stream exchanges the retimer data stream via an orthogonal differential vector signaling code.

7. The device according to claim 1, wherein The system also includes a board management controller for providing a control signal to the processor via a system management bus to configure the channel routing logic.

8. The device according to claim 1, wherein The processor configures the channel routing logic of the first circuit die and the second circuit die through configuration registers.

9. The device according to claim 8, wherein The processor configures the channel routing logic of the first circuit die and the second circuit die during reset.

10. The device according to claim 1, wherein The processor configures the channel routing logic of the first circuit die and the second circuit die to map the retimer data flow between a first upstream pseudo port and a first downstream pseudo port, and to map the retimer data flow between a second upstream pseudo port and a second downstream pseudo port, wherein each retimer data flow is exchanged simultaneously over the inter-die data interface.

11. A method, characterized in that include: receiving data by a set of physical layer transceivers of a first upstream pseudo port of a first circuit die, wherein the first circuit die includes a plurality of upstream pseudo ports, each upstream pseudo port having a respective connection to a respective root complex device of a plurality of root complex devices; routing data received by the set of physical layer transceivers of the first upstream pseudo port to a corresponding set of physical layer transceivers of a first downstream pseudo port of a second circuit die via an inter-die data interface, wherein the second circuit die includes a plurality of downstream pseudo ports, each downstream pseudo port having a respective connection to a respective endpoint of a plurality of endpoints; and The channel routing logic of the first circuit die and the second circuit die is configured by a processor of one of the first circuit die and the second circuit die to map the set of physical layer transceivers of the first upstream pseudo port and the set of physical layer transceivers of the first downstream pseudo port to corresponding adaptation layer ports as part of an adaptation layer protocol.

12. The method according to claim 11, wherein The received data is routed via the inter-die data interface using a universal chiplet interconnect interface.

13. The method according to claim 11, wherein The received data includes (i) a deserialized data codeword, (ii) a receive clock signal, (iii) a reset signal, and (iv) a data enable signal.

14. The method according to claim 11, wherein The adaptation layer ports are statically mapped to a plurality of die-to-die transceivers.

15. The method according to claim 14, wherein The die-to-die transceiver routes received data through the die-to-die interface using orthogonal differential vector signaling codes.

16. The method according to claim 11, wherein Also includes: receiving data by a set of physical layer transceivers in a second upstream pseudo port of the first circuit die; routing data received by the set of physical layer transceivers in the second upstream pseudo port to a corresponding set of physical layer transceivers in the second downstream pseudo port of the second circuit die via an inter-die data interface; as well as The channel routing logic of the first circuit die and the second circuit die is configured to map the set of physical layer transceivers of the second upstream pseudo port and the set of physical layer transceivers of the second downstream pseudo port to the adaptation layer port.

17. The method according to claim 11, wherein Configuring the channel routing logic of the first circuit die and the second circuit die includes writing to configuration registers of the first circuit die and the second circuit die.

18. The method according to claim 17, wherein Writing to the configuration register of one of the first circuit die and the second circuit die is performed through a chiplet-to-chiplet interface.

19. The method according to claim 11, wherein Also includes: The processor receives a control signal, wherein the control signal is related to the configuration of the channel routing logic, and wherein the control signal is sent by a board management controller.

20. The method according to claim 11, wherein The channel routing logic of the first circuit die and the second two circuit dies are configured during reset.

Citation Information

Cited By

  • CXL switching chip of multi-core-grain architecture and cross-core-grain routing switching method thereof

    CN121309513A