System and method for transporting memory-mapped traffic between integrated circuit devices
The described chip-to-chip interface architecture addresses the inflexibility of traditional communication systems by enabling packetized and scalable memory-mapped traffic exchange between heterogeneous IC devices, improving communication efficiency and interoperability within SoC architectures.
Patent Information
- Application Number
- JP2025505553
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-08-02
- Filing Date
- 2023-05-17
- Publication Date
- 2025-09-02
AI Technical Summary
Traditional chip-to-chip communication architectures are inflexible and lack configurability for memory-mapped traffic between heterogeneous integrated circuit devices, leading to inefficiencies in communication and development cycles.
A distributed chip-to-chip interface architecture that enables packetized, scalable, and configurable communication between heterogeneous IC devices, utilizing network-on-chip circuitry and NoC-to-chip bridge circuitry to facilitate memory-mapped traffic exchange with off-chip devices, supporting plug-and-play functionality.
The solution provides a flexible and efficient method for transporting memory-mapped traffic between IC devices, enhancing communication performance and enabling interoperability among different subsystems within a system-on-chip architecture.
Smart Images

Figure 2025528765000001_ABST
Abstract
Description
[Technical Field]
[0001] Examples of the present disclosure generally relate to transporting memory-mapped traffic between integrated circuit devices. [Background technology]
[0002] A system-on-chip (SoC) combines various subsystems within a single IC device. An SoC may include a packet-switched network-on-chip (NoC) to provide intra-chip communication between the subsystems. An SoC may offer various benefits over multi-chip architectures.
[0003] On the other hand, a multi-chip architecture may be useful for substitutability and interchangeability of the various subsystems and for isolating the development cycles of the various subsystems from one another, which may be useful for allowing mixing and matching of different types and versions of the various subsystems.
[0004] Multi-chip architectures utilize a variety of architectures for chip-to-chip or chip-to-chip (C2C) communication. Traditional C2C architectures are fixed or hardened. Traditional C2C architectures are not configurable for communication between heterogeneous IC devices. Traditional C2C architectures handle memory-mapped traffic with limitations. Summary of the Invention
[0005] Techniques for transporting memory-mapped traffic between integrated circuit devices are described. One example is an integrated circuit (IC) chip that includes functional circuitry that exchanges memory-mapped traffic with off-chip devices, network-on-chip (NoC) circuitry that packetizes and depacketizes memory-mapped traffic received from and directed to the functional circuitry and routes the packetized memory-mapped traffic between the functional circuitry and the off-chip devices, and NoC-to-chip bridge (NICB) circuitry that interfaces between the NoC circuitry and the off-chip devices to exchange the packetized memory-mapped traffic with the off-chip devices over a chip-to-chip (C2C) interconnect.
[0006] Another example described herein is an IC device that includes functional circuitry that sends and receives memory-mapped data to and from off-chip devices, NoC circuitry that provides the functional circuitry with a virtual interface to other devices, and NICB circuitry that provides the NoC circuitry with a virtual interface to the off-chip devices.
[0007] Another example described herein is a method that includes generating memory-mapped traffic including a memory access request; packetizing the memory access request; routing the packetized memory access request through an on-chip packet-switched communication network at a first transmission rate; and forwarding the packetized memory access request from the on-chip packet-switched communication network to an off-chip device in a segment via a set of C2C interconnects at a second transmission rate, where the second transmission rate is greater than or equal to the first transmission rate. [Brief explanation of the drawings]
[0008] In a manner in which the above-recited features may be understood in detail, a more particular description briefly summarized above may be made by reference to exemplary implementations, some of which are illustrated in the accompanying drawings. It should be noted, however, that the accompanying drawings illustrate only typical example implementations and therefore should not be considered limiting of the scope thereof. [Figure 1] FIG. 1 is a block diagram of a system, according to an embodiment. [Figure 2] FIG. 1 is a block diagram of a portion of a system, according to an embodiment. [Figure 3] FIG. 2 is a block diagram of a network-on-chip master unit (NMU) circuit, according to an embodiment. [Figure 4] FIG. 2 is a block diagram of a network-on-chip slave unit (NSU) circuit, according to an embodiment. [Figure 5] FIG. 2 is a block diagram of a portion of a chiplet, according to an embodiment. [Figure 6] FIG. 1 is a conceptual diagram of a network-on-chip inter-chip bridge (NICB) infrastructure, according to an embodiment. [Figure 7] FIG. 1 is a conceptual diagram of a framing check architecture for memory-mapped traffic and streaming, according to an embodiment. [Figure 8] FIG. 2 is a block diagram of an integrated circuit chip to illustrate features of a transmit NICB (TNICB) circuit, according to an embodiment. [Figure 9] 9 is another block diagram of the integrated circuit chip of FIG. 8 illustrating features of the receive NICB (RNICB) circuitry, according to an embodiment. [Figure 10] FIG. 2 is a conceptual diagram of a configurable virtual channel (VC) buffer and a VC start ID register in a first configuration, according to an embodiment. [Figure 11] FIG. 10 is a conceptual diagram of a configurable VC buffer and VC start ID register in a second configuration, according to an embodiment. [Figure 12] FIG. 1 is a conceptual diagram of two VCs configured with unequal numbers of FIFO buffer entries, according to an embodiment. [Figure 13] 13 is a table illustrating the simplified read and write example of FIG. 12. [Figure 14] 1 illustrates an exemplary mapping between write requests (WR) streaming requests, according to an embodiment. [Figure 15] 1 illustrates a method for sending memory-mapped traffic to an off-chip device, according to an embodiment. [Figure 16] 1 illustrates a method for receiving and processing memory-mapped traffic from an off-chip device, according to an embodiment. [Figure 17] 10 illustrates a table of address decoder parameters, according to an embodiment.
[0009] For ease of understanding, wherever possible, identical reference numbers have been used to indicate identical elements common to the figures. It is contemplated that elements of one embodiment may be beneficially incorporated in other embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0010] Various features are described below with reference to the drawings. It should be noted that the drawings may or may not be drawn to scale, and that elements of similar structure or function are represented by similar reference numerals throughout the drawings. It should be noted that the drawings are intended only to facilitate the description of the features. They are not intended as an exhaustive description of the features or as limitations on the scope of the claims. In addition, the illustrated example need not have all the aspects or advantages shown. An aspect or advantage described in connection with a particular embodiment is not necessarily limited to that embodiment and may be implemented in any other embodiment, even if not so illustrated or explicitly described.
[0011] Embodiments herein describe a distributed chip-to-chip (C2C) interface architecture for transporting memory-mapped traffic between heterogeneous IC devices in a packetized, scalable, and configurable manner. The C2C interface architecture can serve as a virtualization interface to off-chip devices. The C2C interface architecture can be configured as a plug-and-play architecture. Examples are provided below with reference to C2C communication between an anchor IC device or chip and one or more chiplets. However, the methods and systems disclosed herein are not limited to the anchor / chiplet architecture.
[0012] Unless otherwise specified herein, the terms bandwidth, phit, and flit have the meanings given below.
[0013] A network has a width w and a transmission rate f. The bandwidth (BW) of a network is given by BW=w * The amount of data transferred in one cycle is called a physical unit or fit. The width of the network, w, is also equal to the fit size. The bandwidth of a network may be defined in fits / sec.
[0014] Messages can be broken down into segments of fixed-length entities called packets. Packets can be broken down into message flow control units or flits. A flit is a unit amount of data at which a message is transmitted at the link level. Inter-router forwarding is structured in terms of phits; in contrast, switching technologies operate in terms of flits.
[0015] 1 is a block diagram of a system 100 according to an embodiment. System 100 includes a first integrated circuit (IC) chip, shown here as anchor 102, connected to a plurality of heterogeneous IC chips, shown here as chiplets 104A-104C, via respective C2C interface circuits 118A, 118B, and 118C.
[0016] The anchor 102 may include, but is not limited to, a processor subsystem, a memory subsystem, and / or programmable / configurable circuitry. The chiplets 104A-104C may include respective dedicated functional circuitry, such as, but not limited to, an artificial intelligence engine (AIE) and / or memory.
[0017] 1, chiplet 104A includes functional circuitry 108 that sends and receives memory-mapped traffic 110 to and from off-chip devices, shown here as anchors 102. Memory-mapped traffic 110 may include memory write requests and / or memory read requests, collectively referred to herein as memory access requests.
[0018] For purposes of explanation, memory access requests are described below as originating in functional circuitry 108 and directed to functional circuitry 130 of anchor 102. Functional circuitry 130 may include, but is not limited to, a direct memory access (DMA) controller. The DMA controller can then access memory within anchor 102 and / or memory within chiplets 104B or 104C. In other embodiments, memory access requests may originate in anchor 102.
[0019] Memory-mapped traffic 110 may be formatted according to the Advanced eXtensible Interface (AXI) on-chip communication bus protocol developed by ARM of Cambridge, U.K. Memory-mapped traffic 110 may be formatted according to AXI3 or AXI4 memory-mapped modes, but memory-mapped traffic 110 is not limited to the AXI protocol.
[0020] Chiplet 104A further includes network-on-chip (NoC) circuitry 112 that provides an on-chip packet-switched communication network. The NoC circuitry 112 may include one or more NoC packet switches (NpSSs). The NoC circuitry 112 can present to the functional circuitry 108 a virtual interface to off-chip memory devices.
[0021] Anchor 102, chiplet 104B, and chiplet 104C include respective NoC circuits 114, 116, and 118 to present corresponding on-chip packet-switched communication networks.
[0022] The on-chip packet-switched communication networks are independent of one another. In other words, each packet-switched communication network is local to a respective chip (anchor or chiplet) such that addressing for each NoC terminates within the respective IC chip. Addresses and other control structures (e.g., system management identifiers (SMIDs) and / or destination identifiers (DEST-IDs)) for received transmissions can be regenerated in the respective C2C interface circuits, as described in one or more examples herein.
[0023] C2C interface circuits 118A, 118B, and 118C interface between the NoC circuitry 114 of anchor 102 and the NoC circuitry of each chiplet 104A-104C. C2C interface circuit 118A includes NoC inter-chip bridge (NICB) circuit 120, NICB circuit 122, and interconnect 124. NICB circuit 120 may present a virtual interface to anchor 102 to NoC circuit 112. NICB circuit 122 may present a virtual interface to chiplet 104A to NoC circuit 114.
[0024] C2C interface circuits 118B and 118C may be similar to or identical to C2C interface circuit 118A.
[0025] Anchor 102, chiplets 104A, 104B, and / or 104C may include multiple sets of NICB circuitry.
[0026] FIG. 2 is a block diagram of a portion of a system 100 according to an embodiment.
[0027] In the example of FIG. 2, the NoC circuitry 112 includes an NPS 202, an NoC master unit (NMU) circuitry 204, and an NoC slave unit (NSU) circuitry 206.
[0028] NMU circuitry 204 interfaces between functional circuitry 108 and NPS 202 for memory access requests originating from functional circuitry 108. Among other tasks, NMU circuitry 204 packetizes memory access requests originating from functional circuitry 108 based on the NoC Packet Protocol (NPP) of NoC circuitry 112 and depacketizes responses to the requests.
[0029] NSU circuitry 206 interfaces between functional circuitry 108 and NPS 202 for memory access requests originating from other devices, such as from anchor 102 or another chiplet. Among other tasks, NSU circuitry 206 depacketizes memory access requests originating from other devices and packetizes responses to the requests.
[0030] NoC circuitry 114 includes NPS 212, NMU circuitry 214, and NSU circuitry 216. NMU circuitry 214 interfaces between functional circuitry 130 and NPS 212 for memory access requests originating from functional circuitry 108. Among other tasks, NMU circuitry 214 packetizes memory access requests originating from functional circuitry 130 based on the NPP of NoC circuitry 114.
[0031] NSU circuitry 216 interfaces between functional circuitry 130 and NPS 212 for memory access requests that originate off-chip, such as from chiplet 104A or another chiplet. Among other tasks, NSU circuitry 216 de-packetizes memory access requests that originate off-chip.
[0032] 2, NICB circuitry 120 includes a receive NoC-to-chip bridge (RNICB) circuitry 208 and a transmit NoC-to-chip bridge (TNICB) circuitry 210. NICB circuitry 122 includes an RNICB circuitry 218 and a TNICB circuitry 220.
[0033] For a memory access request originating in functional circuit 108, NMU circuit 204 packets the access request into j-bit flits, NPS 202 communicates the j-bit flits to TNICB circuit 210 via interconnect 126, TNICB circuit 210 communicates the contents of the j-bit flits in segments to RNICB circuit 218 via interconnect 124, RNICB circuit 218 combines the segments to regenerate the j-bit flit and forwards the j-bit flit to NoC circuit 114, NSU circuit 216 depacketizes the memory access request from the j-bit flits, and functional circuit 130 processes the memory access request.
[0034] Functional circuitry 130 may similarly send responses to memory access requests to functional circuitry 108 via NSU circuitry 216 , NPS 212 , TNICB circuitry 220 , RNICB circuitry 208 , NPS 202 , and NMU circuitry 204 .
[0035] The NoC circuitry 112 may include one or more additional PSSs to provide other functional circuits of chiplet 104A with an interface to the packet-switched communication network of the NoC circuitry 112.
[0036] The NoC circuitry 114 may include one or more additional PSSs to provide other functional circuits of the anchor 102 with an interface to the packet-switched communication network of the NoC circuitry 114 .
[0037] In an embodiment, the NoC circuitry 112 and / or the NICB circuitry 120 are provided as pods for chiplets. The pods may be provided as code for inclusion in a manufacturing process or as physical devices. A pod may include one NMU, one NSU, and one NPS (e.g., a hybrid switch). In the NMU, address decoding may be established using a reduced set of remap registers.
[0038] TNICB circuitry 210 and / or TNICB circuitry 220 may perform one or more of the following operations. Flow control; - Signal mapping from NPP to interconnect 124; Error Correction Code (ECC) generation.
[0039] RNICB circuitry 208 and / or RNICB circuitry 218 may perform one or more of the following operations. Flow control; - Signal mapping from interconnect 124 to NPP; ECC checking and poison bit generation; ECC generation for NPS (e.g., to comply with NoC requirements); SMID remapping; Address translation; Generate destination ID; Read interleaving and reordering; Write tracking.
[0040] 3 is a block diagram of an NMU circuit 300, according to an embodiment. NMU circuit 204 and / or NMU circuit 214 may include one or more features of NMU circuit 300.
[0041] 4 is a block diagram of an NSU circuit 400, according to an embodiment. NSU circuit 206 and / or NSU circuit 216 may include one or more features of NSU circuit 400.
[0042] FIG. 5 is a block diagram of a portion of a chiplet 104A, according to an embodiment.
[0043] In the example of FIG. 5, the interconnect 126 includes a first set of j-interconnects 502 that provide j-bit flits from the RNICB circuitry 208 to the NPS 550 of the NoC circuitry 112 .
[0044] The interconnect 126 further includes a second set of j-interconnects 504 that provide j-bit flits from the NPS 552 of the NoC circuit 112 to the TNICB circuit 210 .
[0045] Further, in the example of FIG. 5, interconnect 124 includes a first set of interconnects 508 that provide transmissions to RNICB circuitry 208 .
[0046] The interconnect 124 further includes a second set of interconnects 512 transmitting from the TNICB circuit 210 .
[0047] The remainder of interconnect 124 may or may not be used by other NICB circuits (eg, NICB circuits 554, 556, and 558), such as for streaming data.
[0048] As a non-limiting example, in embodiments The interconnect 502 may include approximately 288 interconnects to provide 288-bit flits from the RNICB circuitry 208 to the NoC circuitry 112 . The interconnect 504 includes approximately 288 interconnects for providing 288-bit flits from the NoC circuitry 112 to the RNICB circuitry 208 . The interconnects 508 may include approximately 80 interconnects, 72 of which may be used to receive 288-bit flits in four 72-bit segments at the RNICB circuit 208. The interconnects 512 may include approximately 80 interconnects, 72 of which may be used to transmit 288-bit flits in four 72-bit segments from the TNICB circuit 210.
[0049] The remaining interconnects of interconnects 508 and 512 may be reserved for mapping, parity, etc.
[0050] In an embodiment, the NoC circuitry and the NICB circuitry 120 communicate over the interconnect 126 at a first transmission rate, and the NICB circuitry 120 and the NICB circuitry 122 communicate over the interconnect 124 at a second transmission rate. The first and second transmission rates may be set so that the bandwidth of the interconnect 126 measured over time is equal to the bandwidth of the interconnect 124. For example, in FIG. 5 , the TNICB circuitry 210 may receive a 288-bit flit from the NoC circuitry over 288 interconnects 508 at a first transmission rate of 1 GHz and may transmit the 288-bit flit in four 72-bit segments over the interconnect 502 at a second transmission rate of 4 GHz.
[0051] 5, interconnect 508 is shown as a pair of k-bit receiving data words DW1 and DW2, shown here as 520 and 522, respectively. Similarly, interconnect 512 is shown as a pair of transmitting data words DW1 and DW2, shown here as 524 and 526, respectively. Continuing with the previous example, each data word can represent approximately 40 interconnects, 36 of which can be used to transmit memory-mapped data. The remaining interconnects can be used or reserved for mapping, parity, etc.
[0052] 5, NICB circuitry 120 includes a protocol processing unit (PHU) 540 that includes receiver circuits 542 and 544 that receive data words (DWs) 520 and 522, respectively. PHU 540 further includes transmit circuits 546 and 548 that transmit DWs 524 and 526, respectively.
[0053] 5, NICB circuits 120, 554, 556, and 558 provide a total of eight receive DWs and eight transmit DWs. For example, at 256 bits / DW, NICB circuits 120, 554, 556, and 558 provide 2 Terabits per second (Tbps) of bandwidth in each direction. In another example, 16 DWs are provided in each direction for a bandwidth of 16 x 256 = 4 Tbps in each direction.
[0054] NICB signal mapping NICB signal mapping is described below with reference to FIG.
[0055] 6 is a conceptual diagram of a NICB infrastructure 600, according to an embodiment. The NICB infrastructure 600 includes a virtual channel 610, an arbiter circuit 612, and an NPP2DW circuit 614 that maps NPP flits to DW0 606 and DW1 608.
[0056] In an embodiment, when a flit 602 is mapped to DW0 606 and DW1 608, the signal mapping is 1:1 in that all bits of the flit 602 are mapped directly to DW0 606 and DW1 608 (e.g., to / from the least significant bits (LSBs) of DWs 606 and 608). The signals may retain their semantic definitions. The C2C protocol layer may use a subset of signals internally, examples of which are provided in the table below.
[0057] [Table 1]
[0058] When mapping flits 602 across DWs 606 and 608, the circuits associated with DWs 606 and 608 may operate in lockstep mode. Arbitration may be performed with respect to one of DWs 606 and 608, and the payload may be carried by the other of DWs 606 and 608.
[0059] On the receiving side, a framing check can be performed for alignment purposes. The framing check can be performed using circuitry also used for flow control streaming mode. An example is provided further below with reference to FIG. 7.
[0060] In addition to other signals, the NICB circuit may use some bits (e.g., 4 bits) of DW0 606 or DW1 to map virtual channel (VC) credits. VC credits may be transported over the C2C interconnect by piggybacking onto adjacent (pairing) DWs. This may be similar to streaming transmission.
[0061] 7 is a conceptual diagram of a framing check architecture 700 for memory-mapped traffic and streaming, according to an embodiment. In addition to normal signaling, the NICB circuitry may use four bits of the DW to map VC credits. As with streaming, credits may be transported to the other side by piggybacking on a neighboring (paired) DW server.
[0062] Transmitter (TNICB) An exemplary TNICB feature is described below with reference to FIG.
[0063] 8 is a block diagram of an IC chip 800. The IC chip 800 may represent an anchor 102 or a chiplet 104A. The IC chip 800 includes a NICB circuit 802 that interfaces between an NoC circuit 804 and an off-chip device 801. The NICB circuit 802 includes an RNICB circuit 806 and a TNICB circuit 808.
[0064] The TNICB circuit 808 includes VC buffers 810 that may be operated as first-in, first-out (FIFO) buffers. Each of the VC buffers 810 buffers may be associated with a respective sending virtual channel (VC).
[0065] Flow control is achieved by sending per-VC credits 814 back to the NoC circuitry 804 when there is a pop from each one of the VC buffers 810. The TNICB circuitry 808 may be initialized with credits to match the corresponding RNICB FIFO size.
[0066] The TNICB circuit 808 further includes an arbiter circuit 812. When an incoming transfer is received from the NoC circuit 804, the arbiter circuit 812 performs a VC decoding / arbitration operation to assign the incoming transfer to one of the VC buffers 810 of the assigned VC. The arbitration may be similar to a deficit round-robin process in the NoC circuit 804. Other types of arbitration, such as most recently used arbitration, may also be used.
[0067] The TNICB circuit 808 further includes a mapping circuit, shown here as an NPP2DW circuit 816, that performs signal mapping.
[0068] Receiver (RNICB) Exemplary RNICB features are described below with reference to FIG.
[0069] FIG. 9 is another block diagram of an IC chip 800 illustrating exemplary features of an RNICB circuit 806, according to an embodiment.
[0070] The RNICB circuit 806 includes VC buffers 910 that may be operated as first-in, first-out (FIFO) buffers. Each of the VC buffers 910 may be associated with a respective receiving VC.
[0071] The RNICB circuitry 806 further includes an ECC check circuitry 902 that generates a poison bit 904 that indicates whether the received data is corrupted.
[0072] The RNICB circuitry 806 further includes a VC decode circuit 906 that decodes the VC from transfers received from the off-chip device 801 and loads the transfers into corresponding ones of the VC buffers 910 .
[0073] The RNICB circuit 806 further includes an arbiter circuit 912, which may include a deficit round-robin arbiter circuit. The arbiter circuit 912 may receive additional input from a read reorder buffer (RAOB). In an embodiment, a RD request can win arbitration only when the RROB has space for at least one chop, and a WR request / data can win arbitration only when the WR buffer has space to hold at least one chop.
[0074] Address map translation and destination ID remapping Address Decoder In an embodiment, the incoming address is dynamically translated to the outgoing address.
[0075] In an embodiment, the RNICB generates the destination ID and SMID.
[0076] In an embodiment, the RNICB may change the priority of a transaction.
[0077] In an embodiment, the RNICB includes a programmable address decoder. Figure 17 shows a table 1700 of exemplary address decoder parameters, according to an embodiment.
[0078] The programmable registers of the RNICB may be similar to the remap registers of the NMU, but with differentiated widths. An exemplary embodiment includes 32 remap registers with 64KB granularity, 32 remap registers with 1MB granularity, 16 remap registers with 64MB granularity, and 16 remap registers with 64MB granularity.
[0079] The address decoder may include a MASK register and an ADD register. The source address can be calculated as follows: (address_hit_remap_register_X) Where Output_address=Input_address & MASK_X+ADD_X, where X indicates the remap register in which the input address found a hit.
[0080] SMID Mapping In an embodiment, the NoC uses the SMID as a unique identifier for each master. Programmable registers can be used for SMID mapping (e.g., up to 64 SMID mappings, or more). Each mapping can take an incoming SMID and map it to an outgoing SMID. If the IDs do not match, the incoming SMID can be used as the outgoing SMID.
[0081] priority The SMID lookup register may be used to set priority bits (e.g., AXPROT bits), which may be retained if the SMID does not match the SMIDs in the (e.g., 64) registers.
[0082] RD Interleaving and ROOB The RNICB circuitry 806 may support read (RD) interleaving read reorder buffering (RAOB), as described below.
[0083] In an embodiment, an incoming DW from an off-chip device 801 has an associated VC identifier (VC-ID). In an embodiment, the ROOB supports up to eight RD VCs. Each VC may have a configurable register to set the maximum number of outstanding requests (e.g., in bytes). The NICB circuit 802 may throttle requests on a VC if the maximum number of outstanding requests is reached.
[0084] In an embodiment, a mapping is created to map an incoming VC to one of the read reordering buffers (RROBs). For example, VC0, VC1, and VC2 may be mapped to RD_REQ. VC0, VC1, and VC2 may also be mapped to RD_RESP in the opposite direction (TNICB).
[0085] For requests appearing on VC0, the NICB circuitry 802 may chop the transactions, interleave them as necessary, and store the RROB information in the RROB. On the responder side, for responses appearing on VC0, the RAOB is updated.
[0086] An entry in the ROOB may include a key generated by concatenating identifiers received in the DW (e.g., VC-ID, SRC-ID, and TAG-ID) and a locally generated identifier (L-TAG-ID). The ordering of responses for a given concatenated identifier {VC-ID, SRC_ID, TAG-ID, L-TAG-ID} may be preserved. The NICB circuit 802 may use the {VC-ID, SRC-ID, TAG-ID} tuple in a manner similar to how an NMU uses identifiers.
[0087] In an embodiment, if (local_tag_id, VC, SRC-ID, TAG-ID) is occupied in the RROB, the NICB circuit 802 may add (local_tag_id, VC, SRC-ID, TAG-ID) to the candidates.
[0088] The NICB circuitry 802 may arbitrate among the candidates using, for example, an NPS arbitration circuit, and may send responses and update status over the C2C interface (i.e., bookkeeping).
[0089] Write Tracking The NICB circuit 802 may apply write tracking to multiple VCs. In an embodiment, the NICB circuit 802 applies tracking to up to two VCs. The NICB circuit 802 may apply write tracking in a manner similar to the write tracking applied by the NMU circuit.
[0090] The NICB circuit 802 can index into the writer tracker using the {SRC_ID,TAG_ID} tuple. PROB can be performed similarly.
[0091] RNIC Buffering In an embodiment, VC buffers such as VC buffer 910 have a fixed depth (i.e., a fixed number of entries) sufficient to cover round-trip latency. VC buffers 910 may each contain, for example, approximately 25 FIFO entries. Over 8 VCs, this corresponds to approximately 200 entries.
[0092] The VC buffer 910 may include one or more optimizations that may be implemented individually and / or in combination with each other.
[0093] In a first optimization, the VC buffer 910 is implemented as a static random access memory (SRAM) cell.
[0094] In another optimization, a buffer pooling technique is utilized to dynamically control the depth of each of the VC buffers 910. The buffer pool allows the allocation of VCs to buffer entries to be dynamically controlled. The buffer pool technique can utilize a programmable register that specifies the next ID. In this way, each VC is assigned a virtual FIFO whose depth can be changed at runtime. Such buffer pooling is referred to herein as dynamically configurable NICB buffering.
[0095] In a further optimization, another set of programmable registers can be used to dynamically specify the starting ID for each VC. This can be useful to provide fine-grained redundancy (i.e., by moving pointers around). This can be useful to bypass relatively individual FIFO entries rather than an entire block of memory cells when one of the memory cells fails.
[0096] 9, IC chip 800 includes controller circuitry for dynamically configuring RNICB buffering and dynamically assigning start IDs. Dynamically configurable RNICB buffering and dynamic assignment of start IDs are further described below with reference to FIGS. 10-13.
[0097] FIG. 10 is a conceptual diagram of a configurable VC buffer and VC Start ID register in a first configuration, according to an embodiment.
[0098] FIG. 11 is a conceptual diagram of a configurable VC buffer and VC Start ID register in a second configuration, according to an embodiment.
[0099] FIG. 12 is a conceptual diagram of two VCs configured with unequal numbers of FIFO buffer entries, according to an embodiment.
[0100] FIG. 13 is a table 1300 illustrating the simplified RD and WR example of FIG.
[0101] 10 shows FIFO buffer entries 1002, next entry IDs 1006, entry IDs 1004, and VC start entries 1010. Each of the FIFO buffer entries 1002 is associated with a respective one of the entry IDs 1004 and a respective one of the next entry IDs 1006.
[0102] The FIFO buffer entries 1002 hold the payload and may be implemented in SRAM. The next entry ID 1006 may be stored in the respective FIFO buffer entries 1002 or in a register. The VC start entry 1010 may be stored in a respective register.
[0103] 10, there are eight VC start entries 1010, each associated with a respective one of eight VCs, shown here as VC0 through VC7. Each VC start entry 1010 indicates the start entry ID for the respective VC.
[0104] In this example, VC0 starts at VC entry 0, VC1 starts at VC entry 5, and VC2 starts at VC entry 10. Thus, VC0 includes five FIFO buffer entries, identified here as entry IDs 0, 1, 2, 3, and 4. VC1 also includes five FIFO buffer entries, identified here as entry IDs 5, 6, 7, 8, and 9.
[0105] Next Entry ID 1006 acts as a pointer to the next available FIFO entry for the respective VC. Next Entry ID 1006 essentially creates a circular linked list for each VC. For example, Next Entry ID 1006 for VC0 cycles through entry IDs 0, 1, 2, 3, 4, then wraps around to 0. Next Entry ID 1006 for VC1 cycles through entry IDs 5, 6, 7, 8, 9, then wraps around to 5.
[0106] The architecture of FIG. 10 allows for programmatic reconfiguration of the VC buffer depth in that the Next Entry ID 1006 and VC Start Entry registers can be reprogrammed to increase or decrease the VC depth.
[0107] 10 also allows for changing the VC start entry 1010 without necessarily changing the depth of any of the VCs (i.e., by changing the next entry ID to maintain the existing depth), which can be useful for bypassing a relatively small set or block of memory cells when a fault is detected within the block.
[0108] In Figure 11, relative to Figure 10, the VC start entry for VC1 has been changed from 5 to 4 in VC start entries 1010. The other VC start entries remain unchanged. In this example, VC0 starts at VC entry 0, VC1 starts at VC entry 4, and VC2 starts at VC entry 10. VC0 contains four FIFO buffer entries, identified here as entry IDs 0, 1, 2, and 3. This is one less entry than in Figure 10. VC1 contains six FIFO buffer entries, identified here as entry IDs 4, 5, 6, 7, 8, and 9. This is one more entry than in Figure 10. The next entry IDs 1006 for VC0 and VC1 are updated accordingly.
[0109] 12, a first VC 1202 (shown here as VC0) has two FIFO buffer entries 1204 and 1206 (shown here as entries 0 and 1). A second VC 1208, shown here as VC1, has three FIFO buffer entries 1210, 1214, and 1216, shown here as entries 2, 3, and 4. The FIFO buffer entries of VC0 and VC1 are used in a circular manner, as shown in the example of FIG.
[0110] In Figure 13, all FIFO buffer entries are initially empty. With each WR / RD, the RD and WR pointers (i.e., RD Next Entry ID 1302 and WR Next Entry ID 1304) change as shown, generating the FULL / EMPTY signals as shown.
[0111] With the buffer pooling methods disclosed herein, the NICB circuitry can operate with fewer overall buffer entries than architectures that use a fixed number of buffer entries, since typically only a few VCs are active at any given time. By tailoring the VC buffers for a particular use case, buffer entries can be multiplexed across multiple VCs.
[0112] Streaming Support For streaming data, the RNICB circuit can support mapping of incoming identifiers (e.g., TDEST-ID) to outgoing destination identifiers (destination IDs). The NoC circuit can be configured or programmed to allow one or more destination IDs to be routed to the TNICB circuit. Thus, each destination ID routed to the TNICB circuit can be mapped to a different destination ID in the RNICB circuit.
[0113] C2C / NICB compression mode In an embodiment, the NICB circuitry operates in a compressed mode in which packetized memory-mapped traffic is mapped to the interconnects 124 in a less than 1:1 manner (eg, k interconnects 124).
[0114] In an embodiment, the NICB circuitry is configurable in each of a plurality of modes. The NICB circuitry may be configurable at boot time in one of a plurality of selectable modes. Exemplary modes are described below with reference to the NICB circuitry 120. However, the NICB circuitry 120 is not limited to the following examples.
[0115] In a first mode, also referred to herein as full mode or uncompressed mode, the NICB circuitry 120 maps packetized memory-mapped traffic 1:1 to the interconnects 124, as described in the example above. In Figure 5, for example, the NICB circuitry 120 maps packetized memory-mapped traffic to DWs 520 and 522 (i.e., two sets of k interconnects 124).
[0116] In a second mode, referred to herein as a compressed mode or high data word (DW) performance mode, the NICB circuitry 120 maps packetized memory-mapped traffic to fewer interconnects 124 than in the high-performance mode (i.e., in a less than 1:1 manner). In Figure 5, for example, the NICB circuitry 120 may map packetized memory-mapped traffic to DWs 520 (i.e., a single set of k interconnects 124). In the example of Figure 5, the NICB circuitry may pack 256-bit flits into one DW.
[0117] The NICB circuitry 122 of the anchor 102 may be configured similarly to a NICB.
[0118] The compressed mode can provide an efficient mapping of NPP traffic from interconnect 126 to interconnect 124 with relatively little additional latency and computation.
[0119] In addition to the above example, in FIG. 5, the compression mode is as follows: The RNICB circuit 208 may receive 288-bit flits in eight 36-bit segments on the DW 520. The TNICB circuit 210 may transmit 288-bit flits in eight 36-bit segments on the DW 524.
[0120] As further described above, in an embodiment, the NoC circuitry and the NICB circuitry 120 communicate over the interconnect 126 at a first transmission rate, and the NICB circuitry 120 and the NICB circuitry 122 communicate over the interconnect 124 at a second transmission rate. The first and second transmission rates may be set such that the bandwidth of the interconnect 126 measured over time is equal to the bandwidth of the interconnect 124. For example, in FIG. 5 , in compressed mode, the TNICB circuitry 210 may receive 288-bit flits from the NoC circuitry over the 288 interconnects 508 at a first transmission rate of 1 GHz and may transmit the 288-bit flits in eight 36-bit segments over the interconnect 502 at a second transmission rate of 8 GHz.
[0121] A high performance mode may be useful for maximizing the performance of a single interface, perhaps at the expense of reducing the overall bandwidth of the C2C interface, while a high DW performance mode may be useful for maximizing DW utilization, perhaps at the expense of lower performance for each memory-mapped interface.
[0122] Exemplary mapping semantics are described below.
[0123] For a write request (WR REQ) originating from the functional circuit 108, the TNICB circuit 210 may separate the WR REQ from the write data (WR DATA) and send the WR REQ and WR DATA as separate flits.
[0124] In the case of a WR REQ flit, the TNICB circuit 210 may save the TAG-ID of the write request for association with the write response, calculate the ECC on the WR REQ flit, and send the WR REQ flit to the RNICB circuit 218 via the interconnect 124. Upon receiving the WR REQ flit, the RNICB circuit 218 stores the VC-ID and Tag-ID and performs the WR processing.
[0125] For a WR data flit, the TNICB circuit 210 may calculate a strobe (STRB) and ECC to include in the WR data (e.g., 256 bits) and send the WR data flit to the RNICB circuit 218 via the interconnect 124.
[0126] In an embodiment, the TNICB circuit 210 attempts to encode the STRB into a predetermined number of bits (e.g., 10 bits, start-end). If the encoding is successful, the TNICB circuit 210 transmits the encoded STRB, ECC, and WR data in a single WR data flit over the interconnect 124. If the encoding is unsuccessful (i.e., the STRB is encoded with more than the predetermined number of bits), the TNICB circuit 210 transmits the encoded STRB, ECC, and WR data in multiple WR data flits, along with a flag included in at least a first WR data flit of the multiple WR data flits.
[0127] For example, the TNICB circuit 210 may set the VC-ID of the first WR data flit to a predetermined or reserved value to indicate that it is the first of multiple WR data flits. The TNICB circuit 210 may include as many bits of WR data as will fit in the first WR data flit and place the remaining bits of WR data in subsequent WR data flits.
[0128] In the anchor 102, the RNICB circuit 218 detects a predetermined or stored VC-ID in the first WR data flit, stores the WR data of the first WR data flit, and waits for subsequent WR data flits. When the RNICB circuit 218 receives the subsequent WR data flits, the RNICB circuit 218 concatenates the WR data flits, performs WR data processing, and forwards the concatenated WR data flits to the NoC circuit 114 via the interconnect 132.
[0129] For a write request involving streaming data in compressed mode (STRM data), the TNICB circuit 210 may attempt to encode the handshaking signaling (e.g., TKEEP) into a predetermined number of bits (e.g., 6 bits). If unsuccessful, the TNICB circuit 210 may transmit the streaming data over the interconnect 124 in multiple flits, as described above with respect to WR data. The TNICB circuit 210 may, for example, set the VC-ID of the first streaming data flit to a predetermined or reserved value to indicate that the current word is the first set of bits of streaming data (e.g., the first 288 bits) and provide the remaining bits of streaming data to subsequent streaming data flits.
[0130] In the anchor 102, the RNICB circuit 218 detects the predetermined or stored VC-ID in the first streaming data flit, stores the data of the first streaming data flit, and waits for the subsequent streaming data flit. When the RNICB circuit 218 receives the subsequent streaming data flit, the RNICB circuit 218 concatenates the streaming data flits, performs STRM processing, and forwards the concatenated streaming data flit to the NoC circuit 114 via the interconnect 132.
[0131] The TNICB circuit 210 may remove one or more predetermined bits (e.g., the DST and DST parity bits) from the streaming data flit, recalculate the ECC based on the remaining bits, and send the flit to the RNICB circuit 218 over the interconnect 124 without a STRM traffic type indication. In this example, the RNICB circuit 218 uses the VC mapping to determine the STRM traffic type.
[0132] For STRM data, the C2C interface circuit 118A may support a predetermined number of handshake bits (eg, 4 bits of TDEST and 3 bits of TID).
[0133] The TNICB circuit 210 may encode one or more handshake signals (e.g., VALID) into one or more VC numbers (e.g., VC numbers 9-14). This may be useful to assist in deriving the STRM traffic type. In this example, the RNICB circuit 218 may map the handshake signals back to the corresponding WR VCs, as shown in FIG. 14.
[0134] Figure 14 shows an example mapping 1400 between a WR 1402 and a STRM 1406, according to an embodiment. In the example of Figure 14, the WR 1402 is configured on VCs 3 and 6. The corresponding VCs for the STRM 1406 are 9 and 10, where 9 maps back to 3 and 10 maps back to 6.
[0135] When responding to a write request involving streaming data in compressed mode, the TNICB circuit 220 may look up the source identifier (SRC ID) and recreate the DW flit. The TNICB circuit 220 may send a write response (WR RESP) only for the last chop. The RNICB circuit 208 may send the flit to the NoC circuit 112.
[0136] Read request (RD REQ) In an embodiment, for a read request (RD REQ) originating at functional circuit 108, TNICB circuit 210 saves the incoming identifier (e.g., Tag-ID), calculates a local key, such as by appending a source identifier (e.g., SRC-ID) and the least significant bit of the incoming identifier, and stores the source identifier indexed by the incoming identifier. TNICB circuit 210 then creates a read request using the local key as the new tag ID, recalculates the ECC, and sends the flit over interconnect 124.
[0137] In the anchor 102, the RNICB circuit 218 performs the RD REQ processing.
[0138] TNICB circuitry 220 receives a response flit (e.g., an RD-RESP flit) from functional circuitry 130 via NoC circuitry 114. In an embodiment, TNICB circuitry 220 strips off or removes the source identifier (e.g., an SRC-ID) in the RD-RESP flit, recalculates the ECC, and transmits the flit to chiplet 104A via interconnect 124.
[0139] In chiplet 104A, RNICB circuitry 208 receives the RD-RESP flit, uses the incoming identifier (e.g., TAG-ID) to recover the original Tag-ID and SRC-ID, updates the RD-RESP flit, recalculates the ECC, and transmits the RD-RESP flit to the NoC circuitry via interconnect 126.
[0140] Credit and valid. The NICB circuit can use credit and valid signals to control the loading of the VC buffer. In an embodiment, only one credit signal is transmitted in a given cycle. In this example, the credit and valid signals can each be encoded into a predetermined number of allocated bits (e.g., four allocated bits each).
[0141] In compressed mode, the NICB circuit may map flit types to respective VCs. The NICB circuit may also send a WR REQ flit followed by all data flits corresponding to the WR REQ flit. In an embodiment, the NICB circuit does not perform interleaving within a VC. In another embodiment, the NICB circuit performs interleaving within a VC. In an embodiment, the NICB circuit performs one or more of the above features even in full mode or uncompressed mode.
[0142] Additional exemplary embodiments of the uncompressed / full mode and the compressed mode are provided below.
[0143] In an embodiment, in full mode, a j-bit flit is mapped to two DWs (i.e., a 1:1 mapping), and in compressed mode, a j-bit flit is mapped to one DW (i.e., a less than 1:1 mapping).
[0144] In an embodiment, in full mode, a WR flit occupies two DWs and data can accompany the WR in the same flit. In compressed mode, a WR request and WR data can be sent in consecutive flits.
[0145] In an embodiment, in full mode, valid bits (e.g., 8 bits) and credit bits (e.g., 8 bits) may be used for each flit. In compressed mode, the video bits and credit bits may be encoded (e.g., as bits 1 through 15). In an embodiment, a valid value of 0 indicates "invalid" and a credit value of 0 indicates "no credit." A VC-ID value (e.g., VC-15) may be reserved to indicate a partial transfer. A set of VC-IDs (e.g., VC-IDs 9 through 14) may be reserved for STRM mapping.
[0146] In an embodiment, in full mode, WR data STRBs are not encoded. In compressed mode, WR data STRBs may be encoded (e.g., to 10 bits). If encoding is unsuccessful, the data flit may be transmitted in two consecutive flits.
[0147] In embodiments, the STRM may incorporate one or more of the aforementioned features.
[0148] NICB-LITE In an embodiment, the NICB circuit is configured as an “optical” circuit (NICB-LITE) that utilizes a reduced number of VCs (e.g., 4 VCs vs. 8 VCs). The NICB circuit may be manufactured as a NICB-LITE circuit or may be reconfigurable between full VC requirements and a reduced number of VCs. In an embodiment, the NICB-LITE circuit operates exclusively in compressed mode. Alternatively, the NICB-LITE circuit may operate in full mode. The NICB-LITE circuit may include or utilize one or more features described above with respect to compressed mode and / or full mode. The NICB-LITE circuit may be useful for chiplets / applications that do not necessarily require full VC requirements. In an embodiment, the NICB-LITE circuit utilizes a reduced-size RROB (e.g., 16 KB vs. 8 KB).
[0149] 15 illustrates a method 1500 for transmitting memory-mapped traffic to an off-chip device, according to an embodiment. Method 1500 is described below with reference to chiplet 104A. However, method 1500 is not limited to the example of chiplet 104A.
[0150] At 1502, functional circuitry 108 generates memory-mapped traffic including a memory access request.
[0151] At 1504, the NMU circuit 204 packetizes the memory access request.
[0152] At 1506, the NPS 202 routes the packetized memory access request to the TNICB circuit 210 at the first transmission rate.
[0153] At 1508, the TNICB circuitry 210 forwards the packetized memory access request in the segment from the NPS 202 to the anchor 102 over the first set of interconnects 124 at a second transmission rate, where the second transmission rate is greater than the first transmission rate. The TNICB circuitry 210 may utilize one or more features described in other examples herein.
[0154] At 1510, the RNICB circuitry 208 receives responses in segments from the anchors 102 at a second transmission rate over the second set of interconnects 124 and combines the segments into flits. The RNICB circuitry 208 may utilize one or more features described in other examples herein.
[0155] At 1512, the NPS 202 routes the flit to the NMU circuitry 204 at the first transmission rate.
[0156] At 1514, the NMU circuit 204 de-packetizes the flit.
[0157] At 1516, the functional circuitry 108 processes the de-packetized flits.
[0158] 16 illustrates a method 1600 for receiving and processing memory-mapped traffic from an off-chip device, according to an embodiment. Method 1600 is described below with reference to anchor 102. However, method 1600 is not limited to the anchor 102 example.
[0159] At 1602, the RNICB circuitry 218 receives packetized memory access requests in segments from chiplet 104A at a first transmission rate over the first set of interconnects 124 and combines the received segments into flits. The RNICB circuitry 218 may utilize one or more features described in other examples herein.
[0160] At 1604, the NPS 212 routes the flit to the NSU circuitry 216 at a second transmission rate, where the first transmission rate is greater than the second transmission rate.
[0161] At 1606, the NSU circuitry 216 de-packetizes the flits and forwards the de-packetized flits to the functional circuitry 130 for processing.
[0162] At 1608 , the NSU circuitry 216 packetizes the response from the functional circuitry 130 .
[0163] At 1610, the NPS 212 routes the packetized response to the TNICB circuit 220 at the second transmission rate.
[0164] At 1612, the TNICB circuit 220 forwards the packetized response in segments to chiplet 104A over the second set of C2C interconnects at the first transmission rate. The TNICB circuit 220 may utilize one or more features described in other examples herein.
[0165] The techniques described above can be illustrated without limitation in the non-limiting examples provided below.
[0166] Example 1: An integrated circuit (IC) chip comprising: functional circuits configured to exchange memory-mapped traffic with off-chip devices; NoC (network-on-chip) circuitry configured to packetize and depacketize memory-mapped traffic received from and directed to the functional circuits and route the packetized memory-mapped traffic between the functional circuits and the off-chip devices; and NoC-to-chip bridge (NICB) circuitry configured to interface between the NoC circuitry and the off-chip devices, including exchanging the packetized memory-mapped traffic with the off-chip devices via a chip-to-chip (C2C) interconnect.
[0167] The IC chip of embodiment 1, wherein the NICB circuitry is further configured to perform address map translation and destination identifier remapping based on a system management identifier (SMID) associated with the packetized memory-mapped traffic.
[0168] Example 3: The IC chip of Example 1, wherein the NoC circuit is further configured to route the packetized memory access requests to the NICB circuit at a first transmission rate, and the NICB circuit is further configured to transmit the packetized memory access requests in segments at a second transmission rate over the C2C interconnect in a 1:1 mapping, wherein the second transmission rate is greater than the first transmission rate.
[0169] Example 4: The IC chip of example 1, wherein the NoC circuit is further configured to route packetized memory access requests to the NICB circuit at a first transmission rate, and the NICB circuit is further configured to transmit the packetized memory accesses in segments over the C2C interconnect with a less than 1:1 mapping at a second transmission rate, the second transmission rate being greater than the first transmission rate.
[0170] Example 5: The IC chip of example 4, wherein the NICB circuitry is further configured to: map a command portion of the packetized memory access request to a first flit, map a data portion of the packetized memory access request to one or more additional flits, and transmit the first flit and the one or more additional flits in segments over the C2C interconnect at a second transmission rate.
[0171] Example 6: The IC chip of Example 5, wherein the packetized memory access request includes a write request and data, and wherein the NICB circuitry is further configured to: place the write request in a first flit, calculate an error code based on the first flit, calculate a strobe based on the data, encode the strobe, place the data, the error code, and the encoded strobe in one or more additional flits, and transmit the first flit and the one or more additional flits in segments over the C2C interconnect.
[0172] Example 7: The IC chip of Example 6, wherein the NICB circuit is further configured to arrange the data, the error code, and the coded strobe into a single additional flit if the length of the coded strobe is less than a threshold, arrange the data, the error code, and the coded strobe across multiple additional flits if the length of the coded strobe exceeds the threshold, and arrange a predetermined identifier in a first additional flit of the additional flits to indicate that the first additional flit is one of the multiple additional flits of the write request.
[0173] Example 8: The IC chip of Example 1, wherein the NICB is further configured to receive a packetized memory access request from the NoC circuit, load the packetized memory access request into a first virtual channel (VC) buffer of a plurality of virtual channel (VC) buffers, map contents of the first VC buffer to a first set of segments, include an identifier with each segment of the first set of segments, and transmit the first set of segments to an off-chip device via the C2C interconnect.
[0174] Example 9: The IC chip of Example 8, wherein the NICB circuit is further configured to receive responses to memory access requests from off-chip devices as a second set of segments, load the second set of segments into a second one of the VC buffers based on identifiers included within the segments of the second set of segments, concatenate the second set of segments as flits, and forward the flits to the NoC circuit, wherein the NoC circuit is further configured to route the flits to functional circuits and de-packetize the flits.
[0175] Example 10: The IC chip of example 1, wherein the NICB circuitry is further configured to receive the first and second flits over the C2C interconnect and concatenate the first and second flits based on determining that the first flit includes a predetermined virtual channel identifier.
[0176] Example 11: The IC chip of Example 1, wherein the packetized memory access request includes a read request, and the NICB circuitry is further configured to receive a response to the read request from an off-chip device in a segment, load the segment into an entry of each of a first virtual channel (VC) buffer of a plurality of VC buffers based on an identifier included in the segment, reorder and concatenate the segments to provide a flit, and communicate the flit to the NoC circuitry.
[0177] Example 12: The IC chip of Example 1, wherein the NICB circuit is configurable in each of a full mode and a compressed mode, the NoC circuit is further configured to route packetized memory access requests to the NICB circuit at a first transmission rate, the NICB circuit is configured to transmit the packetized memory access requests in segments of a first length at a second transmission rate with a 1:1 mapping over the C2C interconnect in the full mode, and the NICB circuit is configured to transmit the packetized memory access requests in segments of the second length at a third transmission rate with a mapping less than 1:1 over the C2C interconnect in the compressed mode.
[0178] Example 13: The IC chip of Example 12, wherein the third transmission rate is greater than the second transmission rate, the second transmission rate is greater than the first transmission rate, and the first length is greater than the second length.
[0179] Example 14: The IC chip of example 1, further comprising: a controller circuit configured to dynamically configure a depth of each of the plurality of VC buffers and a starting location of the VC buffer.
[0180] Example 15: The IC chip of Example 1, wherein the NICB circuitry is further configured to map the streaming data across the first and second flits, include a predetermined identifier in the first flit to indicate that the second flit follows, and transmit the first flit and the second flit to an off-chip device via the C2C interconnect.
[0181] Example 16: The IC chip of example 1, wherein the NICB circuitry is further configured to receive the flit over the C2C interconnect and determine that the flit is a streaming data flit based on a virtual channel mapping encoded within the flit.
[0182] Example 17: An integrated circuit (IC) device comprising: a functional circuit configured to send and receive memory-mapped data to and from an off-chip device; a network-on-chip (NoC) circuit configured to provide a virtual interface to the off-chip device for the functional circuit; and a NoC inter-chip bridge (NICB) circuit configured to provide a virtual interface to the off-chip device for the NoC circuit.
[0183] Example 18: A method comprising: generating, in an integrated circuit (IC) chip, memory-mapped traffic including a memory access request; packetizing the memory access request; routing the packetized memory access request through an on-chip packet-switched communication network at a first transmission rate; and forwarding the packetized memory access request in a segment from the on-chip packet-switched communication network to an off-chip device via a set of chip-to-chip (C2C) interconnects at a second transmission rate, wherein the second transmission rate is greater than the first transmission rate.
[0184] Example 19: The method of example 18, wherein the forwarding includes performing address map translation and destination identifier remapping based on a system management identifier (SMID) associated with the packetized memory-mapped traffic.
[0185] Example 20: The method of Example 18, wherein the forwarding includes mapping the contents of the packetized memory access request to the C2C interconnect in a 1:1 mapping in full mode, or mapping the contents of the packetized memory access request to the C2C interconnect in a mapping less than 1:1 in compressed mode.
[0186] In the foregoing, reference is made to the embodiments presented in this disclosure. However, the scope of the disclosure is not limited to the specific described embodiments. Instead, any combination of the described features and elements, whether associated with different embodiments or not, is contemplated for implementing and practicing the contemplated embodiments. Moreover, while the embodiments disclosed herein may achieve advantages over other possible solutions or prior art, whether or not a particular advantage is achieved by a given embodiment does not limit the scope of the disclosure. Accordingly, the foregoing aspects, features, embodiments, and advantages are merely illustrative and are not considered elements or limitations of the appended claims unless expressly recited in the claims.
[0187] As will be appreciated by one skilled in the art, embodiments disclosed herein may be embodied as a system, method, or computer program product. Accordingly, aspects may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, all of which may be generally referred to herein as a "circuit," "module," or "system." Furthermore, aspects may take the form of a computer program product embodied in one or more computer-readable medium(s) having computer-readable program code embodied therein.
[0188] Any combination of one or more computer-readable media may be utilized. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (non-exhaustive list) of computer-readable storage media include an electrical connection having one or more wires, a portable computer floppy disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of this specification, a computer-readable storage medium is any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.
[0189] A computer-readable signal medium may include a propagated data signal in which computer-readable program code is embodied, for example, in baseband or as part of a carrier wave. Such a propagated signal may take any of a variety of forms, including, but not limited to, electromagnetic, optical, or any suitable combination thereof. A computer-readable signal medium is not a computer-readable storage medium but may be any computer-readable medium that can communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.
[0190] The program code embodied on the computer readable medium may be transmitted using any suitable medium, including but not limited to wireless, wireline, fiber optic cable, RF, etc., or any suitable combination of the foregoing.
[0191] Computer program code for carrying out operations of aspects of the present disclosure may be written in any combination of one or more programming languages, including, for example, object-oriented programming languages such as Java, Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" programming language or similar programming languages. The program code may execute entirely on the user's computer, partially on the user's computer as a standalone software package, partially on the user's computer, partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet Service Provider).
[0192] Aspects of the present disclosure are described below with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments presented in the present disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, such that the instructions, executed by the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram blocks.
[0193] These computer program instructions may also be stored on a computer-readable storage medium, and the instructions may direct a computer, programmable data processing apparatus, and / or other device to function in a particular manner to produce an article of manufacture including instructions that implement the functions / acts specified in the flowchart and / or block diagram blocks.
[0194] Computer program instructions may also be loaded into a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to create a computer-implemented process, such that the instructions executing on the computer or other programmable apparatus provide a process for implementing the functions / acts specified in the flowchart and / or block diagram blocks.
[0195] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowcharts or block diagrams may represent a module, segment, or portion of instructions, including one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may in fact be executed substantially concurrently, or the blocks may be executed in the reverse order, depending on the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, may be implemented by a dedicated hardware-based system that performs the specified functions or acts or a combination of dedicated hardware and computer instructions.
[0196] While the above is directed to particular examples, other and further examples may be devised without departing from the basic scope thereof, which scope is determined by the following claims.
Claims
1. An integrated circuit (IC) chip comprising: functional circuitry configured to exchange memory-mapped traffic with off-chip devices; network-on-chip (NoC) circuitry configured to packetize and depacketize memory-mapped traffic received from and directed to the functional circuitry, and to route the packetized memory-mapped traffic between the functional circuitry and the off-chip device; and a NoC chip-to-chip bridge (NICB) circuit configured to interface between the NoC circuitry and the off-chip device, including exchanging the packetized memory-mapped traffic with the off-chip device over a chip-to-chip (C2C) interconnect.
2. The NICB circuit includes:
10. The IC chip of claim 1, further configured to perform address map translation and destination identifier remapping based on a system management identifier (SMID) associated with the packetized memory mapped traffic.
3. the NoC circuitry is further configured to route packetized memory access requests to the NICB circuitry at a first transmission rate; the NICB circuitry is further configured to transmit the packetized memory access requests in segments over the C2C interconnect in a 1:1 mapping at a second transmission rate; 2. The IC chip of claim 1, wherein the second transmission rate is greater than the first transmission rate.
4. the NoC circuitry is further configured to route packetized memory access requests to the NICB circuitry at a first transmission rate; the NICB circuitry is further configured to transmit the packetized memory accesses in segments over the C2C interconnect with a less than 1:1 mapping at a second transmission rate; 2. The IC chip of claim 1, wherein the second transmission rate is greater than the first transmission rate.
5. The NICB circuit includes: mapping a command portion of the packetized memory access request to a first flit; mapping data portions of the packetized memory access requests to one or more additional flits; The IC chip of claim 4 , further configured to transmit the first flit and the one or more additional flits in segments over the C2C interconnect at the second transmission rate.
6. The packetized memory access request includes a write request and data, and the NICB circuit: placing the write request in the first flit; calculating an error code based on the first flit; calculating a strobe based on said data; encoding the strobe; placing the data, the error code, and the encoded strobe into the one or more additional flits; The IC chip of claim 5 , further configured to transmit the first flit and the one or more additional flits in segments over the C2C interconnect.
7. The NICB is receiving a packetized memory access request from the NoC circuit; loading the packetized memory access request into a first virtual channel (VC) buffer of a plurality of virtual channel (VC) buffers; mapping the contents of said first VC buffer to a first set of segments, and including an identifier with each segment of said first set of segments; The IC chip of claim 1 , further configured to transmit the first set of segments to the off-chip device via the C2C interconnect.
8. the NICB circuit is further configured to receive a response to the memory access request from the off-chip device as a second set of segments, load the second set of segments into a second one of the VC buffers based on identifiers included in the segments of the second set of segments, concatenate the second set of segments as a flit, and forward the flit to the NoC circuit; The IC chip of claim 7 , wherein the NoC circuitry is further configured to route the flits to the functional circuitry and de-packetize the flits.
9. The NICB circuit includes: receiving first and second flits over the C2C interconnect; The IC chip of claim 1 , further configured to concatenate the first flit and the second flit based on a determination that the first flit includes a predetermined virtual channel identifier.
10. The packetized memory access request includes a read request, and the NICB circuit receiving a response to the read request from the off-chip device in a segment; loading the segment into an entry of a respective one of a plurality of virtual channel (VC) buffers based on an identifier contained within the segment; reordering and joining the segments to provide a frit; The IC chip of claim 1 , further configured to communicate the flits to the NoC circuitry.
11. the NICB circuit is configurable in each of a full mode and a compressed mode; the NoC circuitry is further configured to route packetized memory access requests to the NICB circuitry at a first transmission rate; the NICB circuit is configured, in the full mode, to transmit the packetized memory access requests in segments of a first length over the C2C interconnect in a 1:1 mapping at a second transmission rate; the NICB circuit is configured, in the compressed mode, to transmit the packetized memory accesses in segments of a second length over the C2C interconnect at a third transmission rate with a less than 1:1 mapping; the third transmission rate is greater than the second transmission rate; the second transmission rate is greater than the first transmission rate; The IC chip of claim 1 , wherein the first length is greater than the second length.
12. 2. The IC chip of claim 1, further comprising: a controller circuit configured to dynamically configure the depth of each of a plurality of VC buffers and the starting location of said VC buffers.
13. The NICB circuit includes: Mapping the streaming data across the first and second flits; including a predetermined identifier in the first flit to indicate that the second flit follows; The IC chip of claim 1 , further configured to transmit the first and second flits to the off-chip device via the C2C interconnect.
14. The NICB circuit includes: receiving a flit over the C2C interconnect; The IC chip of claim 1 , further configured to determine that the flit is a streaming data flit based on a virtual channel mapping encoded within the flit.
15. 1. A method, comprising: generating memory-mapped traffic including memory access requests; packetizing the memory access request; routing the packetized memory access requests through an on-chip packet-switched communications network at a first transmission rate; forwarding the packetized memory access requests in segments over a set of chip-to-chip (C2C) interconnects from the on-chip packet-switched communications network to an off-chip device at a second transmission rate; The method wherein the second transmission rate is greater than the first transmission rate.