Global System Interconnect for Integrated Circuits
A global ring interconnect (GCR) in integrated circuits addresses the high cost and interface limitations of NoC by providing a low-cost communication infrastructure for control and data exchange, enhancing efficiency and reducing reliance on NoC interfaces.
Patent Information
- Application Number
- JP2025537908
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-28
- Filing Date
- 2023-10-19
- Publication Date
- 2026-01-21
AI Technical Summary
Existing network on chip (NoC) communications in integrated circuits are high-performance but costly, and many hardware blocks lack NoC interfaces, necessitating alternative low-cost communication methods for control and data exchange.
Implementing a global ring interconnect (GCR) that couples with local rings to provide a low-cost sideband communication infrastructure for integrated circuits, allowing hardware blocks to communicate through a global communication ring (GCR) for control, error handling, and monitoring without requiring NoC interfaces.
The GCR enables efficient, low-cost communication and control across hardware blocks with minimal configuration, supporting error handling, interrupt passing, and monitoring, while reducing dependency on NoC interfaces.
Smart Images

Figure 2026502202000001_ABST
Abstract
Description
[Technical Field]
[0001] Examples of the present disclosure generally relate to ring interconnects in integrated circuits. [Background technology]
[0002] A system on chip (SoC) (e.g., a field programmable gate array (FPGA), a programmable logic device (PLD), or an application specific integrated circuit (ASIC)) may include a packet network structure known as a network on chip (NoC) to route data packets between logic blocks (e.g., programmable logic blocks, processors, memories, etc.) within the SoC.
[0003] Because NoCs are in-band communications, typically most (or all) primary communications occur through the NoC. Therefore, NoCs are high-performance but also very expensive. Furthermore, many hardware blocks within an SoC do not have NoC interfaces, such as transceivers, input / output (I / O) elements, voltage / temperature monitors, or phase-locked loops (PLLs), because they do not require high performance. These hardware blocks may still need to communicate with each other or with a platform management controller. Also, some blocks that have NoC interfaces still prefer to have their control performed through a sideband interface to allow exclusivity in control and reduce intrusion into their in-band traffic. Summary of the Invention
[0004] Techniques are described for defining a global ring coupled to a local ring in an integrated circuit. One example is an integrated circuit including: a global ring including a plurality of switches; a plurality of local rings distributed throughout the IC, each of the plurality of local rings coupled to the global ring by a respective one of a plurality of switches; and a plurality of nodes coupled to the plurality of local rings, wherein the global ring is configured to route packets received from a first node of the plurality of nodes coupled to a first ring of the plurality of local rings to a second node of the plurality of nodes coupled to a second ring of the plurality of local rings.
[0005] Another example is a method of transmitting a packet between nodes communicatively coupled by a ring in an IC, the method including receiving only a first portion of the packet over a number of clock cycles at a first node coupled to the ring, determining whether the first node is the destination of the packet based on the first portion upon determining that the first node is the destination of the packet, receiving a remainder of the packet over a number of additional clock cycles at the first node, and forwarding a null packet to a next node in the ring in parallel with receiving the remainder of the packet during the number of additional clock cycles.
[0006] Another example is an integrated circuit including: a global ring including a plurality of switches; a plurality of local rings, each of the plurality of local rings coupled to the global ring by a respective one of the plurality of switches; and a plurality of nodes coupled to the plurality of local rings, wherein the global ring and at least two of the plurality of switches are used to route packets received from a first node coupled to a first ring of the plurality of local rings to a second node coupled to a second ring of the plurality of local rings. [Brief explanation of the drawings]
[0007] In a manner in which the above-recited features may be understood in detail, a more particular description briefly summarized above may be made by reference to exemplary implementations, some of which are illustrated in the accompanying drawings. It should be noted, however, that the accompanying drawings illustrate only typical example implementations and therefore should not be considered limiting of the scope thereof. [Figure 1] 1 is a block diagram of a global ring interconnecting multiple local rings in an integrated circuit, according to an example. [Figure 2] 1 illustrates the transmission of a packet on a ring, according to an example. [Figure 3] 1 illustrates a switch between a global ring and a local ring, according to an example. [Figure 4] 1 illustrates nodes connected to a ring, according to an example. [Figure 5] 1 shows the microarchitecture of a ring-connected node according to an example. [Figure 6] 1 shows a controller interface, by way of example. [Figure 7] 1 is a flowchart for transmitting a packet on a ring, according to an example.
[0008] For ease of understanding, wherever possible, identical reference numbers have been used to indicate identical elements common to the figures. It is contemplated that elements of one embodiment may be beneficially incorporated in other embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0009] Various features are described below with reference to the drawings. It should be noted that the drawings may or may not be drawn to scale, and that elements of similar structure or function are represented by similar reference numerals throughout the drawings. It should be noted that the drawings are intended only to facilitate the description of features. They are not intended as an exhaustive description of the specification or as limitations on the scope of the claims. In addition, the illustrated example need not have all the aspects or advantages shown. An aspect or advantage described in connection with a particular embodiment is not necessarily limited to that embodiment and may be implemented in any other embodiment, even if not so illustrated or explicitly described.
[0010] Embodiments herein describe an integrated circuit (IC), such as an SoC, FPGA, or ASIC, that includes a global ring that interconnects multiple local rings. In one embodiment, the global and local rings form a sideband communication interconnect between distributed blocks on the IC, through which the blocks can communicate with each other and with a controller or application processor. This interconnect is a low-cost ring interconnect (relative to an NoC), called a global communication ring (GCR), which can be used for global communication and control across associated hardware blocks with minimal or no configuration. That is, the GCR can operate when the IC is first powered on without waiting to be configured. In some embodiments, the GCR can be used to communicate errors, interrupts, secure control, sideband communication, and monitoring capabilities across the IC.
[0011] 1 is a block diagram of a global ring 105 interconnecting multiple local rings 110A-D within an integrated circuit 100, according to one example. The combination of global ring 105 and local rings 110A-D is an example of a GCR. In one embodiment, global ring 105 and local rings 110A-D communicatively couple hardware blocks (e.g., nodes) distributed throughout IC 100.
[0012] As shown, some nodes (i.e., nodes 115A-D) are connected to the global ring 105, and other nodes (i.e., nodes 120A-L) are connected to the local rings 110A-D. Additionally, some or all of the nodes 115, 120 may be connected to an NoC (not shown) within the IC 100. For example, nodes 115, 120 that are transceivers or I / O elements may not be connected to an NoC, but their controller interfaces may be coupled to both the NoC and one of the global ring 105 or the local rings 110. Additionally, there may be other interconnections other than rings and NoCs to which the nodes 115, 120 may be connected.
[0013] Global ring 105 is connected to local ring 110 through switch 125. In this example, there is a one-to-one relationship in which one switch is used to route packets between local ring 110 and global ring 105. In other words, switch 125 provides an interface by which packets inserted into local ring 110 can be sent to global ring 105 and by which packets in global ring 105 can be routed to local ring 110. Furthermore, one or more of local rings 110 may include multiple sub-local rings (e.g., a third level in the hierarchy) each connected to local ring 110 using a respective switch.
[0014] In one embodiment, both nodes 115, 120 can insert packets onto and receive packets from their respective rings. However, in other embodiments, some of the nodes 115, 120 may only transmit packets onto or only receive packets from the rings. Furthermore, while FIG. 1 shows nodes 120 coupled to the global ring 105 that can insert and / or receive packets, in other implementations, there may be no nodes connected to the global ring 105. In that case, the global ring 105 may be used to route packets between local rings 110.
[0015] Packets can be routed through the GCR using several different scenarios. In one example, node 120 can insert onto local ring 110 a packet destined for nodes 120 on the same local ring. The packet can then traverse local ring 110 (in a dedicated direction) until it reaches the destination node. Other nodes in local ring 110 that are not the packet's destination simply forward the packet to the next node in ring 110.
[0016] In another example, node 120 may insert onto local ring 110 a packet destined for node 115 on global ring 105. In that case, the packet traverses local ring 110 (in a dedicated direction) until it reaches switch 125, which then inserts the packet onto global ring 105. The packet then traverses the global ring (in a dedicated direction) until it reaches destination node 115.
[0017] In another example, a node 120 may insert onto the local ring 110 a packet destined for a node 120 on another local ring 110. In that case, the packet traverses the local ring 110 until it reaches a switch 125, which then inserts the packet onto the global ring 105. The packet then traverses the global ring (in a dedicated direction) until it reaches the switch 125 that corresponds to the local ring 110 that contains the destination node 120. The switch 125 inserts the packet onto the local ring, and the packet traverses the local ring 110 until it reaches the destination node 120.
[0018] In another example, a node 115 in the global ring 105 inserts a packet onto the global ring 105 that is destined for another node 115 connected to the global ring 105. In that case, the packet traverses the global ring 105 in a dedicated direction until it reaches the destination node 115. Other nodes in the global ring 105 that are not the packet's destination simply forward the packet to the next node in the ring 105.
[0019] In another example, a node 115 in the global ring 105 inserts a packet onto the global ring 105 that is destined for a node 120 connected to the local ring 110. In that case, the packet traverses the global ring 105 in a dedicated direction until it reaches a switch in the local ring that contains the destination node 120. The switch 125 inserts the packet onto the local ring 110, and the packet circulates around the ring until it reaches the destination node 120.
[0020] In one embodiment, the GCR provides an infrastructure for a controller within the IC 100 to pass eFuse information to different hardware blocks (e.g., nodes 115, 120) on the IC 100 before those blocks are used. The eFuse information can include IP validation, repair information, and potential keys with approved / compliant performance. For example, the controller can be one of the nodes connected to the global ring 105 or one of the local rings 110.
[0021] In one embodiment, the GCR provides interconnect for hardware blocks on IC 100 and chiplets that may be connected to IC 100 to pass interrupts to a controller and processor subsystem (PS) on IC 100 to enable firmware and software interaction. Chilets may be separate ICs that are connected to IC 100 through a substrate (e.g., an interposer).
[0022] In one embodiment, the GCR provides an interconnect for passing error information from nodes distributed on the IC and from chiplets to an error aggregation module in a controller on the IC 100.
[0023] In one embodiment, the GCR provides interconnection between nodes on an IC and controllers on chiplets to pass control information between them.
[0024] In one embodiment, the GCR provides an interconnect for a controller on the IC 100 to broadcast information to a specific group of nodes 115, 120 on the IC 100. One example of a group of nodes may be a configuration interface manager (CIM) responsible for configuring regions of the IC 100. By having the CIM communicate over sideband during initialization, the CIM can be programmed to limit its use of regions on the device or even disabled. In one embodiment, the IC 100 is the base die of a multi-die device with other ICs stacked on top of the base die. In that case, the CIM can configure each 3D slice of the stacked multi-die device.
[0025] In one embodiment, the GCR provides an interconnect for continuous communication between a root system monitor (SysMon) in a controller within IC 100 and satellite SysMons distributed throughout IC 100 or root SysMons on chiplets coupled to IC 100. In one embodiment, SysMons is responsible for monitoring voltages and temperatures.
[0026] In one embodiment, the GCR provides a communication interconnect to allow various agents to be implemented to communicate with the enhancement blocks in place of the programmable FPGA fabric within an IC or chiplet.
[0027] In one embodiment, packet propagation on the GCR ring is deterministic and stall-free. This assumes that no additional pipeline stages are required on the path from one node to another. If the speed cannot be met, additional pipeline stages can be included between nodes.
[0028] FIG. 2 illustrates the transmission of packets on a ring, according to an example. In one embodiment, the global ring and local ring each transmit four data signals and one control signal in parallel. However, other embodiments may transmit more (or fewer) data signals in parallel. The example shown in FIG. 2 includes a 48-bit long GCR packet. Assuming the ring transmits four bits of data (also called a nibble) in each cycle, a GCR packet uses 12 cycles or nibbles to be transmitted or received at a node. In one embodiment, each packet includes a header (e.g., the first four bits) that describes the type of packet. The remaining payload of the packet may include a source address and a destination address, depending on the packet type.
[0029] In one embodiment, the ring is formed from GCR links that connect each node or switch to another node or switch in the ring. These links can include four data signals (wires) and a packet marker signal. Figure 2 shows how packets are communicated over the links. As shown, the packet marker signal identifies that a new packet will start in the next cycle. The four bits sampled in the next cycle contain the packet's header, which identifies the packet type. In this example, each packet takes 12 cycles to traverse an entire node.
[0030] 3 illustrates a switch 125 between a global ring and a local ring, according to an example. As described above, the switch connects the local ring to the global ring. In one embodiment, the switch 125 is located on the global ring and examines GCR packets flowing through it. If the packet's destination address is within the address opening associated with the corresponding local ring, the switch 125 buffers the packet and passes it to the first available slot on the local ring.
[0031] If switch 125 cannot buffer more packets and there is another packet destined for its corresponding local ring, the packet continues its path on the global ring until it returns to switch 125, which again attempts to insert the packet into the local ring (assuming its buffer is no longer full). The same forwarding model can be used for packets sourced in the local ring and targeted for a node connected to the global ring or another local ring. That is, if switch 125 cannot buffer the additional packet, the packet can continue to circulate around the local ring until switch 125 has capacity for the additional packet.
[0032] 4 illustrates a node 400 connected to a ring, according to an example. The node 400 can be either a node connected to a global ring or a node connected to a local ring.
[0033] In one embodiment, the GCR node 400 has enough buffering to hold one complete outbound GCR packet (e.g., 48 bits) and capture one complete inbound GCR packet. In that case, it may be preferable for the GCR node 400 to provide buffering for more packets than this. In that case, the node 400 can construct a new packet while the first packet is being transmitted on the ring.
[0034] In one embodiment, each NoC peripheral interconnect (NPI) slave and NPI root includes a GCR node. In one example, the NPI is another interconnect that can configure the nodes connected to it. In one embodiment, each of the GCR nodes is also connected to an NPI interconnect and uses the same address for both interconnects. However, it is possible for a node to be on an NPI interconnect and not on the GCR ring.
[0035] The NPI root may be contained in a CIM. The NPI root in the CIM block for a controller in an IC may have a special GCR node that acts on behalf of the entire controller, the PS, and the root SysMon.
[0036] To initialize the GCR (e.g., in response to a GCR reset signal), a GCR node 400 corresponding to a controller in an IC can inject packet markers along with null packets into the global ring for a certain number of predetermined cycles. During the initialization phase, GCR switches (e.g., switch 125 of FIG. 3) broadcast null packets and packet markers to both their egress ports (e.g., the egress port coupled to the global ring and the egress port coupled to its corresponding local ring). In one embodiment, this continues until a special service packet is transmitted over the global ring that reaches the GCR switch. At that point, the switch enters an operational mode or phase in which it routes packets based on their destination.
[0037] When node 400 receives a packet, the node stores a first portion (e.g., 16 bits) of the packet that is forwarded in a first number of cycles (e.g., 4 cycles). Using that information, node 400 can determine whether the packet is targeted to node 400. If the packet is not targeted to node 400, the packet begins to leave the node in the fifth cycle; otherwise, a general null packet or a SysMon null packet is pushed onto the ring in the fifth cycle, depending on whether the node is a SysMon block.
[0038] During operation (e.g., after initialization), packet slots may be cleared or reused by different nodes. In one embodiment, this responsibility belongs to the node 400 that is the packet's destination. In some cases, when the packet is a broadcast operation by a controller in the IC, the controller may be responsible for clearing the packet. In one embodiment, the last two bits of the packet's fourth nibble, bits [15:14] of the GCR packet, are labeled as count[1:0] and are used to track packets that have not been captured for some reason and have been alive for some time in the GCR ring. Each time a packet passes through the controller or enters the local ingress port of the GCR switch (sourced from a node on the local GCR ring), if the count[1:0] value is not 0x11, count[1:0] is incremented and the packet continues traveling on the ring.
[0039] If the controller determines that count[1:0] in the packet is 0x11, an error can be generated and logged internally within the controller along with the packet information, and the slot on the ring is cleared and converted to its corresponding null packet or carries another outstanding packet of the same type.
[0040] For GCR switches, if count[1:0] is 0x11, the packet is automatically passed to the global ring and routed to the controller, which generates an error.
[0041] FIG. 5 illustrates the microarchitecture of a node connected to a ring, according to an example. Node 500 can be either a node connected to the global ring or a node connected to a local ring. In one embodiment, the first four nibbles (16 bits) of each packet are buffered in node 500. Assuming there are a maximum of 1024 64KB apertures in the NPI on the IC, nibble 0 can identify the packet type, and nibbles 1, 2, and 3 contain a 10-bit aperture exponent. In this way, the first four nibbles of every packet identify whether the inbound packet is targeted to the current node 500, or if these nibbles are blank (and a null packet is of the appropriate type), node 500 can use this packet to send an outbound packet.
[0042] In other words, by buffering and evaluating the first portion of the packet, the node can determine that the packet is not a null packet and whether the node is the packet's destination, or, if the packet is a null packet, what type of packet can be inserted in its place. That is, some spots on the ring can be reserved for specific types of traffic. For example, when initializing the GCR, one in every four null packets can be reserved for a specific type of traffic (or to be used by a specific type of node). For example, there can be SysMon null packets and non-SysMon null packets inserted into the ring during initialization. If node 500 receives a SysMon null packet but is a non-SysMon node (or does not have a SysMon packet), it cannot replace the null packet. Instead, it forwards the SysMon null packet to the next node / switch in the ring. In this embodiment, node 500 receives a null packet and, if the null packet is of the appropriate type, can replace it with a packet it wants to send to another node in the GCR. In this way, portions of the GCR's bandwidth can be reserved for different types of traffic (eg, SysMon traffic and non-SysMon traffic).
[0043] In FIG. 5, node 500 includes two output packet buffers for storing packets waiting to be sent to another node connected to the GCR.
[0044] Node 500 includes two input capture packet buffers in which packets targeted to node 500 can be stored before they are retrieved and the buffers are marked as empty.
[0045] In one embodiment, if there is an inbound packet targeted to node 500 and none of the input capture buffers are available, node 500 can forward the packet on its egress port to a downstream node, and the packet will circulate around the ring before arriving back at node 500. Ideally, by that point, the input capture packet buffer will have room to receive the packet. Thus, node 500 can have multiple opportunities to capture packets destined for node 500.
[0046] When a packet is captured by node 500, the outbound packet is marked as null and is of the same type as the packet captured by node 500. That is, node 500 replaces the received packet with an outbound null packet of the same type. In this way, the bandwidth reservation of the GCR is maintained between different packet types as they are inserted into and then removed from the GCR.
[0047] 6 illustrates an example controller interface 600. That is, FIG. 6 illustrates an interface 600 used to couple a controller (e.g., a platform management controller) in an IC to a GCR. In one embodiment, the controller interface 600 is a node coupled to a global ring, while in other embodiments, the interface 600 may be a node coupled to a local ring of the GCR.
[0048] In one embodiment, controller interface 600 receives errors from a remote node via a generic interrupt controller (GIC). If the IC is part of a multi-die device combined into chiplets or stacked ICs, interface 600 can receive errors from other dies and chiplets in the device and route them to an error aggregation module (EAM) in the controller. In addition to errors, interrupts may also be communicated by remote nodes to processors in the system via the GIC. In this embodiment, interrupts are captured from a GCR in controller interface 600 as shown and passed from there to a GIC associated with, for example, a particular processor in the system.
[0049] In one embodiment, for example, controller interface 600 for a platform management controller receives interrupts from remote blocks within the IC, as well as interrupts from other dies and chiplets (if part of a multi-die device), and routes these interrupts to the controller and PS.
[0050] In one embodiment, the controller interface 600 allows software and hardware events to be packetized and propagated through the GCR to various destinations connected to the GCR, which may be distributed throughout the IC.
[0051] In one embodiment, the controller interface 600 allows the controller to broadcast events to all CIMs in the IC, with each CIM responsible for configuring segments on the device.
[0052] In one embodiment, controller interface 600 provides an infrastructure for the controller to pass eFuse information to a node on an IC before the node is configured for use. The eFuse information can include IP enable signals with approved / compliant performance, repair information, and potential keys.
[0053] In one embodiment, controller interface 600 allows the root SysMon in the controller to communicate with satellite SysMons on an IC or root SysMons in chiplets attached to the IC.
[0054] 7 is a flowchart of a method 700 for transmitting packets on a ring, according to an example. The method 700 can be used in an IC with multiple rings (e.g., the GCR shown in FIG. 1) or an IC with only one ring.
[0055] In block 705, a node on the ring (e.g., node 115 on global ring 105 or node 120 on local ring 110 in FIG. 1) receives a first portion of the packet over multiple clock cycles. Using the above example of FIG. 2, a node may receive four bits of the packet every clock cycle. The first 16 bits (e.g., the first four clock cycles) may have enough information for the node to perform the remainder of method 700. Thus, a node may not have received the entire packet when performing many of the blocks of method 700.
[0056] At block 710, the node determines whether the packet is a null packet based on the first received portion. If the packet is not null (i.e., the packet was placed on the ring by another node and is not simply a null placeholder packet), method 700 proceeds to block 715, where the node determines whether it is the destination of the packet.
[0057] If the node is not the destination, method 700 proceeds to block 717, where the node forwards the packet to the next node in the ring. That is, the node can begin forwarding the first portion of the packet to the next node. Once the node receives the remainder of the packet over additional clock cycles, it can forward the portion of the packet that it has already received to the next node. Thus, the number of clock cycles used to receive the first portion is the delay or buffer between when the node begins receiving the packet from the upstream node in the ring and when the node begins transmitting the packet to the downstream node in the ring.
[0058] However, if the node is the destination of the packet, method 700 instead proceeds to block 720, where the node receives the remainder of the packet over multiple clock cycles. For example, the first portion of the packet may be received in the first four clock cycles, and the remainder of the packet may be received over an additional eight clock cycles. The packet is then removed from the ring. Once received, the node can process the packet or forward the data in the packet to other circuitry within the IC using a different communication path.
[0059] In block 725, the node forwards the same type of null packet to the next node in the ring (i.e., the downstream node). Notably, block 725 can occur in parallel with block 720. For example, if the node determines that it is the destination in block 715, then in the next clock cycle, the node can begin forwarding the null packet to the next node using its egress port, while the node continues to receive the remainder of the packet in parallel at its ingress port.
[0060] Returning to block 710, if the received packet is a null packet, method 700 proceeds to block 730, where the node determines whether it has a packet to send on the ring. This packet may be to a destination node on the same ring or to a destination node on a different ring connected to the current ring by a switch as described in FIG. 1.
[0061] If the node does not have any packets to send, method 700 proceeds to block 725, where the node begins forwarding a null packet of the same type as the received null packet to the next node in the ring, giving the next node an opportunity to send any packets it may have.
[0062] However, if the node has a packet it wants to send, method 700 proceeds to block 735, where the node determines whether the packet is the same type as a null packet (e.g., both the null packet and the actual packet are SysMon type packets). If no, method 700 proceeds to block 725, where the node is not allowed to insert the packet onto the ring and instead forwards the null packet. However, if the packet is the same type as the null packet, the method proceeds to block 740, where the node replaces the null packet with an actual packet of the same type. That is, after receiving the first portion of the null packet, the node can begin sending the actual packet to the next hop without having to wait until it receives the remainder of the null packet.
[0063] In the foregoing, reference is made to the embodiments presented in this disclosure. However, the scope of the disclosure is not limited to the specific described embodiments. Instead, any combination of the described features and elements, whether associated with different embodiments or not, is contemplated for implementing and practicing the contemplated embodiments. Moreover, while the embodiments disclosed herein may achieve advantages over other possible solutions or prior art, whether or not a particular advantage is achieved by a given embodiment does not limit the scope of the disclosure. Accordingly, the foregoing aspects, features, embodiments, and advantages are merely illustrative and are not considered elements or limitations of the appended claims unless expressly recited in the claims.
[0064] As will be appreciated by one skilled in the art, embodiments disclosed herein may be embodied as a system, method, or computer program product. Accordingly, aspects may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, all of which may be generally referred to herein as a "circuit," "module," or "system." Furthermore, aspects may take the form of a computer program product embodied in one or more computer-readable medium(s) having computer-readable program code embodied therein.
[0065] Any combination of one or more computer-readable media may be utilized. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (non-exhaustive list) of computer-readable storage media include an electrical connection having one or more wires, a portable computer floppy disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of this specification, a computer-readable storage medium is any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.
[0066] A computer-readable signal medium may include a propagated data signal in which computer-readable program code is embodied, for example, in baseband or as part of a carrier wave. Such a propagated signal may take any of a variety of forms, including, but not limited to, electromagnetic, optical, or any suitable combination thereof. A computer-readable signal medium is not a computer-readable storage medium but may be any computer-readable medium that can communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.
[0067] The program code embodied on the computer readable medium may be transmitted using any suitable medium, including but not limited to wireless, wireline, fiber optic cable, RF, etc., or any suitable combination of the foregoing.
[0068] Computer program code for carrying out operations of aspects of the present disclosure may be written in any combination of one or more programming languages, including, for example, object-oriented programming languages such as Java, Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" programming language or similar programming languages. The program code may execute entirely on the user's computer, partially on the user's computer as a standalone software package, partially on the user's computer, partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet Service Provider).
[0069] Aspects of the present disclosure are described below with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments presented in the present disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, such that the instructions, executed by the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram blocks.
[0070] These computer program instructions may also be stored on a computer-readable storage medium, and the instructions may direct a computer, programmable data processing apparatus, and / or other device to function in a particular manner to produce an article of manufacture including instructions that implement the functions / acts specified in the flowchart and / or block diagram blocks.
[0071] Computer program instructions may also be loaded into a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to create a computer-implemented process, such that the instructions executing on the computer or other programmable apparatus provide a process for implementing the functions / acts specified in the flowchart and / or block diagram blocks.
[0072] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowcharts or block diagrams may represent a module, segment, or portion of instructions, including one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may in fact be executed substantially concurrently, or the blocks may be executed in the reverse order, depending on the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, may be implemented by a dedicated hardware-based system that performs the specified functions or acts or a combination of dedicated hardware and computer instructions.
[0073] While the above is directed to particular examples, other and further examples may be devised without departing from the basic scope thereof, which scope is determined by the following claims.
Claims
1. 1. An integrated circuit (IC), comprising: a global ring including a plurality of switches; a plurality of local rings distributed throughout the IC, each of the plurality of local rings coupled to the global ring by a respective one of the plurality of switches; a plurality of nodes coupled to the plurality of local rings, wherein the global ring is configured to route a packet received from a first node of the plurality of nodes coupled to a first ring of the plurality of local rings to a second node of the plurality of nodes coupled to a second ring of the plurality of local rings.
2. The IC of claim 1 , wherein the global ring and the plurality of local rings allow packet flow in only one direction.
3. 2. The IC of claim 1, wherein the packet traverses through at least two of the plurality of switches when routed between the first ring and the second ring by the global ring.
4. Upon receiving the packet, the second node: receiving only a first portion of the packet over a plurality of clock cycles; determining whether the second node is the destination of the packet based on the first portion; receiving a remainder of the packet over a number of additional clock cycles upon determining that the second node is the destination of the packet; 2. The IC of claim 1, configured to forward a null packet to a next node in the second ring in parallel with receiving the remaining portion of the packet during the plurality of additional clock cycles.
5. Before the packet is received at the second node, the packet is received at a third node coupled to the second ring, and the third node: receiving only the first portion of the packet over a plurality of clock cycles; determining whether the third node is the destination of the packet based on the first portion; if the third node determines that it is not the destination of the packet, it receives the remaining portion of the packet over a number of additional clock cycles; 5. The IC of claim 4, configured to forward the first portion of the packet to the second node in parallel with receiving the remaining portions of the packet during the plurality of additional clock cycles.
6. An IC, a global ring including a plurality of switches; a plurality of local rings, each of the plurality of local rings coupled to the global ring by a respective one of the plurality of switches; a plurality of nodes coupled to the plurality of local rings, wherein the global ring and at least two of the plurality of switches are used to route a packet received from a first node coupled to a first ring of the plurality of local rings to a second node coupled to a second ring of the plurality of local rings.
7. 7. The IC of claim 1, wherein a second packet inserted by the first node and destined for a third node coupled to the first ring is transmitted to the third node by the first ring without traversing the global ring.
8. 8. The IC of claim 7, further comprising: a fourth node coupled to the global ring, wherein a third packet inserted by the first node and destined for the fourth node is transmitted to the fourth node by the first ring and the global ring.
9. 7. The IC of claim 6, wherein at least one of the plurality of local rings includes a plurality of sub-local rings connected to the at least one local ring by a plurality of respective switches.
10. Upon receiving the packet, the second node: receiving only a first portion of the packet over a plurality of clock cycles; determining whether the second node is the destination of the packet based on the first portion; receiving a remainder of the packet over a number of additional clock cycles upon determining that the second node is the destination of the packet; 7. The IC of claim 6, configured to forward a null packet to a next node in the second ring in parallel with receiving the remaining portion of the packet during the plurality of additional clock cycles.
11. Before the packet is received at the second node, the packet is received at a third node coupled to the second ring, and the third node: receiving only the first portion of the packet over a plurality of clock cycles; determining whether the third node is the destination of the packet based on the first portion; if the third node determines that it is not the destination of the packet, it receives the remaining portion of the packet over a number of additional clock cycles; 11. The IC of claim 10, configured to forward the first portion of the packet to the second node in parallel with receiving the remaining portions of the packet during the plurality of additional clock cycles.
12. 1. A method for transmitting packets between nodes communicatively coupled by a ring in an IC, comprising: receiving only a first portion of a packet over a plurality of clock cycles at a first node coupled to the ring; determining whether the first node is the destination of the packet based on the first portion; receiving a remainder of the packet at the first node over a number of additional clock cycles upon determining that the first node is the destination of the packet; and forwarding a null packet to a next node in the ring in parallel with receiving the remaining portion of the packet during the plurality of additional clock cycles.
13. before receiving the packet at the first node; receiving only the first portion of the packet over a plurality of clock cycles at a second node coupled to the ring; determining whether the second node is the destination of the packet based on the first portion; receiving the remaining portion of the packet at the second node over a number of additional clock cycles upon determining that the second node is not the destination of the packet; 13. The method of claim 12, further comprising: during the plurality of additional clock cycles, transferring the first portion of the packet from the second node to the first node in parallel with receiving the remaining portion of the packet.
14. receiving only a first portion of a second packet over a plurality of clock cycles at the first node; determining whether the second packet is a null packet based on the first portion; upon determining that the second packet is a null packet, determining whether the first node has a third packet ready to transmit on the ring to a destination node; 13. The method of claim 12, further comprising: when the first node determines that the third packet is ready, replacing the null packet on the ring with the third packet.
15. receiving only a first portion of a second packet over a plurality of clock cycles at the first node; determining whether the second packet is a null packet based on the first portion; upon determining that the second packet is a null packet, determining whether the first node has a third packet ready to transmit on the ring to a destination node; 9. The method of claim 8, further comprising: if the first node determines that it does not have a packet ready to send on the ring, forwarding the null packet to a next node in the ring.