Global system interconnect for integrated circuits

By introducing a global ring interconnect structure in the system on chip, the problem of high NoC communication costs is solved, low-cost global communication and controller interconnection is realized, and efficient communication of all hardware blocks is supported.

CN120418787APending Publication Date: 2025-08-01XILINX INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380088582.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-28
Filing Date
2023-10-19
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

NoC communication in existing systems on chip is expensive and not suitable for all hardware blocks, resulting in some hardware blocks being unable to communicate efficiently, and control communication is often carried out through the sideband interface, affecting in-band traffic.

Method used

The global ring interconnect structure is introduced, including global ring and local ring, which connects hardware blocks through switches to achieve low-cost ring interconnection, supports sideband communication and control functions, and provides a global communication ring (GCR) to pass errors, interrupts, security control and monitoring capabilities within the IC.

Benefits of technology

It realizes low-cost global communication, supports communication of all hardware blocks, reduces dependence on NoC, improves communication efficiency and controller interconnection capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120418787A_ABST
    Figure CN120418787A_ABST
Patent Text Reader

Abstract

Embodiments herein describe an integrated circuit (IC) that includes a global ring that interconnects a plurality of local rings distributed throughout the IC. In one embodiment, the global ring is connected to the local ring using respective switches. The global ring (and the switch) interconnects the local rings such that a node coupled to one of the local rings may communicate with a node connected to another local ring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Examples of the present disclosure generally relate to ring interconnects in integrated circuits. Background Art

[0002] A system-on-chip (SoC) (e.g., a field-programmable gate array (FPGA), a programmable logic device (PLD), or an application-specific integrated circuit (ASIC)) may include a packet network structure called a network-on-chip (NoC) to route data packets between logic blocks (e.g., programmable logic blocks, processors, and memories, etc.) in the SoC.

[0003] Since the NoC is in-band communication, typically most (or all) of the primary communication is done through the NoC. Thus, the NoC is high-performance but also very expensive. Additionally, many hardware blocks in the SoC do not have a NoC interface, such as transceivers, input / output (I / O) elements, voltage / temperature monitors, or phase-locked loops (PLLs), because they do not require high performance. These hardware blocks may still need to communicate with each other or with a platform management controller. Further, some blocks with a NoC interface still tend to perform their control through a sideband interface to allow exclusivity in control and reduce interference with their in-band traffic. Summary of the Invention

[0004] Techniques for defining a global ring coupled to local rings in an integrated circuit are described. One example is an integrated circuit that includes: a global ring that includes a plurality of switches; a plurality of local rings distributed throughout the IC, wherein each of the plurality of local rings is coupled to the global ring through a respective one of the plurality of switches; and a plurality of nodes coupled to the plurality of local rings, wherein the global ring is configured to route a packet received from a first node among the plurality of nodes coupled to a first ring of the plurality of local rings to a second node among the plurality of nodes coupled to a second ring of the plurality of local rings.

[0005] Another example is a method for sending a packet between nodes communicatively coupled by a ring in an IC. The method includes: receiving only a first portion of a packet at a first node coupled to the ring over a plurality of clock cycles; determining, based on the first portion, whether the first node is the destination of the packet; receiving, at the first node, the remaining portion of the packet over a plurality of additional clock cycles when it is determined that the first node is the destination of the packet; and forwarding an empty packet to the next node in the ring in parallel with receiving the remaining portion of the packet during the plurality of additional clock cycles.

[0006] Another example is an integrated circuit that includes: a global ring that includes a plurality of switches; a plurality of local rings, where each of the plurality of local rings is coupled to the global ring through a respective one of the plurality of switches; and a plurality of nodes that are coupled to the plurality of local rings, where the global ring and at least two of the plurality of switches are used to route a packet received from a first node coupled to a first ring of the plurality of local rings to a second node coupled to a second ring of the plurality of local rings. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] To enable a detailed understanding of the manner in which the above-described features are implemented, a more specific description briefly summarized above can be obtained by reference to specific examples of implementations, some of which are illustrated in the accompanying drawings. However, it should be noted that the drawings only illustrate typical examples of implementations and should not be considered as limiting their scope.

[0008] Figure 1 is a block diagram of a global ring of a plurality of local rings in an interconnected integrated circuit according to an example.

[0009] Figure 2 illustrates sending a packet in a ring according to an example.

[0010] Figure 3 illustrates a switch between a global ring and a local ring according to an example.

[0011] Figure 4 illustrates a node connected to a ring according to an example.

[0012] Figure 5 illustrates a microarchitecture of a node connected to a ring according to an example.

[0013] Figure 6 illustrates a controller interface according to an example.

[0014] Figure 7 is a flowchart for sending a packet on a ring according to an example.

[0015] For ease of understanding, where possible, the same reference numerals are used to denote the same elements common to the drawings. It is contemplated that elements of one example can be advantageously incorporated into other examples. DETAILED DESCRIPTION

[0016] Various features are described hereinafter with reference to the accompanying drawings. It should be noted that the drawings may or may not be drawn to scale, and elements of similar structure or function are represented by like reference numerals in all the drawings. It should be noted that the drawings are only intended to facilitate the description of the features. They are not intended as an exhaustive description of the present specification nor as a limitation on the scope of the claims. Additionally, the examples illustrated need not have all the aspects or advantages shown. Aspects or advantages described in connection with a particular example are not necessarily limited to that example and may be practiced in any other example even if not so illustrated or not so explicitly described.

[0017] Embodiments herein describe an integrated circuit (IC), such as a SoC, FPGA, ASIC, etc., that includes a global ring interconnecting a plurality of local rings. In one embodiment, the global ring and the local rings form a sideband communication interconnect between distributed blocks on the IC, through which the blocks can communicate with each other and with a controller or an application processor. The interconnect is a low-cost ring interconnect (as opposed to a NoC), which is referred to as a Global Communication Ring (GCR), and the GCR can be used for global communication and control across related hardware blocks with minimal or no configuration. That is, the GCR can operate when the IC is first powered on without waiting to be configured. In some embodiments, the GCR can be used to convey errors, interrupts, security controls, sideband communication, and monitoring capabilities throughout the IC.

[0018] Figure 1 is a block diagram of a global ring 105 of a plurality of local rings 110A to 110D in an interconnected integrated circuit 100 according to an example. The combination of the global ring 105 and the local rings 110A to 110D is an example of a GCR. In one embodiment, the global ring 105 and the local rings 110A to 110D communicatively couple hardware blocks (e.g., nodes) distributed throughout the IC 100.

[0019] As shown, some nodes (i.e., nodes 115A to 115D) are connected to the global ring 105, while other nodes (i.e., nodes 120A to 120L) are connected to the local rings 110A to 110D. Additionally, some or all of the nodes 115, 120 may also be connected to a NoC (not shown) in the IC 100. For example, nodes 115, 120 that are transceivers or I / O elements may not be connected to the NoC, but a controller interface may be coupled to both the NoC and one of the global ring 105 or the local rings 110 simultaneously. Additionally, there may be other interconnects in addition to the rings and the NoC to which the nodes 115, 120 can be connected.

[0020] The global ring 105 is connected to the local ring 110 via a switch 125. In this example, there is a one-to-one relationship in which one switch is used to route packets between the local ring 110 and the global ring 105. In other words, the switch 125 provides an interface at which packets inserted into the local ring 110 can be transmitted to the global ring 105 and packets in the global ring 105 can be routed to the local ring 110. Additionally, one or more of the local rings in the local ring 110 may include a plurality of sub-local rings (e.g., a third level in a hierarchy), and each of the plurality of sub-local rings is connected to the local ring 110 using a corresponding switch.

[0021] In one embodiment, the nodes 115, 120 can both insert packets into and receive packets from the corresponding ring. However, in other embodiments, some of the nodes 115, 120 may only send packets to the ring or only receive packets from the ring. Additionally, although Figure 1 illustrates the node 120 coupled to the global ring 105 that can insert and / or receive packets, in other implementations, there may not be any nodes connected to the global ring 105. In such a case, the global ring 105 can be used to route packets between the local rings 110.

[0022] Several different scenarios can be used to route packets via the GCR. In one example, the node 120 can insert a packet on the local ring 110 that is destined for the node 120 on the same local ring. In this case, the packet can (in a dedicated direction) traverse the local ring 110 until it reaches the destination node. Other nodes in the local ring 110 that are not the destination of the packet simply forward the packet to the next node in the ring 110.

[0023] In another example, the node 120 can insert a packet on the local ring 110 that is destined for the node 115 on the global ring 105. In this case, the packet (in a dedicated direction) traverses the local ring 110 until it reaches the switch 125, which then inserts the packet into the global ring 105. The packet then (in a dedicated direction) traverses the global ring until it reaches the destination node 115.

[0024] In another example, node 120 may insert a packet destined for node 120 on another local ring 110 onto another local ring 110. In this case, the packet traverses local ring 110 until it reaches switch 125, which then inserts the packet into global ring 105. The packet then traverses the global ring (in the dedicated direction) until it reaches switch 125 corresponding to the local ring 110 that includes the destination node 120. Switch 125 inserts the packet into the local ring, where the packet traverses local ring 110 until it reaches the destination node 120.

[0025] In another example, a node 115 in global ring 105 inserts a packet destined for another node 115 connected to global ring 105 into global ring 105. In this case, the packet traverses global ring 105 in the dedicated direction until it reaches the destination node 115. Other nodes in global ring 105 that are not the destination of the packet simply forward the packet to the next node in ring 105.

[0026] In another example, a node 115 in global ring 105 inserts a packet destined for node 120 connected to local ring 110 into global ring 105. In this case, the packet traverses global ring 105 in the dedicated direction until it reaches the switch of the local ring with the destination node 120. Switch 125 inserts the packet into local ring 110, where the packet traverses the ring until it reaches the destination node 120.

[0027] In one embodiment, the GCR provides infrastructure for the controller in IC 100 to pass eFuse information to different hardware blocks (e.g., nodes 115, 120) on IC 100 before those blocks are used. The eFuse information may include IP enablement, repair information, and potential keys with approval / compatibility capabilities. For example, the controller can be one of the nodes connected to one of the local rings in global ring 105 or local ring 110.

[0028] In one embodiment, the GCR provides interconnection for the hardware blocks and dies on IC 100, and the dies can be connected to IC 100 to pass interrupts to the controller and processor subsystem (PS) on IC 100 to allow firmware and software interaction. The die can be a separate IC connected to IC 100 via a substrate (e.g., an interposer).

[0029] In one embodiment, the GCR provides interconnection to pass error information from the nodes distributed on the IC and from the dies to the error aggregation module in the controller on IC 100.

[0030] In one embodiment, the GCR provides interconnection for nodes on an IC and a controller on a dielet to transfer control information to each other.

[0031] In one embodiment, the GCR provides interconnection for a controller on IC 100 to broadcast information to a specific set of nodes 115, 120 on IC 100. An example of a set of nodes can be a Configuration Interface Manager (CIM) that is responsible for configuring regions of IC 100. By enabling the CIM to communicate via sideband during initialization, the CIM can be programmed or even disabled to restrict the use of regions on the device. In one embodiment, IC 100 is a base die of a multi-die device where other ICs are stacked above the base die. In such a case, the CIM can configure the corresponding 3D slices of the stacked multi-die device.

[0032] In one embodiment, the GCR provides interconnection for continuous communication between a Root System Monitor (SysMon) in a controller in IC 100 and satellite SysMons distributed through IC 100 or root SysMons on a dielet coupled to IC 100. In one embodiment, the SysMons are responsible for monitoring voltage and temperature.

[0033] In one embodiment, the GCR provides a communication interconnection to allow various agents to be implemented to communicate on behalf of programmable FPGA fabric in an IC or a dielet with hardened blocks.

[0034] In one embodiment, the propagation of packets on the GCR ring is deterministic and does not stop. This assumes that no additional pipeline stages are required on the path between one node and another. If the speed cannot be met, additional pipeline stages between nodes can be included.

[0035] Figure 2 An example of sending packets in a ring according to an example is illustrated. In one embodiment, each local ring in the global ring and local rings sends four data signals and one control signal in parallel. However, other embodiments can send more (or fewer) data signals in parallel. Figure 2 The example shown includes a 48-bit long GCR packet. Assuming the ring sends 4 bits of data (also known as a nibble) per cycle, the GCR packet takes 12 cycles or nibbles to be sent or received at a node. In one embodiment, each packet contains a header (e.g., the first four bits) that describes the packet type. The remaining payload of the packet can include a source address and a destination address depending on the packet type.

[0036] In one embodiment, the ring is formed by GCR links that connect each node or switch to another node or switch in a ring. These links can include four data signals (wires) and a packet tag signal. Figure 2 Illustrates how packets are conveyed over the links. As shown, the packet tag signal identifies that a new packet will start in the next cycle. The 4 bits sampled in the next cycle contain a header for the packet that identifies the packet type. In this example, each packet takes 12 cycles to traverse the entire node.

[0037] Figure 3 Illustrates switch 125 between the global ring and a local ring according to the example. As mentioned above, the switch connects the local ring to the global ring. In one embodiment, switch 125 is located on the global ring and examines the GCR packets flowing through it. If the destination address of a packet is within the address aperture associated with the corresponding local ring, switch 125 buffers the packet and delivers it to the first available time slot on the local ring.

[0038] If switch 125 cannot buffer more packets and there is another packet destined for its corresponding local ring, the packet will continue its path on the global ring until it returns to switch 125, where the switch will again attempt to insert the packet into the local ring (assuming its buffer is no longer full). The same transmission model can be used for packets originating from a local ring and destined for a node connected to the global ring or another local ring. That is, if switch 125 cannot buffer additional packets, these packets can continue to circulate around the local ring until switch 125 has capacity for the additional packets.

[0039] Figure 4 Illustrates node 400 connected to the ring according to the example. Node 400 can be a node connected to the global ring or a node connected to a local ring.

[0040] In one embodiment, GCR node 400 has sufficient buffering to hold one complete outbound GCR packet (e.g., 48 bits) and capture one complete inbound GCR packet. In some cases, it may be preferred for GCR node 400 to buffer more packets than this. In such a case, node 400 can construct a new packet while sending the first packet on the ring.

[0041] In one embodiment, each NoC Peripheral Interconnect (NPI) slave device and NPI root includes a GCR node. In one example, the NPI is another interconnect to which nodes can be configured to connect. In one embodiment, each GCR node is also connected to the NPI interconnect and uses the same address for both interconnections. However, a node may be on the NPI interconnect but not on the GCR ring.

[0042] The NPI root can be included in the CIM. The NPI root (which is in the CIM block of the controller in the IC) can have special GCR nodes representing the entire controller, PS, and root SysMon operations.

[0043] To initialize the GCR (e.g., in response to a GCR reset signal), the GCR node 400 corresponding to the controller in the IC can inject packet tags and empty packets into the global ring for a certain number of predetermined cycles. During the initialization phase, the GCR switch (e.g., Figure 3 switch 125 in) broadcasts the empty packets and packet tags to two of their egress ports (e.g., the egress port coupled to the global ring and the egress port coupled to its corresponding local ring). In one embodiment, this continues until a special service packet reaches the GCR switch via the global ring. At this point, the switch enters an operation mode or phase in which the switch routes packets based on the destination of the packet.

[0044] When node 400 receives a packet, the node stores the first part (e.g., 16 bits) of the packet transmitted in the first number of cycles (e.g., 4 cycles). Using this information, node 400 can determine whether the packet is destined for node 400. If the packet is not destined for node 400, then in the fifth cycle, the packet will start to leave the node; otherwise, depending on whether the node is a SysMon block, in the fifth cycle, a general empty packet or a SysMon empty packet is pushed onto the ring.

[0045] During operation (e.g., after initialization), packet time slots are cleared or reused by different nodes. In one embodiment, this responsibility belongs to node 400 which is the destination of the packet. In some cases, when the packet is a broadcast operation by the controller in the IC, the controller may be responsible for clearing the packet. In one embodiment, bits [15:14] of the GCR packet (which are the last 2 bits in the fourth nibble of the packet) are labeled as Count[1:0], and are used to track packets that are not captured for some reason and survive in the GCR ring for some time. Whenever a packet enters (from a node on the local GCR ring) through the controller or at the local ingress port of the GCR switch, if the Count[1:0] value is not 0x11, then Count[1:0] is incremented; and the packet continues its journey on the ring.

[0046] If the controller determines that Count[1:0] in the packet is 0x11, an error can be generated and recorded inside the controller along with the packet information, and the time slot on the ring is cleared and converted to its corresponding empty packet or another unfinished packet of the same type.

[0047] In the case of a GCR switch, if Count[1:0] is 0x11, the packet is automatically passed to the global ring, allowing it to be routed to a controller, which then generates an error.

[0048] Figure 5 Illustrated is the microarchitecture of a node connected to a ring. Node 500 can be a node connected to the global ring or a node connected to a local ring. In one embodiment, the first four nibbles (16 bits) of each packet are buffered in node 500. Assuming there are at most 1024 64KB apertures in the NPI on the IC, nibble 0 can identify the packet type, and nibbles 1, 2, and 3 include a 10-bit aperture index. In this way, the first four nibbles of each packet identify whether the inbound packet is destined for the current node 500, or if these nibbles are blank (and the empty packet is of the appropriate type), node 500 can use the packet to transmit an outbound packet.

[0049] In other words, by buffering and evaluating the first part of the packet, the packet can determine that the packet is not an empty packet, and whether the node is the destination of the packet, or if the packet is an empty packet, it can determine what type of packet can be inserted in its place. That is, some points on the ring can be reserved for specific types of traffic. For example, when initializing the GCR, one out of every four empty packets can be reserved for a specific type of traffic (or used by a specific type of node). For example, there may be SysMon empty packets and non-SysMon empty packets inserted into the ring during initialization. If node 500 receives a SysMon empty packet, but the node is a non-SysMon node (or does not have SysMon packets), the node cannot replace the empty packet. Instead, the node forwards the SysMon empty packet to the next node / switch in the ring. In this embodiment, if node 500 receives an empty packet and the empty packet is of the appropriate type, the node can replace the packet with a packet it wishes to transmit to another node in the GCR. In this way, a portion of the GCR's bandwidth can be reserved for different types of traffic (e.g., SysMon and non-SysMon traffic).

[0050] In Figure 5 Node 500 includes two output packet buffers for storing packets waiting to be transmitted to another node connected to the GCR.

[0051] Node 500 includes two input capture packet buffers that can store packets destined for node 500 before retrieval, and the buffers are marked as empty.

[0052] In one embodiment, if there are inbound packets targeted at node 500 and the input capture buffers are all unavailable, node 500 can forward the packet on its egress port to a downstream node where the packet loops around the ring before returning to node 500. Ideally, by then, there is space in these input capture packet buffers to receive the packet. Thus, node 500 can have multiple opportunities to capture packets intended for node 500.

[0053] When node 500 captures a packet, the outbound packet is marked as empty and has the same type as the packet captured by node 500. That is, node 500 replaces the packet it receives with an outbound empty packet of the same type. In this way, when different packet types are inserted and then removed from the GCR, the bandwidth reservation of the GCR is maintained between these packet types.

[0054] Figure 6 Illustrated is a controller interface 600 according to an example. That is, Figure 6 Illustrated is an interface 600 for coupling a controller (e.g., a platform management controller) in an IC to the GCR. In one embodiment, the controller interface 600 is a node coupled to the global ring, but in other embodiments, the interface 600 can be a node coupled to a local ring of the GCR.

[0055] In one embodiment, the controller interface 600 receives errors from remote nodes via a General Interrupt Controller (GIC). If the IC is part of a multi-die device coupled to a dielet or a stacked IC, the interface 600 can receive errors from other dies and dielets in the device and route these errors to an Error Aggregation Module (EAM) in the controller. In addition to errors, interrupts can also be communicated by remote nodes to the processors in the system via the GIC. In this embodiment, as shown, interrupts are captured from the GCR in the controller interface 600 and passed from there to, for example, the GIC associated with a specific processor in the system.

[0056] In one embodiment, the controller interface 600, for example for a platform management controller, receives interrupts from remote blocks in the IC and from other dies and dielets (if part of a multi-die device) and routes these interrupts to the controller and the PS.

[0057] In one embodiment, the controller interface 600 allows software events and hardware events to be grouped and propagated via the GCR to various destinations connected to the GCR, which can be distributed throughout the IC.

[0058] In one embodiment, the controller interface 600 enables the controller to broadcast events to each CIM in the IC, where each CIM is responsible for configuration on a section of the device.

[0059] In one embodiment, the controller interface 600 provides the controller with an infrastructure to pass eFuse information to nodes on the IC before the nodes are configured for use. The eFuse information may include an IP enable signal with approval / compatibility capabilities, repair information, and potential keys.

[0060] In one embodiment, the controller interface 600 allows the root SysMon within the controller to communicate with satellite SysMons on the IC or the root SysMon in a die attached to the IC.

[0061] Figure 7 is a flowchart of a method 700 for sending packets on a ring. The method 700 can be used in an IC with multiple rings (e.g., Figure 1 the illustrated GCR) or an IC with only one ring.

[0062] At block 705, a node on the ring (e.g., Figure 1 node 115 on the global ring 105 or node 120 on the local ring 110 in Figure 2 receives the first part of a packet over multiple clock cycles. Using the example above Figure 2 in, the node may receive four bits of the packet per clock cycle. The first 16 bits (e.g., the first four clock cycles) may have enough information for the node to execute the remainder of method 700. Thus, when the node executes many of the blocks in method 700, the node may not have received the entire packet.

[0063] At block 710, the node determines whether the packet is an empty packet based on the first received part. If the packet is not empty (i.e., the packet was placed on the ring by another node and is not just an empty placeholder packet), then method 700 proceeds to block 715, where the node determines whether it is the destination of the packet.

[0064] If the node is not the destination, then method 700 proceeds to block 717, where the node forwards the packet to the next node in the ring. That is, the node may start forwarding the first part of the packet to the next node. When the node receives the remainder of the packet over additional clock cycles, the node may forward the part of the packet that it has already received to the next node. Thus, the number of clock cycles used to receive the first part is the delay or buffer between when a node starts receiving a packet from an upstream node in the ring and when the node starts sending the packet to a downstream node in the ring.

[0065] However, if the node is the destination of the packet, method 700 proceeds to block 720, where the node receives the remainder of the packet over multiple clock cycles. For example, the first part of the packet may be received in the first four clock cycles, while the remainder of the packet is received in eight additional clock cycles. The packet is then removed from the ring. Once received, the node can process the packet or forward the data in the packet to other circuits in the IC using a different communication path.

[0066] At block 725, the node forwards an empty packet of the same type to the next node in the ring (i.e., the downstream node). It is noted that block 725 may occur in parallel with block 720. For example, once the node determines at block 715 that it is the destination, the node can start forwarding the empty packet to the next node using the egress port in the next clock cycle, while the node continues to receive the remainder of the packet in parallel at the ingress port.

[0067] Returning to block 710, if the received packet is an empty packet, method 700 proceeds to block 730, where the node determines whether it has a packet to transmit on the ring. The packet may go to a destination node on the same ring or to a destination node on a different ring connected to the current ring by a switch as discussed in Figure 1 If the node does not have a packet it wants to transmit, method 700 proceeds to block 725, where the node starts forwarding an empty packet of the same type as the received empty packet to the next node in the ring. This provides the possibility for the next node to transmit any packet it may have.

[0068] However, if the node does have a packet it wants to transmit, method 700 proceeds to block 735, where the node determines whether the packet is of the same type as the empty packet (e.g., both the empty packet and the actual packet are SysMon type packets). If not, method 700 proceeds to block 725, where the node is not allowed to insert its packet into the ring, but instead forwards the empty packet. But if the packet is of the same type as the empty packet, the method proceeds to block 740, where the node replaces the empty packet with the actual packet of the same type. That is, after receiving the first part of the empty packet, the node can start transmitting the actual packet to the next hop without having to wait until the remainder of the empty packet is received.

[0069]

[0070] ​Previously, reference has been made to embodiments presented in the present disclosure. However, the scope of the present disclosure is not limited to the specifically described embodiments. Instead, any combination of the described features and elements (whether or not they relate to different embodiments) is contemplated for implementing and practicing the contemplated embodiments. Additionally, although the embodiments disclosed herein may achieve advantages over other possible solutions or over the prior art, whether a particular advantage is achieved by a given embodiment does not limit the scope of the present disclosure. Accordingly, the foregoing aspects, features, embodiments, and advantages are illustrative only and are not to be considered elements or limitations of the appended claims unless expressly recited therein.

[0071] As will be understood by those skilled in the art, the embodiments disclosed herein may be embodied as a system, method, or computer program product. Accordingly, aspects may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, all of which may herein be generally referred to as a “circuit,” “module,” or “system.” Additionally, aspects may take the form of a computer program product embodied in one or more computer-readable media having computer-readable program code embodied thereon.

[0072] Any combination of one or more computer-readable media may be utilized. A computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be (by way of example, but not limitation) an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium would include the following: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of this document, a computer-readable storage medium is any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.

[0073] A computer-readable signal medium may include a propagated data signal having computer-readable program code embodied therein (e.g., in baseband or as part of a carrier wave). Such a propagated signal may take any of a variety of forms, including but not limited to electromagnetic, optical, or any suitable combination thereof. A computer-readable signal medium may be any computer-readable medium that is not a computer-readable storage medium and that can communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.

[0074] The program code embodied on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to wireless, wired, fiber optic cable, RF, etc. or any suitable combination of the foregoing.

[0075] The computer program code for performing operations for aspects of the present disclosure can be written in any combination of one or more programming languages, including object-oriented programming languages (such as Java, Smalltalk, C++, etc.) and conventional procedural programming languages (such as the "C" programming language or similar programming languages). The program code can be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0076] Aspects of the present disclosure are described below with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments presented in the present disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, a special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions executed via the processor of the computer or other programmable data processing apparatus create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0077] These computer program instructions can also be stored in a computer-readable medium that can direct a computer, other programmable data processing apparatus, or other device to function in a particular manner, such that the instructions stored in the computer-readable medium produce an article of manufacture including instructions for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0078] The computer program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to produce a computer-implemented method, such that the instructions executed on the computer or other programmable apparatus provide a process for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0079] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various examples of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions that includes one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, depending on the functionality involved, two blocks shown in succession may in fact be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or combinations of dedicated hardware and computer instructions.

[0080] While the foregoing is directed to specific examples, other and additional examples can be devised without departing from the basic scope of the present invention, and the scope of the present invention is determined by the appended claims.

Claims

1. An integrated circuit (IC), the integrated circuit (IC) comprising: A global ring, the global ring including a plurality of switches; A plurality of local rings, the plurality of local rings being distributed throughout the IC, wherein each local ring of the plurality of local rings is coupled to the global ring through a respective one of the plurality of switches; And A plurality of nodes, the plurality of nodes being coupled to the plurality of local rings, wherein the global ring is configured to route a packet received from a first node of the plurality of nodes coupled to a first ring of the plurality of local rings to a second node of the plurality of nodes coupled to a second ring of the plurality of local rings.

2. The IC according to claim 1, wherein the global ring and the plurality of local rings only permit packet flow in one direction.

3. The IC according to claim 1, wherein when the packet is routed by the global ring between the first ring and the second ring, the packet traverses at least two of the plurality of switches.

4. The IC according to claim 1, wherein when the packet is received, the second node is configured to: Receive only a first portion of the packet within a plurality of clock cycles; Determine whether the second node is the destination of the packet based on the first portion; Receive the remaining portion of the packet within a plurality of additional clock cycles when it is determined that the second node is the destination of the packet; And Forward an empty packet to the next node in the second ring in parallel with receiving the remaining portion of the packet during the plurality of additional clock cycles.

5. The IC according to claim 4, wherein the packet is received at a third node coupled to the second ring before being received at the second node, wherein the third node is configured to: Receive only the first portion of the packet within a plurality of clock cycles; Determine whether the third node is the destination of the packet based on the first portion; Receive the remaining portion of the packet within a plurality of additional clock cycles when it is determined that the third node is not the destination of the packet; And Forward the first portion of the packet to the second node in parallel with receiving the remaining portion of the packet during the plurality of additional clock cycles.

6. An IC, the IC comprising: A global ring, the global ring including a plurality of switches; A plurality of local rings, wherein each local ring of the plurality of local rings is coupled to the global ring through a respective one of the plurality of switches; and A plurality of nodes, the plurality of nodes being coupled to the plurality of local rings, wherein the global ring and at least two of the plurality of switches are used to route a packet received from a first node coupled to a first ring of the plurality of local rings to a second node coupled to a second ring of the plurality of local rings.

7. The IC according to any one of claims 1 to 6, wherein a second packet with a destination of a third node coupled to the first ring inserted by the first node is sent by the first ring to the third node without traversing the global ring.

8. The IC according to claim 7, wherein the IC further comprises: A fourth node coupled to the global ring, wherein a third packet with the fourth node as the destination inserted by the first node is sent to the fourth node by the first ring and the global ring.

9. The IC according to claim 6, wherein at least one of the plurality of local rings comprises a plurality of sub-local rings, and the plurality of sub-local rings are connected to the at least one local ring through a plurality of corresponding switches.

10. The IC according to claim 6, wherein when receiving the packet, the second node is configured to: Receive only the first part of the packet within a plurality of clock cycles; Determine whether the second node is the destination of the packet based on the first part; When determining that the second node is the destination of the packet, receive the remaining part of the packet within a plurality of additional clock cycles; And Forward an empty packet to the next node in the second ring in parallel with receiving the remaining part of the packet during the plurality of additional clock cycles.

11. The IC according to claim 10, wherein before receiving the packet at the second node, the packet is received at a third node coupled to the second ring, and the third node is configured to: Receive only the first part of the packet within a plurality of clock cycles; Determine whether the third node is the destination of the packet based on the first part; When determining that the third node is not the destination of the packet, receive the remaining part of the packet within a plurality of additional clock cycles; And Forward the first part of the packet to the second node in parallel with receiving the remaining part of the packet during the plurality of additional clock cycles.

12. A method for sending a packet between nodes communicatively coupled by a ring in an IC, the method comprising: Receiving only the first part of the packet within a plurality of clock cycles at a first node coupled to the ring; Determining whether the first node is the destination of the packet based on the first part; When determining that the first node is the destination of the packet, receiving the remaining part of the packet at the first node within a plurality of additional clock cycles; And Forwarding an empty packet to the next node in the ring in parallel with receiving the remaining part of the packet during the plurality of additional clock cycles.

13. The method according to claim 12, wherein the method further comprises: Before receiving the packet at the first node: Receiving only the first part of the packet within a plurality of clock cycles at a second node coupled to the ring; Determining whether the second node is the destination of the packet based on the first part; When determining that the second node is not the destination of the packet, receiving the remaining part of the packet at the second node within a plurality of additional clock cycles; And Forwarding the first part of the packet from the second node to the first node in parallel with receiving the remaining part of the packet during the plurality of additional clock cycles.

14. The method according to claim 12, wherein the method further comprises: At the first node, only receive a first portion of a second packet over a plurality of clock cycles; Determine whether the second packet is an empty packet based on the first portion; When determining that the second packet is an empty packet, determine whether the first node has a third packet ready to be transmitted over the ring to a destination node; And When determining that the first node has the ready third packet, replace the empty packet with the third packet on the ring.

15. The method according to claim 8, the method further comprising: At the first node, only receive a first portion of a second packet over a plurality of clock cycles; Determine whether the second packet is an empty packet based on the first portion; When determining that the second packet is an empty packet, determine whether the first node has a third packet ready to be transmitted over the ring to a destination node; And When determining that the first node does not have a packet ready to be transmitted over the ring, forward the empty packet to the next node in the ring.