INDEPENDENT ON-CHIP MULTI-CHECKING

DE102022109273B4Active Publication Date: 2025-10-30APPLE INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
DE102022109273
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-06-03
Filing Date
2022-04-14
Publication Date
2025-10-30
Estimated Expiration
2042-04-14

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

System, comprehensive: a plurality of processor clusters, wherein a given processor cluster comprises one or more processors; a large number of graphics processing units; a variety of storage controllers configured to control access to storage devices; a multitude of agents; and a multitude of network switches coupled to the multitude of processor clusters, the multitude of graphics processing units, the multitude of memory controllers, and the multitude of agents, wherein: a first subset of the multitude of network switches is interconnected to form a network of a central processing unit (CPU) between the multitude of processor clusters and the multitude of memory controllers, a second subset of the multitude of network switches is interconnected to form an input / output (I / O) network between the multitude of processor clusters, the multitude of agents, and the multitude of memory controllers, a third subset of the multitude of network switches is interconnected to form a loosely sorted network between the multitude of graphics processing units, selected agents of the multitude of agents, and the multitude of memory controllers, The CPU network, the I / O network, and the loosely sorted network are independent of each other. the CPU network and the I / O network are coherent and The network with loosened sorting is incoherent and has reduced sorting constraints compared to the CPU network and the I / O network.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND Technical area

[0001] The embodiments described herein relate to integrated circuits of a system-on-a-chip (SOC) and in particular to interconnections between components in a SOC. Description of the state of the art

[0002] The integrated circuits of a system-on-a-chip (SoC) generally include one or more processors, which serve as the central processing units (CPUs) for a system, along with various other components, such as memory controllers and peripherals. Additional components can be included within the SoC to form a given device. However, as the number of transistors achievable on an integrated circuit die continues to increase, it is possible to integrate increased numbers of processors and other components onto a given SoC, thereby reducing the number of other components required to form the device.

[0003] Increasing the number of processors and other discrete components on a SoC is desirable for improved performance. Additionally, cost savings can be achieved by reducing the number of other components required besides the SoC to complete the device. The device can be more compact (smaller in size) when more of the overall system is integrated into the SoC. Furthermore, integrating more components into the SoC can reduce the overall power consumption of the device.

[0004] On the other hand, increasing the number of processors and other components on the SoC increases the bandwidth requirements between the memory controllers and the components, and can overwhelm the interconnection used for communication on the SoC, potentially leading to increased latency. The lack of available bandwidth and the increased latency can diminish the performance benefits that were intended to be achieved by integrating the components into the SoC.

[0005] US 2020 / 0 153 757 A1 discloses a method for forwarding FLITS in a network-on-chip, wherein routers forward FLITS based on the operating states of routers and communication links, taking into account routing characteristics such as latency or power consumption.

[0006] US 2014 / 0301241A1 discloses a method for the automatic and dynamic construction of multiple heterogeneous Network-on-Chip (NoC) layers with different topologies, wherein system traffic flows are distributed to suitable NoC layers and routes according to their latency and bandwidth requirements to meet performance requirements such as low latency, high bandwidth, and traffic isolation. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] The following detailed description refers to the accompanying drawings, which will now be briefly described. Fig. Figure 1 is a block diagram of a system that includes an embodiment of multiple network interconnection agents. Fig. Figure 2 is a block diagram of an embodiment of a network that uses a ring topology. Fig. Figure 3 is a block diagram of an embodiment of a network that uses a mesh topology. Fig. Figure 4 is a block diagram of an embodiment of a network that uses a tree topology. Fig. Figure 5 is a block diagram of an embodiment of a system-on-a-chip (SOC) that includes multiple networks for one embodiment. Fig. Figure 6 is a block diagram of an embodiment of a system-on-a-chip (SOC) illustrating one of the independent networks that are in Fig. Figure 5 shows one embodiment. Fig. Figure 7 is a block diagram of an embodiment of a system-on-a-chip (SOC) illustrating another of the independent networks that are in Fig. Figure 5 shows one embodiment. Fig. Figure 8 is a block diagram of an embodiment of a system-on-a-chip (SOC), illustrating yet another of the independent networks that are in Fig. Figure 5 shows one embodiment. Fig. Figure 9 is a block diagram of an embodiment of a multi-die system that includes two semiconductor dies. Fig. Figure 10 is a block diagram of an embodiment of an input / output cluster (I / O cluster). Fig. Figure 11 is a block diagram of an embodiment of a processor cluster. Fig. 12 is a pair of tables illustrating virtual channels, traffic types, and networks that are in Fig. Figures 5 to 8 show how they are used for one embodiment. Fig. Figure 13 is a flowchart illustrating one implementation of initiating a transaction on a network. Fig. Figure 14 is a block diagram of an embodiment of a system. Fig. Figure 15 is a block diagram of an embodiment of a computer-accessible storage medium.

[0008] The invention is described in the accompanying set of claims. While embodiments described in this disclosure may be subject to various modifications and alternative forms, specific embodiments are shown by way of example in the drawings and are described in detail herein. It is understood, however, that the drawings and the detailed description thereto are not intended to limit the embodiments of the disclosed particular form, but rather that, on the contrary, all modifications, equivalents, and alternatives that are within the nature and scope of protection of the accompanying patent claims are intended to be covered. The headings used herein serve only organizational purposes and are not intended to limit the scope of the description. DETAILED DESCRIPTION OF EXECUTION FORMS

[0009] In one embodiment, a system-on-a-chip (SOC) can include a plurality of independent networks. The networks can be physically independent (e.g., having dedicated wires and other switching logic that form the network) and logically independent (e.g., communications provided by agents in the SOC can be logically defined to be transmitted on a selected network from among the plurality of networks and cannot be affected by transmissions on other networks). In some embodiments, network switches can be included to transmit packets on a given network. The network switches can be a physical part of the network (e.g., dedicated network switches can be present for each network).In other embodiments, a network switch can be shared between physically independent networks, thus ensuring that communication received on one of the networks remains on that network.

[0010] By providing physically and logically independent networks, high bandwidth can be achieved through parallel communication across these networks. Additionally, different types of traffic can be carried on different networks, allowing a given network to be optimized for a specific traffic type. For example, processors such as central processing units (CPUs) in a system-on-a-chip (SoC) can be sensitive to memory latency and may cache data that is expected to be coherent between the processors and memory. Accordingly, a CPU network can be provided where the CPUs and memory controllers in a system act as agents. This CPU network can be optimized to provide low latency. For example, in one implementation, virtual channels can be provided for low-latency requests and bulk requests.Low-latency requests can be prioritized over bulk requests during routing within the fabric and by memory controllers. The CPU network can also support cache coherence, with defined messages and protocols for coherent communication. Another network can be an input / output (I / O) network. This network can be used by various peripheral devices to communicate with memory. The network can support the bandwidth required by the peripheral devices and can also support cache coherence. However, I / O traffic can sometimes have significantly higher latency than CPU traffic. Separating I / O traffic from CPU-to-memory traffic can reduce the impact of I / O traffic on CPU traffic.The CPUs can also be included as agents on the I / O network to manage coherence and communicate with peripherals. In yet another embodiment, a loosened-sort network can be used. The CPU and I / O networks can both support sorting models among the communications on these networks, providing the sorting expected by the CPUs and peripherals. However, the loosened-sort network may be incoherent and potentially unable to enforce as many sorting constraints. Graphics processing units (GPUs) can use the loosened-sort network to communicate with memory controllers. Thus, the GPUs can have dedicated bandwidth on the networks and are not limited by the sorting required by the CPUs and / or peripherals.Other embodiments can employ any subset of the foregoing networks and / or any additional networks as desired.

[0011] A network switch can be a circuit configured to receive communications on a network and forward those communications toward their destination. For example, communication provided by a processor might be passed to a memory controller, which controls the memory associated with the communication's address. At each network switch, the communication can be forwarded toward the memory controller. If the communication is a read operation, the memory controller can communicate the data back to the source, and each network switch can forward the data on the network toward the source. In one embodiment, the network can support a plurality of virtual channels. The network switch can utilize resources allocated to each virtual channel (e.g.,Virtual channels are dedicated to buffers, allowing communications on the virtual channels to remain logically independent. The network switch can also employ arbitration switching logic to select which buffered communications should be forwarded across the network. Virtual channels can be channels that physically share a network but are logically independent on the network (e.g., communications on one virtual channel do not block the progress of communications on another virtual channel).

[0012] An agent can generally be any device (e.g., processor, peripheral, memory controller, etc.) capable of providing and / or receiving communications on a network. A source agent initiates (provides) a communication, and a destination agent receives (receives) the communication. A given agent can be a source agent for some communications and a destination agent for other communications.

[0013] Now, referring to the characters, Fig. 1 A generic diagram illustrating physically and logically independent networks. Fig. Figures 2-4 are examples of different network topologies. Fig. 5 is an example of a SOC with a large number of physically and logically independent networks. Fig. Figures 6-8 illustrate the different networks of Fig. 5. Separated for added clarity. Fig. Figure 9 is a block diagram of a system, including two semiconductor dies, illustrating the scalability of the networks to multiple instances of the SOC. Fig. 10 and Fig. Eleven are exemplary agents shown in more detail. Fig. 12 shows different virtual channels and communication types and which networks are in Fig. 5, to which the virtual channels and communication types are applicable. Fig. 13 is a flowchart that illustrates a process, and Fig. 14 and Fig. Figure 15 are exemplary embodiments of a system and a computer-accessible storage medium. The following description provides further details based on the drawings.

[0014] Fig. Figure 1 is a block diagram of a system that includes an embodiment of multiple network interconnection agents. Fig. Figure 1 illustrates agents 10A, 10B, and 10C, although any number of agents may be included in different embodiments. Agents 10A-10B are coupled to a network 12A, and agents 10A and 10C are coupled to a network 12B. Any number of networks 12A-12B may also be included in different embodiments. Network 12A includes a plurality of network switches, including network switches 14AA, 14AB, 14AM, and 14AN (collectively, network switches 14A); and similarly, network 12B includes a plurality of network switches, including network switches 14BA, 14BB, 14BM, and 14BN (collectively, network switches 14B). Different networks 12A-12B may include different numbers of network switches 14A-14B. Additionally, networks 12A-12B include physically separate connections (“wires”, “buses” or “twist”) that are in Fig. 1 are illustrated as different arrows.

[0015] Since each network, 12A-12B, has its own physically and logically separate wiring and network switches, networks 12A-12B are physically and logically isolated. Communication on network 12A is unaffected by communication on network 12B, and vice versa. Even the bandwidth at the wiring in the respective networks 12A-12B is separate and independent.

[0016] Optionally, an agent 10A-10C can include or be coupled to a network interface circuit (reference 16A-16C). Some agents 10A-10C can include or be coupled to network interfaces 16A-16C, while others cannot. Network interfaces 16A-16C can be configured to transmit and receive traffic on networks 12A-12B for the corresponding agents 10A-10C. The network interfaces 16A-16C can be configured to convert or modify communications issued by the corresponding agents 10A-10C to conform to the protocol / format of the networks 12A-12B, and to remove modifications or convert received communications into the protocol / format used by the agents 10A-10C.Thus, network interfaces 16A-16C can be used for agents 10A-10C, which are not specifically designed to connect directly to networks 12A-12B. In some cases, an agent 10A-10C can communicate on more than one network (e.g., agent 10A communicates on both networks 12A-12B). Fig. 1) The corresponding network interface 16A can be configured to separate traffic issued by agent 10A to networks 12A-12B according to which network 12A-12B each communication is assigned to; and network interface 16A can be configured to combine traffic received from networks 12A-12B for the corresponding agent 10A. Any mechanism for determining which network 12A-12B should carry a given communication can be used (e.g., based on the type of communication, the destination agent 10B-10C for the communication, an address, etc., in various embodiments).

[0017] Since the network interface circuits are optional and many are not required for agents that directly support networks 12A-12B, the network interface circuits are omitted from the remaining drawings for the sake of simplicity. It is understood, however, that the network interface circuits can be used in any of the illustrated embodiments by any agent, any subset of agents, or even all of the agents.

[0018] In one embodiment, the system can be Fig. 1 be implemented as a SOC, and the one in Fig. The components illustrated in Figure 1 can be formed on a single semiconductor substrate die. The switching logic enclosed in the SOC can include the plurality of agents 10A-10C and the plurality of network switches 14A-14B coupled to the plurality of agents 10A-10C. The plurality of network switches 14A-14B are interconnected to form a plurality of physically and logically independent networks 12A-12B.

[0019] Since networks 12A-12B are physically and logically independent, different networks can have different topologies. For example, a given network can be a ring, a mesh, a tree, a star, a fully interconnected set of network switches (e.g., a switch directly connected to every other switch in the network), a shared bus with multiple agents coupled to the bus, and so on, or hybrids of one or more of these topologies. Each network 12A-12B can employ a topology that provides, for example, the bandwidth and latency attributes desired for that network, or any desired attribute for the network. Thus, the SOC can generally include a first network built according to a first topology and a second network built according to a second topology that differs from the first.

[0020] Fig. Figures 2-4 illustrate exemplary topologies. Fig. Figure 2 is a block diagram of an embodiment of a network that uses a ring topology to couple agents 10A-10C. In the example of Fig. 2 is the ring formed by network switches 14AA-14AH. Agent 10A is connected to network switch 14AA; agent 10B is connected to network switch 14AB, and agent 10C is connected to network switch 14AE.

[0021] In a ring topology, each network switch 14AA-14AH can be connected to any two other network switches 14AA-14AH, and the switches form a ring. Therefore, any network switch 14AA-14AH can reach any other network switch in the ring by transmitting a communication on the ring towards that other network switch. A given communication can pass through one or more intermediate network switches in the ring to reach the destination network switch. When a given network switch 14AA-14AH receives a communication from a neighboring network switch 14AA-14AH on the ring, the given network switch can examine the communication to determine if the destination is an agent 10A-10C to which the given network switch is coupled. If so, the given network switch can terminate the communication and forward it to the agent.If not, the given network switch can forward the communication to the next network switch on the ring (e.g., the other network switch 14AA-14AH that is adjacent to the given network switch and is not the adjacent network switch from which the given network switch received the communication). A network switch adjacent to a given network switch can be a network switch to which the given network switch can directly transmit a communication without the communication passing through any intermediate network switches.

[0022] Fig. Figure 3 is a block diagram of an embodiment of a network that uses a mesh topology to couple agents 10A-10P. As shown in Fig. As shown in Figure 3, the network can include network switches 14AA-14AH. Each network switch 14AA-14AH is coupled to two or more other network switches. For example, network switch 14AA is coupled to network switches 14AB and 14AE; network switch 14AB is coupled to network switches 14AA, 14AF, and 14AC, and so on, as shown in Figure 3. Fig. Figure 3 illustrates this. Thus, different network switches in a mesh network can be coupled to different numbers of other network switches. Furthermore, while the embodiment of Fig. While network 3 has a relatively symmetrical structure, other mesh networks may be asymmetrical, depending, for example, on the various traffic patterns expected to prevail in the network. At each network switch 14AA-14AH, one or more attributes of a received communication can be used to determine the adjacent network switch 14AA-14AH to which the receiving network switch 14AA-14AH transmits the communication (unless an agent 10A-10P, with which the receiving network switch 14AA-14AH is coupled, is the destination of the communication, in which case the receiving network switch 14AA-14AH can terminate the communication on the network and make it available to the destination agent 10A-10P). For example, in one embodiment, the network switches 14AA-14AH can be programmed during system initialization to redirect communications based on various attributes.

[0023] In one embodiment, communications can be rerouted based on the destination agent. The rerouting can be configured to transport communications through the fewest number of network switches (the "shortest path") between the source and destination agents that can be supported in the mesh topology. Alternatively, different communications for a given source agent to a given destination agent can take different paths through the mesh. For example, latency-sensitive communications can be transported over a shorter path, while less critical communications can take a different path to avoid consuming bandwidth on the short path, where the other path might be less congested during use.

[0024] Fig. 3 can be an example of a partially connected mesh: at least some communications can pass through one or more intermediate network switches in the mesh. A fully connected network can have a connection from every network switch to every other network switch, and thus all communications can be transmitted without passing through intermediate network switches. In various embodiments, any degree of interconnection can be used.

[0025] Fig. Figure 4 is a block diagram of an embodiment of a network that uses a tree topology to couple the agents 10A-10E. The network switches 14AA-14AG are connected to each other in this example to form the tree. The tree is a form of hierarchical network in which edge network switches (e.g., 14AA, 14AB, 14AC, 14AD, and 14AG in Figure 4) are connected to each other. Fig. 4), which couple with agents 10A-10E, and intermediate network switches (e.g. 14AE and 14AF in Fig. 4) that only couple with other network switches. A tree network can be used, for example, when a particular agent is frequently a destination for communications initiated by other agents, or is frequently a source agent for communications. Thus, for example, the tree network of Fig. 4. This can be used for Agent 10E, which is a primary source or destination for communications. For example, Agent 10E might be a memory controller, which would frequently be a target for memory transactions.

[0026] There are many other possible topologies that can be used in other implementations. For example, a star topology has a source / destination agent in the "middle" of a network, and other agents can connect to the middle agent directly or through a series of network switches. Like a tree topology, a star topology can be used in a case where the middle agent is frequently a source or destination of communications. A split bus topology can be used, and hybrids of any two or more of these topologies are also possible.

[0027] Fig. Figure 5 is a block diagram of an embodiment of a system-on-a-chip (SOC) 20, which includes multiple networks for one embodiment. In the embodiment of Fig. 5. The SOC 20 includes a variety of processor clusters (P-clusters) 22A-22B, a variety of input / output clusters (I / O clusters) 24A-24D, a variety of memory controllers 26A-26D, and a variety of graphics processing units (GPUs) 28A-28D. As implied by the name (SOC), the components in Fig. The five illustrated components (except for the memories 30A-30D in this embodiment) are integrated on a single semiconductor die or “chip.” However, other embodiments may employ two or more dies coupled or packaged in any desired manner. Additionally, while specific numbers of P-clusters 22A-22B, I / O clusters 24A-24D, memory controllers 26A-26D, and GPUs 28A-28D are shown in the example of Fig. As shown in section 5, the number and arrangement of any of the above components can be varied and can be more or less than those shown. Fig. The number shown is 5. The memory modules 30A-30D are coupled to the SOC 20 and more precisely to the memory controllers 26A-26D, as shown in Fig. 5 shown.

[0028] In the illustrated embodiment, the SOC 20 connects three physically and logically independent networks formed from a plurality of network switches 32, 34 and 36, as shown in Fig. Figure 5 shows a network switch configuration, and an interconnection between them, illustrated by arrows between the network switches and other components. Other embodiments may include more or fewer networks. For example, network switches 32, 34, and 36 may be instances of network switches similar to network switches 14A-14B, as shown above in relation to Fig. 1-4 described. The multitude of network switches 32, 34 and 36 are coupled to the multitude of P-clusters 22A-22B, the multitude of GPUs 28A-28D, the multitude of memory controllers 26A-25B and the multitude of I / O clusters 24A-24D, as described in Fig. Figure 5 shows that the P-clusters 22A-22B, GPUs 28A-28B, memory controllers 26A-26B, and I / O clusters 24A-24D can all be examples of agents communicating on the various networks of SOC 20. Other agents can be included as desired.

[0029] In Fig. Figure 5 is a central processing unit (CPU) network consisting of a first subset of the multitude of network switches (e.g., network switches 32) and interconnections between them, illustrated by short / long dashed lines, as in reference 38. The CPU network connects the P clusters 22A-22B and the memory controllers 26A-26D. An I / O network is formed from a second subset of the multitude of network switches (e.g., network switches 34) and interconnections between them, illustrated by solid lines, as in reference 40. The I / O network connects the P clusters 22A-22B, the I / O clusters 24A-24D, and the memory controllers 26A-26B. A network with loose sorting is formed from a third subset of the multitude of network switches (e.g., the network switches 36) and a connection between them, illustrated as short dashed lines, such as reference sign 42.The loosened-sort network couples GPUs 28A-28D and memory controllers 26A-26D. In one embodiment, the loosened-sort network can also couple selected I / O clusters 24A-24D. As mentioned above, the CPU network, the I / O network, and the loosened-sort network are independent of each other (e.g., logically and physically independent). In one embodiment, the protocol on the CPU network and the I / O network support cache coherence (e.g., the networks are coherent). The loosened-sort network may not support cache coherence (e.g., the network is incoherent). The loosened-sort network also has reduced sorting constraints compared to the CPU network and the I / O network. For example, in one embodiment, a set of virtual channels and subchannels within the virtual channels is defined for each network.For CPU and I / O networks, communications between the same source and destination agents, and within the same virtual channel and subchannel, can be sorted. For the network with looser sorting, communications between the same source and destination agents can be sorted. In one embodiment, only communications to the same address (given a certain granularity, such as a cache block) between the same source and destination agents can be sorted. Because the network with looser sorting enforces less strict sorting, higher bandwidth can be achieved on average. This is because, for example, transactions can be completed in a different order if newer transactions are ready to complete before older ones.

[0030] The interconnection between network switches 32, 34, and 36 can have any form and configuration in various embodiments. For example, in one embodiment, the interconnection can consist of unidirectional point-to-point links (e.g., buses or serial links). Packets can be transmitted on the links, and the packet format can include data specifying the virtual channel and subchannel in which a packet travels, a memory address, source and destination agent identifiers, data (if applicable), and so on. Several packets can constitute a given transaction. A transaction can be a complete communication between a source agent and a destination agent.For example, depending on the protocol, a read transaction might include a read request packet from the source agent to the destination agent, one or more coherence message packets between the caching agent and the destination agent and / or source agent (if the transaction is coherent), a data response packet from the destination agent to the source agent, and possibly a completion packet from the source agent to the destination agent. A write transaction might include a write request packet from the source agent to the destination agent, one or more coherence message packets as in a read transaction (if the transaction is coherent), and possibly a completion packet from the destination agent to the source agent. The write data might be included in the write request packet in one embodiment, or it might be transferred from the source agent to the destination agent in a separate write data packet.

[0031] The arrangement of agents in Fig. In one embodiment, 5 can specify the physical arrangement of agents on the semiconductor die that forms the SOC 20. That is to say, Fig. 5 can be considered the surface area of ​​the semiconductor die, and the locations of various components in Fig. 5 their physical locations can be approximated by the area. Thus, for example, the I / O clusters 24A-24D can be located in the semiconductor die area that is defined by the top of the SOC 20 (as in Fig. 5 aligned) is shown. The P-clusters 22A-22B can be located in the area shown by the section of SOC 20 below and between the arrangement of I / O clusters 24A-24D, as shown in Fig. 5. The GPUs 24A-28D can be arranged in a central orientation and oriented towards the area shown by the underside of the SOC 20 as shown in Fig. 5 aligned extend. The memory controllers 26A-26D can be located on the areas represented by the right and left sides of the SOC 20, as shown in Fig. 5 should be arranged in a straight line.

[0032] In one embodiment, the SOC 20 can be configured to couple directly with one or more other instances of the SOC 20, thereby creating a logically coupled network on the instances where an agent on one die can communicate logically with an agent on another die in the same way that the agent communicates within another agent on the same die. While the latency may differ, the communication can be performed in the same way. Thus, as in Fig. As illustrated in Figure 5, the networks extend to the underside of SOC 20, as shown in Fig. 5 aligned. Interface switching logic (e.g. serializer / deserializer circuits (SERDES circuits)) that are in Fig. Figure 5, not shown, can be used to communicate with another die across the die boundary. Thus, the networks can be scalable to two or more semiconductor dies. For example, the two or more semiconductor dies can be configured as a single system, in which the presence of multiple semiconductor dies is transparent to software running on the single system. In one embodiment, as a aspect of software transparency for the multi-die system, the delays in die-to-die communication can be minimized, such that die-to-die communication typically exhibits no significant additional latency compared to intra-die communication. In other embodiments, the networks can be closed networks that communicate only intra-die.

[0033] As mentioned above, different networks can have different topologies. In the embodiment of Fig. For example, in version 5, the CPU and I / O networks implement a ring topology, and the loosened sorting can implement a mesh topology. However, other topologies can be used in other embodiments. Fig. 6, Fig. 7 and Fig. 8 illustrate sections of the SOC 30, which include the various networks: CPU ( Fig. 6), I / O ( Fig. 7) and relaxed sorting ( Fig. 8). As in Fig. 6 and Fig. As shown in Figure 7, network switches 32 and 34 form a ring when paired with the corresponding switches on another die. If only a single die is used, a connection can be established between the two network switches 32 or 34 on the underside of the SOC 20, as shown in Figure 7. Fig. 6 and Fig. 7 aligned (e.g., via an external connection to the pins of the SOC 20). Alternatively, the two network switches 32 or 34 on the underside can have links between them, which can be used in a single-die configuration, or the network can operate with a daisy-chain topology.

[0034] Similarly, in Fig. Figure 8 shows the connection of the network switches 36 in a mesh topology between the GPUs 28A-28D and the memory controllers 26A-26D. As mentioned previously, in one embodiment, one or more of the I / O clusters 24A-24D can also be coupled to the loosely collated network. For example, the I / O clusters 24A-24D, which include video peripherals (e.g., a display controller, a memory scaler / rotator, a video encoder / decoder, etc.), can have access to the loosely collated network for video data.

[0035] The network switches 36 are located near the bottom of the SOC 30, as shown in Fig. Aligned with 8, connections can include those that can be redirected to another instance of SOC 30, thus allowing the mesh network to span multiple dies, as discussed above with respect to the CPU and I / O networks. In a single-die configuration, paths extending out of the die cannot be used. Fig. Figure 9 is a block diagram of a two-die system where each network spans the two SOC dies 20A-20B, thus forming networks that are logically identical despite spanning two dies. Network switches 32, 34, and 36 have been omitted for simplicity. Fig. 9 removed, and the loosely sorted network was simplified to a line, although in one embodiment it can be a mesh. The I / O network 44 is shown as a solid line, the CPU network 46 is shown as an alternating long and short dashed line, and the loosely sorted network 48 is shown as a dashed line. The ring structure of networks 44 and 46 is also shown in Fig. 9 evident. While two this in Fig. As shown in Figure 9, other embodiments can employ more than two dies. The networks can be connected to each other via daisy-chaining in various embodiments, be fully connected with point-to-point links between each die pair, or have any other connection structure.

[0036] In one embodiment, the physical separation of the I / O network from the CPU network can help the system provide low-latency memory access through processor clusters 22A-22B, as I / O traffic can be offloaded to the I / O network. The networks use the same memory controllers for memory access, so the memory controllers can be configured to favor memory traffic from the CPU network over memory traffic from the I / O network to some extent. Processor clusters 22A-22B can also be part of the I / O network to access device space in I / O clusters 24A-24D (e.g., using programmed input / output (PIO) transactions). However, memory transactions initiated by processor clusters 22A-22B can be transmitted over the CPU network.Thus, CPU clusters 22A-22B can be examples of an agent coupled to at least two of the multitude of physically and logically independent networks. The agent can be configured to generate a transaction to be transferred and to select one of the at least two of the multitude of physically and logically independent networks over which the transaction should be transferred, based on a type of transaction (e.g., memory or PIO).

[0037] Different networks can include different numbers of physical and / or virtual channels. For example, the I / O network can have multiple request channels and termination channels, while the CPU network can have one request channel and one termination channel (or vice versa). The requests transmitted on a given request channel, when more than one is present, can be determined in any desired way (e.g., by request type, request priority, to balance bandwidth across physical channels, etc.). Similarly, the I / O and CPU networks can include a virtual snoop channel to transmit snoop requests, but the loosely sorted network might not include the virtual snoop channel because it is incoherent in this configuration.

[0038] Fig. Figure 10 is a block diagram of an embodiment of an input / output cluster (I / O cluster) 24A, which is illustrated in more detail. Other I / O clusters 24B-24D may be similar. In the embodiment of Fig. The I / O cluster 24A includes peripheral devices 50 and 52, a peripheral device interface controller 54, a local interconnect 56, and a bridge 58. The peripheral device 52 can be coupled to an external component 60. The peripheral device interface controller 54 can be coupled to a peripheral device interface 62. The bridge 58 can be coupled to a network switch 34 (or to a network interface that couples to the network switch 34).

[0039] Peripherals 50 and 52 can include any set of additional hardware functionality (e.g., beyond CPUs, GPUs, and memory controllers) that is included in the SoC 20. For example, peripherals 50 and 52 can include video peripherals, such as an image signal processor configured to process image capture data from a camera or other image sensor, video encoders / decoders, scalers, rotators, mixers, a display controller, etc. The peripherals can include audio peripherals, such as microphones, speakers, microphone and speaker interfaces, audio processors, digital signal processors, mixers, etc. The peripherals can include network peripherals, such as media access controllers (MACs). The peripherals can include other types of memory controllers, such as non-volatile memory controllers.Some peripheral devices 52 may include an in-chip component and an external component 60. The peripheral device interface controller 54 may include interface controllers for various interfaces 62 that are outside the SOC 20, including interfaces such as Universal Serial Bus (USB), Peripheral Component Interconnect (PCI), including PCI Express (PCIe), serial and parallel ports, etc.

[0040] The local interconnection 56 can be an interconnection on which the various peripheral devices 50, 52, and 54 communicate. The local interconnection 56 can differ from the one in Fig. The system-wide interconnection shown in Figure 5 (e.g., of the CPU, I / O, and loosened networks) can be distinguished. Bridge 58 can be configured to convert communications on the local interconnection to communications on the system-wide interconnection and vice versa. In one embodiment, Bridge 58 can be coupled to one of the network switches 34. Bridge 58 can also manage sorting among the transactions issued by peripherals 50, 52, and 54. For example, Bridge 58 can use a cache coherence protocol supported on the networks to ensure the sorting of transactions for peripherals 50, 52, and 54, and so on. Different peripherals 50, 52, and 54 may have different sorting requirements, and Bridge 58 can be configured to adapt to these different requirements.In some embodiments, the bridge 58 can also implement various performance-enhancing features. For example, the bridge 58 can pre-retrieve data for a given request. The bridge 58 can capture a coherent copy of a cache block (e.g., in the exclusive state) to which one or more transactions are directed from the peripherals 50, 52, and 54, to allow the transactions to complete locally and to enforce sorting. The bridge 58 can speculatively capture an exclusive copy of one or more cache blocks to which subsequent transactions are directed and can use the cache block to complete the subsequent transactions if the exclusive state is successfully maintained until the subsequent transactions can be completed (e.g., after sorting constraints with earlier transactions have been satisfied).Thus, in one embodiment, multiple requirements within a cache block can be met by the cached copy. Further details can be found in preliminary US patent applications Nos. 63 / 170,868, filed April 5, 2021, 63 / 175,868, filed April 16, 2021, and 63 / 175,877, filed April 16, 2021. These patent applications are incorporated herein by reference in their entirety. To the extent that any of the incorporated material conflicts with the material expressly set forth herein, the material expressly set forth herein shall prevail.

[0041] Fig. Figure 11 is a block diagram of a particular embodiment of a processor cluster 22A. Other embodiments may be similar. In the embodiment of Fig. 10. The processor cluster 22A includes one or more processors 70 coupled to a last-level cache (LLC) 72. The LLC 72 may include interface switching logic for connecting to the network switches 32 and 34 to transfer transactions on the CPU network and the I / O network as appropriate.

[0042] The Processors 70 can include any switching logic and / or microcode configured to execute instructions defined in an instruction set architecture implemented by the Processors 70. The Processors 70 can have any microarchitecture implementation, any performance level, and any performance characteristics, etc. For example, Processors 70 can be in-order execution, out-of-order execution, superscalar, superpipelined, etc.

[0043] The LLC 72 and all caches within the processors 70 can have any capacity and configuration, such as set-associative, directly mapped, or fully associative. The cache block size can be any desired size (e.g., 32 bytes, 64 bytes, 128 bytes, etc.). The cache block can be the unit of mapping and unmapping in the LLC 70. Additionally, the cache block can be the unit over which coherence is maintained in this embodiment. The cache block can also be referred to as a cache row in some cases. In one embodiment, a distributed, directory-based coherence scheme can be implemented with a coherence point at each memory controller 26 in the system, the coherence point being valid for memory addresses mapped to that memory controller. The directory can track the state of cache blocks.which are cached in any coherent agent. The coherence scheme can be scalable to many memory controllers across potentially multiple semiconductor dies. For example, the coherence scheme can employ one or more of the following features: exact directory for snoop filtering and race resolution at coherent and memory agents; sort point (access order) determined at the memory agent, serialization point migrated between coherent and memory agents; secondary collection of completion (invalidation acknowledgment) at the requesting coherent agent, tracked with completion count provided by a memory agent; fill / snoop and snoop / victim acknowledgment race resolution, handled at the coherent agent by directory state.which is provided by the storage agent; distinct primary / secondary shared states to support race resolution and limit in-flight snoops to the same address / destination; absorption of conflicting snoops on a coherent agent to avoid a deadlock without additional nack / conflict / retry messages or actions; serialization minimization (one additional message latency per accessor to transfer ownership through a conflict chain); message minimization (messages directly between relevant agents and no additional messages to handle conflicts / races (e.g., no messages back to a storage agent)); store-conditional without over-invalidation on race failures; exclusive ownership request with intent,to modify the entire cache line with minimized data transmission (only in the dirty case) and associated cache / directory states; distinct snoop-back and snoop-forward message types to handle both cacheable and non-cacheable streams (e.g., 3-hop and 4-hop protocols). Additional details can be found in the preliminary US patent application No. 63 / 077,371, filed on September 11, 2020. That patent application is incorporated herein by reference in its entirety. To the extent that any of the incorporated material conflicts with the material expressly set forth herein, the material expressly set forth herein shall prevail.

[0044] Fig. 12 is a pair of Tables 80 and 82, which illustrate virtual channels and traffic types and the networks that are in Fig. Figures 5 to 8 show their use in one embodiment. As shown in Table 80, the virtual channels can include the virtual bulk channel, the virtual low-latency channel (LLT), the real-time (virtual RT channel), and the virtual non-DRAM message (VCP) channel. The virtual bulk channel can be the default virtual channel for memory accesses. The virtual bulk channel can, for example, receive a lower quality of service than the virtual LLT and RT channels. The virtual LLT channel can be used for memory transactions that require low latency for high-performance operation. The virtual RT channel can be used for memory transactions that have latency and / or bandwidth requirements for proper operation (e.g., video streams).The VCP channel can be used to separate traffic that is not directed to memory, in order to prevent interference with memory transactions.

[0045] In one embodiment, the virtual mass and LLT channels can be supported on all three networks (CPU, I / O, and loosened sorting). The virtual RT channel can be supported on the I / O network but not on the CPU or loosened sorting networks. Similarly, the virtual VCP channel can be supported on the I / O network but not on the CPU or loosened sorting networks. In one embodiment, the virtual VCP channel on the CPU and loosened sorting networks can be supported only for transactions directed to the network switches on that network (e.g., for configuration) and thus cannot be used during normal operation. Therefore, as illustrated in Table 80, different networks can support different numbers of virtual channels.

[0046] Table 82 illustrates different traffic types and which networks carry each traffic type. Traffic types can include coherent memory traffic, non-coherent memory traffic, real-time memory traffic (RT memory traffic), and VCP (non-memory) traffic. The CPU and I / O networks can both carry coherent traffic. In one embodiment, coherent memory traffic provided by processor clusters 22A-22B can be carried on the CPU network, while the I / O network can carry coherent memory traffic provided by I / O clusters 24A-24D. Non-coherent memory traffic can be carried on the loosely sorted network, and RT and VCP traffic can be carried on the I / O network.

[0047] Fig. Figure 13 is a flowchart illustrating one embodiment of a procedure for initiating a transaction on a network. In one embodiment, an agent can generate a transaction to be transferred (Block 90). The transaction is to be transferred on one of a plurality of physically and logically independent networks. A first network of the plurality of physically and logically independent networks is built according to a first topology, and a second network of the plurality of physically and logically independent networks is built according to a second topology that differs from the first. One of the plurality of physically and logically independent networks is selected on which the transaction is to be transferred based on a transaction type (Block 92). For example, processor clusters 22A-22B can transfer coherent memory traffic on the CPU network and PIO traffic on the I / O network.In one embodiment, the agent can select a virtual channel from a variety of virtual channels supported on the selected network of the variety of physically and logically independent networks based on one or more attributes of the transaction that differs from the type (Block 94). For example, a CPU can select the LLT virtual channel for a subset of memory transactions (e.g., the oldest memory transactions that are cache misses, or a number of cache misses up to a threshold number, after which the bulk channel can be selected). A GPU can select between the LLT and bulk virtual channels based on the urgency with which the data is needed. Video devices can use the RT virtual channel as needed (e.g., the display controller can output frame data read operations on the RT virtual channel).The virtual VCP channel can be selected for transactions that are not memory transactions. The agent can transmit a transaction packet across the selected network and virtual channel. In one embodiment, transaction packets in different virtual channels can take different paths through the networks. In another embodiment, transaction packets can take different paths based on the type of transaction packet (e.g., request versus response). In yet another embodiment, different paths can be supported for both different virtual channels and different types of transactions. Other embodiments can use one or more additional attributes of transaction packets to determine a path through the network for these packets.Viewed from another perspective, network switches constitute the network and can redirect packets that differ based on the virtual channel, type, or any other attributes. A different path can refer to traversing at least one segment between network switches that is not traversed by the other path, even though the transaction packets move from the same source to the same destination using the different paths. The use of different paths can provide load balancing in the networks and / or reduced latency for transactions. computer system

[0048] Next, referring to Fig. Figure 14 shows a block diagram of an embodiment of a system 700. In the illustrated embodiment, the system 700 includes at least one instance of a system-on-a-chip (SOC) 20 coupled with one or more peripheral devices 704 and an external memory 702. A power supply unit (PMU) 708 is provided, which supplies voltages to the SOC 10 and one or more voltages to the memory 702 and / or the peripheral devices 154. In some embodiments, more than one instance of the SOC 20 may be included (and more than one memory 702 may also be included). In one embodiment, the memory 702 may be the Fig. Include the 5 illustrated memory locations 30A-30D.

[0049] The 704 peripherals can include any desired switching logic, depending on the type of system 700. For example, in one embodiment, the system 704 can be a mobile device (e.g., a personal digital assistant (PDA), a smartphone, etc.), and the 704 peripherals can include devices for various types of wireless communication, such as Wi-Fi, Bluetooth, cellular networks, global positioning systems, etc. The 704 peripherals can also include additional storage, including RAM, solid-state storage, or disk storage. The 704 peripherals can include user interface devices, such as a display screen, including touchscreens or multi-touchscreens, keyboards or other input devices, microphones, speakers, etc.In other embodiments, the System 700 can be any type of computing system (e.g., desktop personal computer, laptop, workstation, nettop, etc.).

[0050] The 702 external memory can include any type of memory. For example, the 702 external memory can be SRAM, dynamic RAM (DRAM), such as synchronous DRAM (SDRAM), double-speed SDRAM (DDR, DDR2, DDR3, etc.), RAMBUS DRAM, lower-power versions of DDR DRAM (e.g., LPDDR, mDDR, etc.), and so on. The 702 external memory can include one or more memory modules to which the memory devices are attached, such as single-row memory modules (SIMMs), dual-row memory modules (DIMMs), and so on. Alternatively, the 702 external memory can include one or more memory devices attached to the SOC 20 in a chip-on-chip or package-on-package implementation.

[0051] As illustrated, System 700 is shown to have applications in a wide range of fields. For example, System 700 can be used as part of the chips, switching logic, components, etc., of a desktop computer 710, laptop computer 720, tablet computer 730, cellular or mobile phone 740, or television 750 (or a set-top box coupled with a television). Also illustrated are a smartwatch and a health monitoring device 760. In some embodiments, a smartwatch can include a variety of general-purpose computing functions. For example, a smartwatch can provide access to email, a cellular mobile service, a user calendar, and so on. In various embodiments, a health monitoring device can be a dedicated medical device or otherwise include dedicated health-related functionality.For example, a health monitoring device can monitor a user's vital signs, track a user's proximity to other users for the purpose of maintaining epidemiological distance, track contacts, facilitate communication with an emergency service in the event of a health emergency, and so on. In various embodiments, the aforementioned smartwatch may include some or no health monitoring-related functions. Other wearable devices are also being considered, such as devices worn around the neck, devices that can be implanted in the human body, glasses designed to provide an augmented and / or virtual reality experience, and so forth.

[0052] The System 700 can also be used as part of a cloud-based service (770). For example, the devices mentioned above and / or other devices can access computing resources in the cloud (i.e., remotely located hardware and / or software resources). Furthermore, the System 700 can be used in one or more other devices within a dwelling besides those mentioned above. For example, devices within the dwelling can monitor and detect conditions that require attention. For example, various devices within the dwelling (e.g., a refrigerator, a cooling system, etc.) can monitor the device's status and provide an alert to the dwelling owner (or, for example, a repair facility) should a specific event be detected.Alternatively, a thermostat can monitor the temperature in the apartment and automate settings of a heating / cooling system based on a history of responses to different conditions by the apartment owner. Also in . Fig. Figure 14 illustrates the application of System 700 to various modes of transportation. For example, System 700 can be used in the control and / or entertainment systems of airplanes, trains, buses, rental cars, passenger vehicles, watercraft ranging from private boats to cruise ships, (rental or owned) scooters, and so on. In various cases, System 700 can be used to provide automated guidance (e.g., self-driving vehicles), general system control, and other things. Any number of these other embodiments are possible and are being considered. It should be noted that the in Fig. The 14 illustrated devices and applications are for illustrative purposes only and are not intended to be limiting. Other devices are possible and will be considered. Computer-readable storage medium

[0053] Now, referring to Fig. Figure 15 shows a block diagram of an embodiment of a computer-readable storage medium 800. In general terms, a computer-accessible storage medium can include any storage medium that can be accessed by a computer during use to provide instructions and / or data to the computer. For example, a computer-accessible storage medium can include storage media such as magnetic or optical media, e.g., disks (fixed or portable), tapes, CD-ROM, DVD-ROM, CD-R, CD-RW, DVD-R, DVD-RW, or Blu-ray. Storage media can further include volatile or non-volatile storage media, such as RAM (e.g., synchronous dynamic RAM (SDRAM), Rambus DRAM (RDRAM), static RAM (SRAM), etc.), ROM, or flash memory. The storage media can be physically enclosed within the computer to which the storage media provide instructions / data.Alternatively, the storage media can be connected to the computer. For example, the storage media can be connected to the computer via a network or a wireless connection, such as network storage. The storage media can be connected via a peripheral interface, such as the Universal Serial Bus (USB). Generally, the computer-accessible storage medium can store 800 data in a non-transitory manner, where non-transitory in this context can refer to the fact that the instructions / data are not transmitted on a signal. For example, non-transitory storage can be volatile (and lose the stored instructions / data in response to a shutdown), or it can be non-volatile.

[0054] The computer-accessible storage medium 800 in Fig. A database 804, representative of SOC 20, can be stored. In general, database 804 can be a database that can be read by a program and used directly or indirectly to construct the hardware comprising SOC 20. For example, the database can be a behavioral-level or register-transfer (RTL) description of the hardware functionality in a high-level design (HDL) language such as Verilog or VHDL. The description can be read by a synthesis tool, which can synthesize it to generate a netlist containing a list of gates from a synthesis library. The netlist includes a set of gates that also represent the functionality of the hardware comprising SOC 20. The netlist can then be placed and routed to generate a dataset describing geometric shapes to be applied to masks.The masks can then be used in various semiconductor fabrication steps to produce a semiconductor circuit or circuits that conform to the SOC 20. Alternatively, the database 804 on the computer-accessible storage medium 800 can be, as desired, the netlist (with or without a synthesis library) or the data set.

[0055] While the computer-accessible storage medium 800 stores a representation of the SOC 10, other embodiments can store a representation of any section of the SOC 20, as desired, including any subset of the data contained therein. Fig. The 5 components shown are carried. Database 804 can represent any section of the foregoing.

[0056] In one embodiment, a system comprises a plurality of processor clusters, a plurality of memory controllers, a plurality of graphics processing units (GPUs), a plurality of agents, and a plurality of network switches coupled to the plurality of processor clusters, GPUs, memory controllers, and agents. A given processor cluster comprises one or more processors. The memory controllers are configured to control access to memory devices. A first subset of the plurality of network switches is interconnected to form a central processing unit (CPU) network between the plurality of processor clusters and the plurality of memory controllers.A second subset of the multitude of network switches is interconnected to form an input / output (I / O) network between the multitude of processor clusters, the multitude of agents, and the multitude of memory controllers. A third subset of the multitude of network switches is interconnected to form a loosely collated network between the multitude of graphics processing units, selected agents, and the multitude of memory controllers. The CPU network, the I / O network, and the loosely collated network are independent of each other. The CPU network and the I / O network are coherent. The loosely collated network is incoherent and has reduced collation constraints compared to the CPU network and the I / O network.In one embodiment, at least one of the CPU network, the I / O network, and the loosely sorted network has a number of physical channels that differs from the number of physical channels on another of the CPU network, the I / O network, and the loosely sorted network. In one embodiment, the CPU network is a ring network. In one embodiment, the I / O network is a ring network. In one embodiment, the loosely sorted network is a mesh network. In one embodiment, a first agent of the plurality of agents comprises an I / O cluster comprising a plurality of peripheral devices. In one embodiment, the I / O cluster further comprises a bridge coupled to the plurality of peripheral devices and also coupled to a first network switch in the second subset.In one embodiment, the system further comprises a network interface circuit configured to convert communications from a given agent into communications for a given network of a CPU network, the I / O network, and the loosened-sort network, wherein the network interface circuit is coupled to one of the plurality of network switches in the given network.

[0057] In one embodiment, a system-on-a-chip (SOC) comprises a semiconductor die on which switching logic is implemented. The switching logic includes a plurality of agents and a plurality of network switches coupled to the plurality of agents. The plurality of network switches are interconnected to form a plurality of physically and logically independent networks. A first network of the plurality of physically and logically independent networks is constructed according to a first topology, and a second network of the plurality of physically and logically independent networks is constructed according to a second topology that differs from the first topology. In one embodiment, the first topology is a ring topology. In another embodiment, the second topology is a mesh topology. In one embodiment, coherence is enforced on the first network.In one embodiment, the second network is a network with loose sorting. In another embodiment, at least one of the plurality of physically and logically independent networks implements a first number of physical channels, and at least one other of the plurality of physically and logically independent networks implements a second number of physical channels, the first number being different from the second number. In another embodiment, the first network includes one or more first virtual channels, and the second network includes one or more second virtual channels. At least one of the first virtual channels is different from the second virtual channels.In one embodiment, the SOC further comprises a network interface circuit configured to convert communications from a given agent of the plurality of agents into communications for a given network of the plurality of physically and logically independent networks. The network interface circuit is coupled to one of the plurality of network switches in the given network. In one embodiment, a first agent of the plurality of agents is coupled to at least two of the plurality of physically and logically independent networks. The first agent is configured to generate a transaction to be transmitted. The first agent is configured to select one of the at least two of the plurality of physically and logically independent networks over which the transaction is to be transmitted, based on a transaction type.In one embodiment, one of the at least two networks is an I / O network on which I / O transactions are transferred.

[0058] In one embodiment, a method comprises generating a transaction in an agent coupled to a plurality of physically and logically independent networks, wherein a first network of the plurality of physically and logically independent networks is constructed according to a first topology, and a second network of the plurality of physically and logically independent networks is constructed according to a second topology that differs from the first topology; and selecting one of the plurality of physically and logically independent networks on which the transaction is to be transmitted based on a transaction type. In one embodiment, the method further comprises selecting, based on one or more transaction attributes that differ from the transaction type, a virtual channel from a plurality of virtual channels supported on the one of the plurality of physically and logically independent networks.

[0059] The present disclosure includes references to “an” embodiment or groups of “embodiments” (e.g., “some embodiments” or “various embodiments”). Embodiments are different implementations or instances of the disclosed concepts. References to “embodiment,” “an embodiment,” “a particular embodiment,” and the like do not necessarily refer to the same embodiment. A large number of possible embodiments are considered, including those specifically disclosed, as well as modifications or alternatives that fall within the nature or scope of the disclosure.

[0060] This disclosure may discuss potential advantages that may arise from the disclosed embodiments. Not all implementations of these embodiments necessarily exhibit all potential advantages. Whether an advantage is achieved for a particular implementation depends on many factors, some of which are outside the scope of this disclosure. Indeed, there are several reasons why an implementation that falls within the scope of the claims may not exhibit some or all of the disclosed advantages. For example, a particular implementation might include different switching logic outside the scope of the disclosure, which, in conjunction with one of the disclosed embodiments, eliminates or reduces one or more of the disclosed advantages. Furthermore, a suboptimal design implementation of a particular implementation (e.g.,Implementation techniques or tools) negate or diminish disclosed benefits. Even assuming a qualified implementation, the attainment of benefits may still depend on other factors, such as the environmental circumstances in which the implementation is deployed. For example, inputs provided to a particular implementation may prevent one or more problems addressed in this disclosure from occurring on a particular occasion, thereby potentially preventing the benefit of its solution from being achieved. Due to the existence of possible factors outside of this disclosure, it is expressly intended that all potential benefits described herein are not to be construed as limitations on claims that must be satisfied to prove infringement.Rather, the identification of such potential benefits is intended to illustrate the type(s) of improvement available to designers who benefit from this disclosure. Describing such benefits in a permissive manner (e.g., stating that a particular benefit “may occur”) is not intended to cast doubt on whether such benefits can actually be achieved, but instead to acknowledge the technical reality that achieving such benefits often depends on additional factors.

[0061] Unless otherwise stated, embodiments are not limiting. This means that the disclosed embodiments are not intended to limit the scope of protection of claims formulated on the basis of this disclosure, even if only a single example is described with respect to a particular feature. The disclosed embodiments are intended to be illustrative and not limiting, unless the disclosure contains statements to the contrary. The application is thus intended to allow claims to cover disclosed embodiments as well as the alternatives, modifications, and equivalents that are obvious to a person skilled in the art who benefits from this disclosure.

[0062] For example, features in this application may be combined in any suitable way. Accordingly, during the further pursuit of this application (or an application claiming priority thereof), new claims may be formulated to any such combination of features. In particular, with reference to the accompanying claims, features from dependent claims may be combined with those of other dependent claims where appropriate, including claims dependent on other independent claims. Similarly, features from respective independent claims may be combined where appropriate.

[0063] While the accompanying dependent claims may be formulated such that each depends on a single other claim, additional dependencies are also considered. All combinations of features in the dependent claim that are consistent with this disclosure are considered in this disclosure and may be claimed in this or any other application. In summary, combinations are not limited to those specifically enumerated in the accompanying claims.

[0064] Where appropriate, consideration is also given to the possibility that claims formulated in one format or statutory type (e.g. establishment) should support corresponding claims of another format or statutory type (e.g. procedure).

[0065] Since this disclosure is a legal document, various terms and phrases may be subject to regulatory and legal interpretation. It is hereby announced that the following paragraphs, as well as definitions provided in the disclosure, shall be used in determining how claims formulated based on this disclosure are to be interpreted.

[0066] References to a singular form of an element (i.e., a noun or noun phrase preceded by "a" or "the") should, unless the context clearly indicates otherwise, mean "one or more." Thus, a reference to "an element" in a claim, without accompanying context, does not exclude additional instances of the element. A "multitude" of elements refers to a set of two or more of the elements.

[0067] The word “can / can” is used here in a permissive sense (i.e. having the potential to be able to) and not in an obligatory sense (i.e. must / must).

[0068] The terms “comprehensive” and “inclusive” and forms thereof are open and mean “including without being limited to”.

[0069] When the term “or” is used in this revelation in reference to a list of options, it is generally understood to be used in an inclusive sense, unless the context indicates otherwise. Thus, a statement of “x or y” is equivalent to “x or y or both” and therefore covers 1) x but not y, 2) y but not x, and 3) both x and y. On the other hand, a phrase such as “either x or y, but not both” makes it clear that “or” is used in an exclusive sense.

[0070] A statement of "w, x, y, or z, or any combination thereof" or "at least one of ... w, x, y, and z" is intended to cover all possibilities, from a single element to the total number of elements in the sentence. For example, in the sentence [w, x, y, z], these phrases cover every single element of the sentence (e.g., w, but not x, y, or z), any two elements (e.g., w and x, but not y or z), any three elements (e.g., w, x, and y, but not z), and all four elements. The phrase "at least one of ... x, y, and z" thus refers to at least one element of the sentence [w, x, y, z], thereby covering all possible combinations in this list of elements. This phrase must not be interpreted as requiring the presence of at least one instance of w, at least one instance of x, at least one instance of y, and at least one instance of z.

[0071] Various “labels” may appear before nouns or noun phrases in this disclosure. Unless the context indicates otherwise, different labels used for a feature (e.g., “first circuit,” “second circuit,” “particular circuit,” “given circuit,” etc.) refer to different instances of the feature. Additionally, when applied to a feature, the labels “first,” “second,” and “third” do not imply any type of ordering (e.g., spatial, temporal, logical, etc.) unless otherwise specified.

[0072] The phrase "based on" or is used to describe one or more factors that influence a determination. This term does not exclude the possibility that additional factors may influence the determination. That is, a determination may be based solely on stated factors or on the stated factors as well as other, unspecified factors. Consider the phrase "determine A based on B." This phrase indicates that B is a factor used to determine A or that influences the determination of A. This phrase does not exclude the possibility that the determination of A may also be based on another factor, such as C. This phrase is also intended to cover an embodiment in which A is determined solely based on B. As used herein, the phrase "based on" is synonymous with the phrase "based at least partially on."

[0073] The phrases "in response to" and "in reaction to" describe one or more factors that trigger an effect. This phrase does not exclude the possibility that additional factors may influence or otherwise trigger the effect, either together with the stated factors or independently of them. That is, an effect may occur solely in response to these factors, or it may occur in response to the stated factors as well as other, unspecified factors. Consider the phrase "to perform A in response to B." This phrase indicates that B is a factor that triggers the performance of A or that triggers a particular outcome for A. This phrase does not exclude the possibility that performing A may also occur in response to another factor, such as C. Nor does it exclude the possibility that performing A may occur in response to both B and C.This phrase is also intended to cover an embodiment in which A is carried out solely based on B. As used herein, the phrase "based on" is synonymous with the phrase "based at least partially on". Similarly, the phrase "in response to" is synonymous with the phrase "at least partially in response to".

[0074] Within this disclosure, various entities (which may be variously referred to as "units," "circuits," other components, etc.) may be described or claimed to be "configured" to perform one or more tasks or operations. This phrase—[entity] configured to [perform one or more tasks]—is used herein to refer to a structure (i.e., something physical). In particular, this phrase is used to indicate that this structure is arranged to perform the one or more tasks during operation. A structure may be described as "configured to" perform a task even if the structure is not currently operating.Thus, an entity described or specified as "configured to" perform a task refers to something physical, such as a device, a circuit, a system with a processing unit and memory that stores program instructions executable to implement the task, etc. This phrase is not used here to refer to something intangible.

[0075] In some cases, various units / circuits / components herein may be described as performing a set of tasks or operations. It is understood that these entities are "configured to" perform these tasks / operations, even if this is not specifically stated.

[0076] The term "configured to" should not be interpreted as "configurable to." For example, an unprogrammed FPGA would not be considered "configured to" perform a specific function. However, this unprogrammed FPGA can be "configurable to" perform that function. After appropriate programming, the FPGA can then be described as "configured to" perform the specific function.

[0077] For the purposes of US patent applications based on this disclosure, a claim that a structure is "configured to" perform one or more functions shall expressly not rely on the application of 35 USC § 112(f) for that claim element. If, during the grant proceedings of a US patent application based on this disclosure, the applicant wishes to rely on the application of section 112(f), they shall specify claim elements using the construct "means to" [perform a function].

[0078] Various “circuits” can be described in this disclosure. These circuits, or “switching logic,” constitute hardware that includes various types of circuit elements, such as combinational logic, clocked storage devices (e.g., flip-flops, registers, latches, etc.), finite automata, memory (e.g., random-access memory, embedded dynamic random-access memory), programmable logic arrays, and so on. Switching logic can be custom-designed or taken from standard libraries. In various implementations, switching logic may, as appropriate, include digital components, analog components, or a combination of both. Certain types of circuits may be referred to more generally as “units” (e.g., a decoding unit, an arithmetic logic unit (ALU), a functional unit, a memory management unit (MMU), etc.).Such units also refer to circuits or switching logic.

[0079] The disclosed circuits / components and other elements illustrated in the drawings and described herein thus include hardware elements such as those described in the preceding paragraph. In many cases, the internal arrangement of hardware elements within a particular circuit can be specified by describing the function of that circuit. For example, a particular "decoding unit" can be described as performing the function of "processing an opcode of an instruction and redirecting that instruction to one or more of a plurality of functional units," meaning that the decoding unit is "configured to" perform this function. This functional description is sufficient for a person skilled in the art of computing to further define a set of possible structures for the circuit.

[0080] In various embodiments, as discussed in the preceding paragraph, circuits, units, and other elements defined by the functions or operations they are configured to implement, the arrangement of such circuits / units / components in relation to one another, and the way in which they interact, constitute a microarchitecture definition of the hardware, which is ultimately fabricated in an integrated circuit or programmed into an FPGA to form a physical implementation of the microarchitecture definition. Thus, the microarchitecture definition is recognized by a person skilled in the art as a structure from which many physical implementations can be derived, all of which fall within the broader structure described by the microarchitecture definition.This means that a person skilled in the art, presented with the microarchitecture definition provided according to this disclosure, can implement the structure without undue experimentation and by applying average skills by encoding the description of the circuits / units / components in a hardware description language (HDL), such as Verilog or VHDL. The HDL description is often expressed in a way that may appear functional. However, to a person skilled in the art, this HDL description is the way in which the structure of a circuit, unit, or component is transformed to the next level of implementation detail. Such an HDL description may take the form of behavioral code (which is not usually synthesizable), register-transfer language (RTL) code (which, unlike behavioral code, is usually synthesizable), or structural code (e.g.,a netlist specifying logic gates and their connectivity). The HDL description can then be synthesized against a library of cells designed for a given integrated circuit fabrication technology and modified for timing, power, and other reasons to produce a final design database. This database is then sent to a foundry to generate masks and ultimately fabricate the integrated circuit. Some hardware circuits, or sections thereof, can also be user-designed in a schematic editor and captured in the integrated circuit design along with synthesized switching logic. The integrated circuits can include transistors and other circuit elements (e.g., passive elements such as capacitors, resistors, inductors, etc.) and connect the transistors and circuit elements.Some embodiments can implement multiple integrated circuits coupled together to implement the hardware circuits, and / or discrete elements can be used in some embodiments. Alternatively, the HDL design can be synthesized into a programmable logic array, such as a field-programmable gate array (FPGA), and implemented in the FPGA. This decoupling between the design of a group of circuits and the subsequent low-level implementation of those circuits typically leads to a scenario where the circuit or logic designer never specifies a particular set of structures for low-level implementation beyond a description of what the circuit is configured for, as this process is performed at a different stage of the circuit implementation process.

[0081] The fact that many different low-level combinations of circuit elements can be used to implement the same circuit specification results in a large number of equivalent structures for that circuit. As indicated, these low-level circuits can vary according to changes in the manufacturing technology, the foundry chosen for production, the library of cells provided for a particular project, and so on. In many cases, the choices made by different design tools or methodologies for manufacturing these various implementations can be arbitrary.

[0082] Furthermore, for a single implementation of a particular functional specification of a circuit, it is common to include a large number of devices (e.g., millions of transistors) for a given embodiment. Accordingly, the sheer volume of this information makes it impractical to provide a complete low-level specification of the structure used to implement a single embodiment, let alone the vast array of equivalent possible implementations. For this reason, the present disclosure describes a structure of circuits using the functional shorthand notation commonly employed in industry.

[0083] Numerous variations and modifications become apparent to the person skilled in the art once the foregoing disclosure is fully understood. It is intended that the following claims be interpreted to include all such variations and modifications.

Claims

[1] System, encompassing: a plurality of processor clusters, wherein a given processor cluster comprises one or more processors; a large number of graphics processing units; a variety of storage controllers configured to control access to storage devices; a multitude of agents; and a multitude of network switches coupled to the multitude of processor clusters, the multitude of graphics processing units, the multitude of memory controllers, and the multitude of agents, wherein: a first subset of the multitude of network switches is interconnected to form a network of a central processing unit (CPU) between the multitude of processor clusters and the multitude of memory controllers, a second subset of the multitude of network switches is interconnected to form an input / output (I / O) network between the multitude of processor clusters, the multitude of agents, and the multitude of memory controllers, a third subset of the multitude of network switches is interconnected to form a loosely sorted network between the multitude of graphics processing units, selected agents of the multitude of agents, and the multitude of memory controllers, The CPU network, the I / O network, and the loosely sorted network are independent of each other. the CPU network and the I / O network are coherent and The network with loosened sorting is incoherent and has reduced sorting constraints compared to the CPU network and the I / O network. [2] System according to claim 1, wherein at least one of the CPU network, the I / O network and the loosened sorting network has a number of physical channels that differs from a number of physical channels of another of the CPU network, the I / O network and the loosened sorting network. [3] System according to claim 1, wherein the CPU network is a ring network. [4] System according to claim 1, wherein the I / O network is a ring network. [5] System according to claim 1, wherein the network with loosened sorting is a mesh network. [6] System according to claim 1, wherein a first agent of the plurality of agents comprises an I / O cluster comprising a plurality of peripheral devices. [7] System according to claim 6, wherein the I / O cluster further comprises a bridge coupled to the plurality of peripheral devices and further coupled to a first network switch in the second subset. [8] System according to claim 1, further comprising a network interface circuit configured to handle communications from a given agent of the plurality of agents in communications for a given network consisting of CPU network, I / O network and the network with loosened sorting, wherein the network interface circuit is coupled to one of the plurality of network switches in the given network. [9] System-on-a-Chip (SOC), comprising a semiconductor die on which switching logic is formed, wherein the switching logic comprises: a plurality of processor clusters, wherein a given processor cluster comprises one or more processors; a large number of graphics processing units; a variety of storage controllers configured to control access to storage devices; a plurality of input / output clusters (I / O clusters), wherein a given I / O cluster comprises one or more peripheral devices; and a multitude of network switches coupled to the multitude of processor clusters, the multitude of graphics processing units, the multitude of memory controllers, and the multitude of I / O clusters, wherein: a first subset of the multitude of network switches is interconnected to form a network of a central processing unit (CPU) between the multitude of processor clusters and the multitude of memory controllers, a second subset of the multitude of network switches is interconnected to form an input / output (I / O) network between the multitude of processor clusters, the multitude of I / O clusters, and the multitude of memory controllers, a third subset of the multitude of network switches is interconnected to form a network with loosened sorting between the multitude of graphics processing units, selected I / O clusters of the multitude of I / O clusters and the multitude of memory controllers, The CPU network, the I / O network, and the loosely sorted network are independent of each other. the CPU network and the I / O network are coherent and The network with loosened sorting is incoherent and has reduced sorting constraints compared to the CPU network and the I / O network. [10] SOC according to claim 9, wherein at least one of the CPU network, the I / O network and the loosened sorting network has a number of physical channels that differs from a number of physical channels of another of the CPU network, the I / O network and the loosened sorting network. [11] SOC according to claim 9, wherein the CPU network is a ring network. [12] SOC according to claim 9, wherein the I / O network is a ring network. [13] SOC according to claim 9, wherein the loosely sorted network is a mesh network. [14] SOC according to claim 9, wherein the given I / O cluster further comprises a bridge coupled to the plurality of peripheral devices and further coupled to a first network switch in the second subset. [15] SOC according to claim 9, further comprising a network interface circuit configured to convert communications from a given agent into communications for a given network consisting of CPU network, I / O network and the loosened sorting network, wherein the network interface circuit is coupled to one of the plurality of network switches in the given network. [16] SOC according to claim 9, wherein: the CPU network includes one or more first virtual channels and the I / O network includes one or more second virtual channels; and at least one of the one or more second virtual channels differs from the one or more first virtual channels. [17] SOC according to claim 9, wherein a first processor cluster of the plurality of processor clusters is configured to generate a transaction to be transferred, and wherein the first processor cluster is configured to select one of the CPU network or I / O network on which the transaction is to be transferred based on a type of transaction. [18] SOC according to claim 9, wherein: the network with loosened sorting includes a large number of virtual channels; and Transactions in various virtual channels take different paths through the network with loosened sorting from a given source to a given destination. [19] SOC according to claim 9, wherein: The network is configured with loose sorting to redirect transactions from a given source to a given destination, and where different paths intended from the given source to the given destination are selected for the transactions based on one or more attributes of the transactions. [20] Procedures, including: Generating a transaction in one of a large number of processor clusters in a system, where a given processor cluster comprises one or more processors, and the system further includes the following: a large number of graphics processing units; a variety of storage controllers configured to control access to storage devices; a multitude of agents; and a multitude of network switches coupled to the multitude of processor clusters, the multitude of graphics processing units, the multitude of memory controllers, and the multitude of agents, wherein: a first subset of the multitude of network switches is interconnected to form a network of a central processing unit (CPU) between the multitude of processor clusters and the multitude of memory controllers, a second subset of the multitude of network switches is interconnected to form an input / output (I / O) network between the multitude of processor clusters, the multitude of agents, and the multitude of memory controllers, a third subset of the multitude of network switches is interconnected to form a loosely sorted network between the multitude of graphics processing units, selected agents of the multitude of agents, and the multitude of memory controllers, The CPU network, the I / O network, and the loosely sorted network are independent of each other. the CPU network and the I / O network are coherent and The network with relaxed sorting is incoherent and has reduced sorting constraints compared to the CPU network and the I / O network; and Transaction transfer based on a transaction type. [21] The method of claim 20, further comprising: Generating a second transaction by a given graphics processing unit of the plurality of graphics processing units; and Transferring the second transaction to the loosened-sort network, where a path of the second transaction through the loosened-sort network is based on one or more attributes of the second transaction. [22] The method of claim 20, further comprising: Generating a second transaction by a given graphics processing unit of the plurality of graphics processing units; and Transferring the second transaction to the loosened sorting network, wherein a path of the second transaction through the loosened sorting network is based on which of the multitude of virtual channels the second transaction is transferred on.

Citation Information

Patent Citations

  • Multiple heterogeneous noc layers

    US20140301241A1

  • Routing Flits in a Network-on-Chip Based on Operating States of Routers

    US20200153757A1