Communication fabric for increased bandwidth
By employing physical network topology and multi-port network interface design in the communication architecture, the problems of communication latency and power consumption are solved, and bandwidth and data transmission efficiency are improved. It is suitable for communication between agents that require ordered and unordered protocols.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-23
- Publication Date
- 2026-04-07
AI Technical Summary
Existing technologies suffer from high communication latency, multiple protocol conversions, and high power consumption in communication architectures, especially when data transmission occurs between agents that require ordered protocols and those that do not, which negatively impacts performance.
A physical network topology design is adopted, which physically separates the source agent and the destination agent and couples them through multiple communication paths. Data buffering is performed using network interfaces with more input channels than output channels. Multi-port network interfaces and I/O interfaces are combined to achieve flexible processing of ordered and unordered protocols.
It increases communication bandwidth, reduces power consumption, and optimizes data transmission efficiency, especially in communication between agents that require ordered protocols and those that do not.
Smart Images

Figure CN121816733A_ABST
Abstract
Description
Technical Field
[0001] The embodiments described herein relate to computing systems, including, for example, systems-on-a-chip (SoCs). More specifically, embodiments relating to techniques for increasing bandwidth through communication architectures are disclosed. Background Technology
[0002] Related Art Network architecture interconnects can provide a high-bandwidth, low-latency transport layer between various agents network-coupled across integrated circuits or multi-chip systems. Such interconnect architectures can be designed with various dedicated channels for transferring data between each agent in the various agents, such as a central processing unit (CPU), graphics processing unit (GPU), neural processing engine, memory system, etc. To support a unified memory space, high-bandwidth network architectures can employ fully buffered network switches. Various types of peripheral circuits (e.g., "agents") can be included in these systems and can employ various communication protocols. These agents can be coupled to the network switch within the network switch via one or more forms of networking interfaces.
[0003] In some systems, a communication architecture may include each agent connected to one of several network interfaces, which in turn may be coupled to the communication architecture. This technique may involve high communication latency, multiple protocol conversions, high-level data buffering, and high-level power consumption. Furthermore, protocols with ordering rules can be applied to agents that do not require ordered protocols, further degrading performance in some cases. Attached Figure Description
[0004] The following detailed embodiments are described with reference to the accompanying drawings, which will now be briefly described.
[0005] FIG. 1 A block diagram illustrating an implementation of a system comprising a communication network coupled to multiple agent circuits via a network interface and input / output (I / O) interfaces is provided.
[0006] FIG. 2 It shows FIG. 1 A block diagram of a part of the implementation scheme of the system, in which multiple protocols are used.
[0007] FIG. 3 Depicting FIG. 1 A block diagram of a part of the implementation of a system in which an ordered protocol is enforced for some proxy circuits.
[0008] FIG. 4 A block diagram illustrating another embodiment of a system comprising a communication network coupled to multiple agent circuits via network interfaces and I / O interfaces is shown.
[0009] FIG. 5 A flow diagram illustrating an embodiment of a method for communicating a data transaction using a network interface and an I / O interface is shown.
[0010] FIG. 6 A flow diagram illustrating an embodiment of a method for communicating a response to a data transaction using a network interface and an I / O interface is shown.
[0011] FIG. 7 A block diagram depicting an embodiment of a system having a communication network with multiple communication pathways is depicted.
[0012] FIG. 8 A block diagram illustrating an embodiment of a system having two SoCs, each having a communication network with multiple communication pathways is depicted.
[0013] FIG. 9 A block diagram illustrating an embodiment of a network having a ring topology is shown.
[0014] FIG. 10 A block diagram depicting an embodiment of a network having a mesh topology is depicted.
[0015] FIG. 11 A block diagram illustrating an embodiment of a system including a network interface having asymmetric input and output lanes is shown.
[0016] FIG. 12 A flow diagram illustrating an embodiment of a method for communicating a data transaction using a network interface having asymmetric input and output lanes is shown.
[0017] FIG. 13 Various embodiments of a system including an integrated circuit utilizing the disclosed technology are depicted.
[0018] FIG. 14 Is a block diagram of an example computer readable medium in accordance with some embodiments.
[0019] While embodiments described in this disclosure can be susceptible to various modifications and alternative forms, specific embodiments are shown by way of example in the drawings and will be described in detail herein. However, it should be understood that the drawings and detailed description thereto are not intended to limit the embodiments to the particular form disclosed, but on the contrary, the intention is to cover all modifications, equivalents and alternatives falling within the spirit and scope of the claims. DETAILED DESCRIPTION
[0020] Various integrated circuits and multi-chip systems can employ multiple communication networks. As used herein, a “communication network” or simply “network” collectively refers to various agents that communicate via a common set of network switches. Such networks can be physically independent (e.g., have dedicated wires and other circuitry that form the network) and logically independent (e.g., communications initiated by an agent in the system can be logically defined to send on a selected network of multiple networks and can not be affected by sending on other networks). In some embodiments, network switches can be included to send packets on a given network. As used herein, an “agent” refers to a functional circuit that can either initiate (originate) a communication on a network or be a destination for a communication on a network. An agent can generally be any circuit that can initiate and / or receive communications on a given network (e.g., CPUs, GPUs, neural processing engines, peripherals, memory controllers, etc.). A source agent generates (initiates) a communication, and a destination agent receives a communication. A given agent can be a source agent for some communications and a destination agent for other communications. In some cases, a communication (also referred to as a “transaction”) between two agents can cross between two or more networks in the network.
[0021] By providing physically and logically independent networks, high bandwidth can be achieved via parallel communications on different networks. Additionally, different traffic can be sent on different networks, and thus a given network can be optimized for a given type of traffic. For example, a multi-core CPU in a system can be sensitive to memory latency and can cache data that is expected to be consistent between cores and memory. Thus, a CPU network can be provided on which cores and memory controllers in the system are agents. Another network can be an input / output (I / O) network. This I / O network can be used by various peripherals (“peripherals”) to communicate with memory. The network can support the bandwidth required by the peripherals and can also support cache coherency. Further, the system can additionally include a relaxed order network. The relaxed order network can be non-coherent and can not enforce as many ordering constraints as the I / O or CPU networks. The relaxed order network can be used by a GPU to communicate with a memory controller. Other embodiments can employ any subset of the above networks and / or any additional networks as desired.
[0022] This combination of networks in a system can be referred to as a "communication fabric," "network fabric," or simply "fabric." In some cases, a "global fabric" can be used to refer to various communication paths that "interleave" across all networks in a system. Thus, a "local fabric" can refer to communication paths that "interleave" across a subset of networks and / or portions of networks. As described above, a communication fabric can include a plurality of agents connected to a network interface, which in turn can be coupled to the communication fabric. In some embodiments, a group of agents can be coupled to an input / output (I / O) interface that resides between the network interface and the group of agents.
[0023] The use of an I / O interface between the agents and the network interface can allow for the use of intellectual property (IP) circuit designs from various IP providers by accommodating a variety of data formats and communication protocols. The I / O interface can support conversion from one or more data formats to a data format supported by the network interface. In addition, such an I / O interface can support an in-order transaction protocol, where transactions are transmitted to the network interface in the order in which they are received from the agents. Any responses to a set of in-order transactions can also be returned to the source agents in the order in which the transactions were transmitted by the source agents.
[0024] The use of an I / O interface between the agents and the network interface can also result in high communication latency, multiple protocol conversions, high level data buffering, resulting in high level power consumption. In addition, the application of ordering rules can be applied to agents that do not require an in-order protocol, further degrading performance in some cases. Thus, a simplified I / O interface is proposed for use with peripheral agents that require an in-order protocol. A multi-port network interface is also proposed for direct communication with a plurality of peripheral agents that do not require the use of an in-order protocol.
[0025] Another novel communication fabric technique disclosed herein includes a physical network topology that places a source agent in one general area of an integrated circuit (IC) while a destination agent can be physically arranged in other areas of the IC. Two or more communication lanes can be utilized in a particular network to couple the source agent and the destination agent. At least some of the network switches coupled to the source agent can be cross-coupled between the communication lanes to enable communication across lanes. The network switches coupled to the destination agent can be coupled only to network switches in the same lane.
[0026] Another novel communication fabric technique involves the use of a network interface with more input lanes than output lanes. Such a network interface can allow for buffering of transactions in the input lanes of the network interface, removing transactions from the network switches, allowing for an increased number of transactions to be routed by the network switches of a given network.
[0027] For ease of discussion, various embodiments in this disclosure are described as being implemented using one or more SoCs. It should be understood that any disclosed SoC can also be implemented using a small-chip-based architecture. Thus, wherever the term “SoC” appears in this disclosure, those references are intended to suggest embodiments in which the same functionality is implemented via a less monolithic architecture, such as via multiple small chips, which in some embodiments can be included in a single package.
[0028] In related discussion, some embodiments are described herein that include more than one SoC. Such architectures should be understood to encompass both homogenous designs (in which each SoC includes the same or nearly the same functionality) and heterogeneous designs (in which the functionality of each SoC differs more). Such disclosure also contemplates embodiments in which different levels of disaggregation are used to implement the functionality of multiple SoCs. For example, the functionality of a first system can be implemented on a single IC, while the functionality of a second system (which can be the same or different than the first system) can be implemented using multiple co-packaged small chips.
[0029] FIG. 1 A block diagram illustrating an embodiment of a system using I / O interfaces in conjunction with network interfaces is exemplified. The system 100 includes a plurality of agent circuits 120a-120h (collectively, 120). A portion of the agent circuits 120 are directly coupled to one of two network interfaces 101a and 101b (collectively, 101), while the remaining agent circuits 120 are coupled to one of the network interfaces 101 via a respective one of I / O interfaces (I / O I / F) 125a-125c (collectively, 125). The network interfaces 101 are coupled to a communication network that includes network switch circuits (NS) 110a and 110b. The network interfaces 101 support communication of data transactions 140 and 144 from respective ones of the agent circuits 120 in conjunction with the communication network 105. The system 100 can be a computing system in whole or in part, such as a desktop or laptop computer, a smartphone, a tablet computer, or a wearable smart device. In some embodiments, the system 100 is a single IC, such as a system-on-a-chip or a multi-die chip.
[0030] As shown, a first set of agent circuits (120d, 120e, and 120f) are configured to communicate data transactions, including data transaction 144, using an ordered protocol, such as a Peripheral Component Interconnect (PCI) protocol. Such an ordered protocol can specify that data transactions are communicated in the order in which they are transmitted from the respective agent circuits. In contrast, a second set of agent circuits (120a-c and 120g-h) are configured to communicate data transactions using a protocol that does not enforce ordering. To support the ordered protocol for the first set of agent circuits, input / output (I / O) interfaces 125a, 125b, and 125c are coupled to agent circuits 120d, 120e, and 120f, respectively, and are configured to enforce the ordered protocol for data transactions (including data transaction 144) transmitted by the first set of agent circuits.
[0031] Communication network 105 includes network switch circuits 110a and 110b. As illustrated, network switch circuit 110a is coupled to network switch circuit 110b, and each network switch circuit can be further coupled to one or more additional network switch circuits in communication network 105.
[0032] As shown, network interface 101a is coupled to agent circuits 120a-c, I / O interface 125a, and network switch circuit 110a. Network interface 101a can be configured to communicate data transactions (including data transaction 140) between agent circuits 120a-c and network switch circuit 110a, as well as between I / O interface 125a and network switch circuit 110a. Similarly, network interface 101b is coupled to agent circuits 120g and 120h, I / O interfaces 125b and 125c, and network switch circuit 110b. Network interface 101b can be configured to communicate data transactions between agent circuits 120g and 120h and network switch circuit 110a, as well as between I / O interfaces 125b and 125c and network switch circuit 110a (including data transaction 144).
[0033] As illustrated, the agent circuit 120a passes data transactions 140 to the network interface 101a in a first order: data transaction 140a, followed by data transaction 140b, and then data transaction 140c. Since the network interface 101a does not need to enforce an ordered protocol, the data transactions 140 can be transmitted to the network switch circuit 110a in any suitable order. For example, each of the data transactions 140 can be directed to a different destination agent circuit, such as a different memory circuit. Thus, the network interface 101a can transmit the data transactions 140 when resources are available to receive a given transaction. Each of these memory circuits can include a respective request queue, and can not be available to receive a given transaction until an entry in the respective request queue is available. Thus, the data transactions 140 can be transmitted in an order that is different from the order in which they are received. As shown, the network interface transmits data transaction 140c first, followed by 140a, and then 140b. In other cases, the data transactions can be reordered based on their respective transaction types. For example, data transactions 140a and 140b can be write requests for a given memory circuit, while data transaction 140c is a read request for the same memory circuit. Read requests can have a higher priority than write requests, and in some cases can also complete faster than write requests. The network interface 101a can reorder the data transactions 140 for various reasons, some of which can increase the bandwidth of the communication network 105 and / or reduce the power consumed by the communication network 105 and / or the destination agent circuits.
[0034] As shown, the agent circuit 120e passes the data transactions 144 to the I / O interface 125b in the first order: data transaction 144a, followed by data transaction 144b, and then data transaction 144c. The I / O interface 125b is configured to follow the same order for the data transactions 144. In some embodiments, the I / O interface 125b can forward each of the data transactions 144 to the network interface 101b one at a time, waiting for an acknowledgement or other form of response before transmitting the subsequent data transaction of the data transactions 144. For example, the I / O interface 125b can transmit transaction 144a to the network interface 101b and wait for a response transmitted via the network interface 101b before transmitting transaction 144b. In this way, the ordered processing of the data transactions 144 can be ensured, but can also be slower compared to other techniques. In other embodiments, the I / O interface 125b can pass the data transactions 144 to the network interface 101b in the order received. However, the network interface 101b can not enforce the same order and can forward the data transactions 144 in a different order. However, the I / O interface 125b can then receive the responses associated with the data transactions 144 and buffer the received responses until the response to the first transaction (data transaction 144a) is received, which can then be forwarded to the agent circuit 120e. The I / O interface 125b can continue to buffer the responses until the response to data transaction 144b is received, and so on. In this way, from the perspective of the agent circuit 120e, the I / O interface 125b can process the transactions in the original order.
[0035] Note that, FIG. 1 The illustrated system 100 is merely an example. The system 100 has been simplified to highlight features relevant to the present disclosure. Elements that are not relevant to the description of the disclosed concepts have been omitted. For example, the system 100 can include various circuitry that is not illustrated, such as one or more memory circuits, clock generator circuits, and power management circuits, etc. Only one communication network is illustrated. In other embodiments, a communication fabric having any suitable number of networks can be included. In various embodiments, the circuitry of the system 100 can be implemented using any suitable combination of sequential and combinatorial logic circuits. In addition, registers and / or memory circuits (such as static random access memory (SRAM)) can be used in these circuits to temporarily hold information such as instructions, data, address values, etc.
[0036] In FIG. 1 communication networks utilizing network interfaces and I / O interfaces are disclosed. Such network interfaces and I / O interfaces can perform additional processing that is not disclosed in the description of the FIG. 1 FIG. 2 Examples of converting between different data formats are depicted in
[0037] Turning toFIG. 2 FIG. 1 illustrates a portion of a block diagram of a system 100. FIG. 1 FIG. 2 Illustrated in FIG. 1 are agents 120a-d, network interface 101a, network switch circuit 110a, and I / O interface 125a. This portion of system 100 is used to depict how a transaction's data format is modified as the transaction progresses from a source agent to a communication network.
[0038] As illustrated, agent circuit 120d is configured to transmit a data transaction 242a to I / O interface 125a using a first protocol (protocol 0). Agent circuit 120a is configured to transmit a data transaction 240a to network interface 101a using a second protocol (protocol 1) that is different from protocol 0. In various embodiments, different protocols can include different data formats as well as different rules for transmitting and receiving data packets. For example, protocol 1 can specify a data packet that has a destination address in a first bit field, a packet size in a second bit field, a source / owner process identifier in a third bit field, and one or more data words specified by the packet size. Rules for transmitting and receiving packets using protocol 1 can include receiving an acknowledgement from a destination agent and one or more rules that bound packet priority relative to other transactions using the same protocol. Protocol 0 can specify similar information in a data packet, but the bit fields can be arranged in a different order. In addition, protocol 0 can specify in-order processing of transactions, and thus can include a packet number in an additional bit field that can be used to determine the order.
[0039] As shown, I / O interface 125a is configured to transmit data transaction 242a received from agent circuit 120d to network interface 101a using protocol 1. Since network interface 101a is coupled to a plurality of different agents, a circuit designer can desire to simplify the design of network interface 101 to support a limited number of protocols. If a particular agent circuit (e.g., agent circuit 120d) does not support one of the limited number of protocols, then a corresponding I / O interface, such as I / O interface 125a, can be included between the agent circuit that does not support the limited protocol (e.g., agent circuit 120d) and network interface 101a.
[0040] Accordingly, I / O interface 125a is configured to convert transaction 242a from protocol 0 to protocol 1, generating data transaction 242b. Such conversion can include rearranging data in various bit fields of protocol 0 into corresponding bit fields of protocol 1, removing extraneous information not included in protocol 1 and adding information not supported by protocol 0 but required by protocol 1, etc. I / O interface 125a transmits data transaction 242b to network interface 101a after conversion.
[0041] The network interface 101a can in turn be configured to convert the received data transactions 240a and 242b to a third protocol (Protocol 2) supported by the network switch circuit 110a and other network switch circuits in the communication network 105. Thus, these conversions of the data transactions 240a and 242b to Protocol 2 can generate transactions 240b and 242c, respectively. Similar to the conversion from Protocol 0 to Protocol 1, the conversion from Protocol 1 to Protocol 2 can include various combinations of rearranging, adding, and removing data among the multiple bit fields in the data packets.
[0042] While transactions are illustrated as moving from the agent circuits 120a and 120d to the communication network 105, the reverse can also be implemented. Thus, the network interface 101a can be configured to convert transactions from Protocol 2 to Protocol 1 from the network switch circuit 110a before being communicated to any of the agent circuits 120a-c or the I / O interface 125a. Likewise, the I / O interface 125a can convert a transaction from Protocol 1 to Protocol 0 before communicating the transaction to the agent circuit 120d from the network interface 101a.
[0043] Note that the system of FIG. 2 is simplified for clarity. In other implementations, any suitable number of agent circuits, I / O interfaces, network interfaces, network switch circuits, etc. can be included. While the network interface 101a is described as converting between Protocol 1 and Protocol 2, other network interfaces can convert between other protocols. For example, FIG. 1 The network interface 101b in can use a fourth protocol for data packets communicated to and from the agent circuits g and h and / or the I / O interfaces 125b and 125c.
[0044] FIG. 2 Depictions of how different protocols are handled for different agent circuits are described. In FIG. 1 the description, a protocol that enforces in-order processing of transactions is disclosed. FIG. 3 Depictions of such implementations are described.
[0045] Continuing to FIG. 3 the same portions of the block diagram of the system 100 of FIG. 3 are illustrated. FIG. 3 Depicted in are the agents 120a-d, the network interface 101a, the network switch circuit 110a, and the I / O interface 125a. This portion of the system 100 is used to depict how the order of a transaction can be adjusted as the transaction progresses from a source agent to the communication network via the I / O interface.
[0046] As illustrated, the agent circuit 120d is configured to transmit a series of data transactions (transactions 342a-342c) to the I / O interface 125a in a first order. The agent circuit 120d can be configured to transmit and receive the data transactions 342a-342c in a particular order. For example, the agent circuit 120d can be configured to utilize a PCI protocol in which data transactions are expected to be processed in the same order in which they are transmitted. The PCI protocol can also expect an acknowledgement or other type of response associated with the transmitted data transactions. As illustrated, the agent circuit 120d transmits transaction 342a, followed by transaction 342b, and then transaction 342c. This is referred to herein as the original order.
[0047] As illustrated, the I / O interface 125a is configured to transmit the data transactions 342a-342c to the network interface 101a in the first order (e.g., 342a, 342b, and then 342c). In some embodiments, the I / O interface 125a can be configured to transmit the data transactions 342a-342c to the network interface 101a using the same original order. However, the network interface 101a can be configured to transmit the data transactions 342a-342c to the communication network 105 in a second order that is different from the first order (i.e., the original order). For example, one or more of the data transactions 342a-342c can be directed to a different destination agent than the other transactions. One of these different destination agents can have a full request queue and be unable to receive new transactions for a period of time, while the other of these destination agents can have available resources and thus be able to receive new transactions. Thus, the network interface 101a can transmit transaction 342c first, while transactions 342a and 342b are delayed due to the unavailable resources of their respective destination agents, because the associated destination agent is available to receive the transaction.
[0048] The network interface 101a can also use other criteria to determine the order for transmitting the data transactions 342a-342c. For example, each of the data transactions 342a-342c can have a respective priority assigned, such as a bulk transaction or a real-time transaction. The bulk transaction can correspond to a standard or lowest priority for data transactions, while the real-time transaction can correspond to a high priority in which the transaction needs to be processed with a minimum amount of latency. In some embodiments, the agent circuit 120d can include a plurality of processor cores, each of which is capable of executing a different software process. The priority can be associated with these different processes, which in turn are assigned to transactions in transactions issued by the respective processes.
[0049] As illustrated, network interface 101a receives the responses to each of data transactions 342a-c from network switch circuit 110a in a third order that is different from the original or second order. As depicted, response 344b, which corresponds to transaction 342b, is received first, followed by response 344c (corresponding to transaction 342c), and then response 344a (corresponding to transaction 342a). The responses can be received in the third order due to the relative amount of time each of these destination agents takes to receive and process a transaction. In addition, network traffic between network switch circuit 110a and network switch circuits coupled to each of these destination agents can delay the transmission of transactions and the receipt of responses. If a given destination agent is located at a far end of communication network 105, the transactions and responses can travel through tens or even hundreds of network switch circuits.
[0050] As shown, network interface 101a is configured to transmit responses 344a-c to I / O interface 125a in the third order in which they are received. For example, network interface 101a can forward each response as it is received. On the other hand, I / O interface 125a is configured to transmit responses 344a-c to agent circuit 120d in the original, first order. To transmit responses 344a-c to agent circuit 120d in the original order, I / O interface 125a is configured to buffer one or more of responses 344a-c as they are received from network interface 101a in the third order. For example, I / O interface 125a can include a buffer circuit that is capable of storing multiple responses. Thus, I / O interface 125a can buffer response 344b and response 344c until response 344a has been received. After response 344a is received, it is forwarded to agent circuit 120d, followed by response 344b, and then response 344c. Thus, agent circuit 120d receives responses 344a-c in the original order, and I / O interface will follow an ordered protocol despite the fact that network interface 101a does not follow an ordered protocol.
[0051] As depicted, agent circuit 120a is configured to transmit a series of data transactions (transactions 340a-c) to network interface 101a in a given order. First is transaction 340a, followed by transaction 340b, and last is transaction 340c. Network interface 101a can use one or more of the criteria described above to determine the order in which to transmit data transactions 340a-c to network switch circuit 110a. As shown, the order is as follows: transaction 340a first, transaction 340c second, and transaction 340b third.
[0052] Responses to these transactions are received at a later point in time in another order. Responses 341a-c are received in the following order: response 341c first, response 341a second, and response 341b third. Since the proxy circuit 120a does not enforce an ordered protocol, the network interface 101a can forward these responses to the proxy circuit 120a in the same order in which the responses are received.
[0053] Note that, FIG. 1 to FIG. 6 Embodiments of the system 100 are merely examples. A limited number of elements are shown to illustrate the disclosed concepts. Any suitable number of each element can be included in other embodiments. Although the I / O interface 125a is described as preserving the original order, in some embodiments, the I / O interface 125a can reorder transactions from the proxy circuit 120d, e.g., moving real-time transactions ahead of one or more bulk transactions. In some embodiments, the I / O interface 125a can still transmit responses to the proxy circuit 120d in the original order.
[0054] In the system 100 of FIG. 7 to FIG. 10 In the system 100 of FIG. 7 Embodiments of such systems are illustrated.
[0055] Continuing to FIG. 1 A block diagram of an embodiment of a system grouping particular sets of proxy circuits with a common network interface circuit. The system 400 includes proxy circuits 420a-h (collectively, 420). The proxy circuits 420a-d are coupled directly to a network interface 401a, while the proxy circuits 420e-h are coupled to a network interface 401b via respective ones of I / O interfaces (I / O I / Fs) 425a-d (collectively, 425). The network interfaces 401a and 401b are coupled to a communication network 405 via network switch circuits (NSs) 410a and 410b, respectively. In a similar manner to the system 100, the system 400 can be included, in whole or in part, in a computing system such as a desktop or laptop computer, a smartphone, a tablet computer, or a wearable smart device. In some embodiments, the system 400 is a single IC such as a system-on-a-chip or a multi-die chip.
[0056] The system 400 illustrates a different arrangement of agent circuitry to network interface. As illustrated, the agent circuitry 420e-420h can be configured to communicate data transactions using an ordered protocol such as described above. Accordingly, the I / O interfaces 425a-425d are coupled to respective ones of the agent circuitry 420e-420h, as well as to the network interface 401b. The I / O interfaces 425a-425d are configured to enforce the ordered protocol when the agent circuitry 420e-420h transmit and receive transactions and / or responses via the network interface 401b. The network interface 401b is further coupled to the network switch circuitry 410b.
[0057] In the manner as previously disclosed, any of the agent circuitry 420e-420h can transmit transactions in a first order using a first protocol, and respective ones of the I / O interfaces 425a-425d can convert the transactions to a second protocol and forward the converted transactions to the network interface 401b in the first order. The network interface 401b can then convert the transactions from the second protocol to a third protocol and can reorder the transactions before passing the transactions to other network switch circuitry and to respective destination agents. Received responses to these transactions can be forwarded to respective I / O interfaces 425a-425d as they are received, regardless of the original order of the transactions. However, the I / O interfaces 425a-425d can buffer received responses that are not acknowledged to be in the first order. When a response to a first transaction in the first order is received, these responses can be forwarded to respective agent circuitry 420e-420h.
[0058] In contrast, the network interface 401a is directly coupled to the agent circuitry 420a-420d and the network switch circuitry 410a. As described above, any of the agent circuitry 420a-420d can transmit transactions in a first order using a second protocol, and the network interface 401a can convert the transactions from the second protocol to a third protocol and can reorder the transactions before passing the transactions to other network switch circuitry and to respective destination agents. Received responses to these transactions can be forwarded to respective agent circuitry 420a-420d as they are received, regardless of the original order of the transactions.
[0059] By grouping agent circuitry having a common attribute to a common network interface, a given network interface can be optimized for the common attribute. However, respective ones of the I / O interface circuitry can still be used for each agent circuitry that follows an ordered protocol. This can allow a single or few I / O interface circuitry designs to be created and reused across multiple IC designs. Separate I / O interface circuitry can also support scalability, as more or less agent circuitry following an ordered protocol can be used in various IC designs.
[0060] Note that,FIG. 7 The illustrated system 400 is merely an example to demonstrate the disclosed technology. The system 400 has been simplified to highlight aspects of the present disclosure. Although a single communication network is illustrated, communication fabrics having any suitable number of networks can be included in other implementations. As described for system 100, the circuitry of system 400 can be implemented using any suitable combination of sequential and combinatorial logic circuits and registers and / or memory circuits.
[0061] The above description is related to FIG. 1 The described circuitry and technology can be performed using various methods. The following describes two methods associated with communicating transactions via a communication network. FIG. 7 The described circuitry and technology can be performed using various methods. The following describes two methods associated with communicating transactions via a communication network.
[0062] Turning now to FIG. 8 , a flowchart of a method implementation for communicating transactions issued by two different agent circuits is illustrated. The method 500 can be performed by any of the systems disclosed herein, such as the systems 100 and 400 of FIG. 8 . The method 500 is described below using the system 100 of FIG. 7 as an example. References to elements in FIG. 7 are included as non-limiting examples.
[0063] At 510, the method 500 begins with a first set of data transactions being communicated by a first agent to an I / O interface using an ordered protocol. For example, the agent circuit 120e can transmit the data transactions 144 to the I / O interface 125b in a first order. The agent circuit 120e can be configured to expect the data transactions 144 to be processed in the first order, as illustrated, which corresponds to the following: data transaction 144a first, data transaction 144b second, and data transaction 144c third. Enforcement of the first order can include expecting responses corresponding to the data transactions 144 to follow the first order.
[0064] The method 500 continues at 520 with a second set of data transactions being communicated by a second agent to a network interface circuit using a non-enforced order protocol. For example, the agent circuit 120a can transmit the data transactions 140 to the network interface 101a in a first order, which corresponds to the following: data transaction 140a first, data transaction 140b second, and data transaction 140c third. Unlike the agent circuit 120e, the agent circuit 120a can be configured to allow the processing of the data transactions 140 to occur in any suitable order. Responses to the data transactions in the data transactions 140 can be received by the agent circuit 120a in any order.
[0065] At 530, the method 500 continues with the first set of data transactions being communicated by the I / O interface to the network interface circuit using the ordered protocol. AsFIG. 8 As shown, I / O interface 125b can forward data transactions 144 to network interface 101b using first order. In some embodiments, I / O interface 125b can delay transmitting subsequent data transactions in data transactions 144 until a previously transmitted transaction 144 has been forwarded by network interface 101b. Such techniques can enable I / O interface 125b to enforce first order. In other embodiments, I / O interface 125b can not delay between delivery of data transactions 144, which can allow network interface 101b to reorder data transactions 144.
[0066] Method 500 continues at 540 with delivering, by the network interface circuit, the first set of data transactions and the second set of data transactions to the communication fabric using an order based on respective destination availability. Network interfaces 101a and 101b can determine availability of the destination agent for each of data transactions 140 and 144, respectively. Network interfaces 101a and 101b can also determine relative priority between respective data transactions 140 and 144. Using such availability and priority information, network interface 101a can determine to deliver data transaction 140c before data transactions 140a and 140b. In some cases, network interface 101b can determine that first order can be maintained, and thus can deliver data transactions 144 using first order. In other embodiments, network interface 101b can reorder data transactions 144 due to, for example, availability of resources in one or more of the destination agents.
[0067] Using an I / O interface with an agent circuit that enforces ordered transactions can allow enforcement to be limited to agent circuits that require ordered processing of transactions. Other solutions can place enforcement of ordered transactions in a network interface that is coupled to multiple agent circuits, some of which can not acknowledge ordered protocols. Forcing ordered processing onto agent circuits that do not acknowledge such protocols can increase latency for these agent circuits to complete transactions, potentially reducing performance bandwidth of the system.
[0068] Note that, FIG. 1 to FIG. 8 The method of 500 includes blocks 510-540. Method 500 can end in block 540, or some or all of the blocks of the method can be repeated. For example, method 500 can repeatedly repeat blocks 510 and / or 520 to deliver additional transactions. In some embodiments, different instances of method 500 can be concurrently executed in system 100. For example, agent circuit 120a and network interface 101b can execute portions of one instance of method 500, while agent circuit 120e, I / O interface 125b, and network interface 101b execute portions of a different instance of method 500.
[0069] Proceeding toFIG. 9 illustrates a flow diagram of an embodiment of a method for receiving responses to previously issued transactions issued by proxy circuitry using an ordered protocol. Similar to method 500, method 600 can be performed by any of the systems disclosed herein, including, for example FIG. 10 systems 100 and 400. Method 600 is described below using system 100 as an example. Reference to elements in FIG. 9 is included as non-limiting examples. The operations of method 600 can occur after an instance of method 500 has been performed. FIG. 9
[0070] Method 600 begins at 610: receiving, by network interface circuitry via a communication fabric, responses to a first set and a second set of data transactions in an order based on respective destination response times. As FIG. 9 illustrated, for example, proxy circuitry 120a issues data transactions 340a-c, and proxy circuitry 120d issues data transactions 342a-c. Each proxy circuitry transmits their respective transactions in the order a-b-c. When these transactions are forwarded to their respective destination proxies, the transactions can be reordered. At a later point in time, network interface 101a receives responses to each of the issued transactions. Network interface 101a receives responses 341a-c in the order c-a-b, and receives responses 344a-c in the order b-c-a.
[0071] At 620, method 600 continues with passing, by network interface circuitry to an I / O interface, responses to the first set of data transactions based on the order of receipt. For example, network interface 101a passes responses 344a-c to I / O interface 125a in the order of receipt (e.g., b-c-a).
[0072] Method 600 continues at 630 with passing, by the I / O interface to the first proxy, responses to the first set of data transactions using an ordered protocol. As FIG. 10 depicted, I / O interface 125a forwards responses 344a-c to proxy circuitry 120d in the original order (e.g., a-b-c). To accomplish this, I / O interface 125a can buffer responses 344b and 344c until response 344a is received. After response 344a is received, it can be forwarded to proxy circuitry 120d, followed by response 344b, and then response 344c, presenting responses 344a-c to proxy circuitry 120d in the original issuance order.
[0073] At 640, the method 600 continues with the network interface circuit transferring the responses to the second set of data transactions to the second agent based on the order of receipt. As described above, the agent circuit 120a can not be configured to enforce an ordered protocol. Thus, the network interface 101a can forward the responses to the agent circuit 120a in the same order in which the responses were received (e.g., response 341c first, response 341a second, and response 341b third).
[0074] Note that the method 600 includes blocks 610-640. The method 600 can end in block 640, or some or all of the blocks of the method can be repeated. For example, the method 600 can return to block 610 to receive additional responses to other previously issued transactions. Various instances of the methods 500 and 600 can be executed concurrently.
[0075] In FIG. 10 , systems and techniques for using I / O interfaces in a communication network including multiple agent circuits are disclosed, where some of the agent circuits can enforce an ordered protocol and where some of the agent circuits can not enforce an ordered protocol. Other techniques for increasing bandwidth in a communication network can be implemented. FIG. 10 One such technique is described in
[0076] Moving to FIG. 10 , a block diagram of an embodiment of a system including a multi-lane communication network is depicted. As illustrated in the figure, the system 700 includes two sets of network switch circuits. The network switch circuits 710a-710g are included in a communication lane 707a. The network switch circuits 712a-712g are included in a communication lane 707b. The system 700 also includes multiple agent circuits 720a-720f and multiple memory circuits 725a-725h. The communication lanes 707a and 707b are included in a communication network 705. In a similar manner as described for the system 100 in FIG. 1 to FIG. 6 , the system 700 can be, in whole or in part, a laptop or desktop computer, a smart phone, a tablet computer, or a wearable smart device, among others. In some embodiments, the system 700 can be a single IC, such as a system on a chip, or can be a multi-die chip.
[0077] As illustrated, the agent circuits 720a-f can be configured to initiate memory transactions as well as receive memory transactions. The agent circuits 720a-f can be any circuit (e.g., CPU, GPU, neural processing engine, peripheral device, memory controller, etc.) that can initiate and / or receive communications over the communication network 705. The memory circuits 725a-h can be configured to respond to memory transactions from the agent circuits 720a-f. For example, the memory circuits 725a-h can include any suitable combination of volatile and non-volatile memory units including, for example, static random access memory (SRAM), dynamic random access memory (DRAM), flash memory, hard disk drive, register file, etc. A given memory circuit of the memory circuits 725a-h can be configured to process a given memory transaction based on an address accessed by the given memory transaction, including read transactions to access values stored in the given memory circuit 725a-h and write transactions to store values supplied by a respective one of the agent circuits 720a-f.
[0078] As shown, the communication network 705 includes network switch circuits 710a-c coupled to the agent circuits 720a-c, respectively, and network switch circuits 712a-c coupled to the agent circuits 720d-f, respectively. The communication network 705 also includes network switch circuits 710d-g coupled to the memory circuits 725a-d and network switch circuits 712d-g coupled to the memory circuits 725e-h. The communication network 705 is configured to communicate memory transactions via the communication pathways 707a and 707b.
[0079] As depicted, the communication pathway 707a includes the network switch circuits 710a-g and thus couples to an appropriate subset of the agent circuits 720a-f (e.g., the agent circuits 720a-c) as well as an appropriate subset of the memory circuits 725a-h (e.g., the memory circuits 725a-d). In a similar manner, the communication pathway 707b includes the network switch circuits 712a-g and thus couples to an appropriate subset of the agent circuits 720d-f and to an appropriate subset of the memory circuits 725e-h. Note that the appropriate subsets of agent circuits and memory circuits are mutually exclusive between the communication pathways 707a and 707b.
[0080] Network switch circuits 710a-c in communication path 707a are coupled to respective ones of network switch circuits 712a-c in communication path 707b. In some embodiments, network switch circuits 710a-c and network switch circuits 712a-c can be coupled to form a mesh subnetwork between communication paths 707a and 707b. In contrast, network switch circuits 710d-g in communication path 707a are isolated from network switch circuits 712a-g in communication path 707b. Similarly, network switch circuits 712d-g in communication path 707b are isolated from network switch circuits 710a-g in communication path 707a. As used herein, "isolated from" means the absence of a direct link from an isolated network switch circuit in one communication path to a network switch circuit in a different communication path, and vice versa.
[0081] In some embodiments, agent circuits 720a-f can be physically located in one particular region of an IC die, such as at one end of the die. Accordingly, network switch circuits 710a-c and 712a-c can be placed in the same region. Moreover, these network switch circuits 710a-c and 712a-c can be coupled to form a mesh network to enable transaction passing between any set of agent circuits 720a-f. The close physical proximity between these agent circuits can reduce latency of communications between the agent circuits. In some cases, agent circuits 720a-f can frequently communicate with each other, and thus, reducing latency (as compared to having the agent circuits more spread apart from each other on the IC die) can improve performance of system 700.
[0082] In contrast, memory circuits 725a-h can not communicate among each other. Rather, memory circuits 725a-h can more frequently communicate with respective ones of agent circuits 720a-f. Accordingly, memory circuits 725a-h can be placed farther apart from each other across other regions of the IC die. If multiple agent circuits 720a-f are accessing respective ones of memory circuits 725a-h, using multiple communication paths including communication paths 707a and 707b can help alleviate traffic congestion. The use of communication paths 707a and 707b can help distribute memory transactions across different sets of network switch circuits 710a-g and 712a-g, reducing latency due to multiple transactions having to go through a common network switch circuit among network switch circuits 710a-g and 712a-g.
[0083] Although the communication network 705 is divided into communication lanes 707a and 707b, any of the agent circuits 720a-f can transmit and receive memory transactions to any of the memory circuits 725a-h. For example, the agent circuit 720b coupled to the network switch circuit 710b in the communication lane 707a can be configured to initiate a particular memory transaction to the memory circuit 725h coupled to the network switch circuit 712g in the communication lane 707b. To transmit the memory transaction, the agent circuit 720b can pass the particular memory transaction to the network switch circuit 710b. The network switch circuit 710b, in turn, can be configured to send the particular memory transaction to the network switch circuit 712b in the communication lane 707b. As illustrated, the network switch circuit 710b can be configured to send the particular memory transaction directly to the network switch circuit 712b without using an intermediate network switch circuit.
[0084] The network switch circuit 712b can be configured to send the particular memory transaction to the network switch circuit 712g in the communication lane 707b. To access the memory circuit 725h, the network switch circuit 712b can also be configured to transmit the particular memory transaction via one or more intermediate network switch circuits (e.g., the network switch circuits 712c and 712f in the illustrated example) in the communication lane 707b.
[0085] Note that, FIG. 11 The illustrated system 700 is merely an example. Like the system 100 in FIG. 11 has been simplified to highlight features relevant to the present disclosure and omit elements that are not relevant. For example, the system 700 can include various circuitry not illustrated, such as clock generator circuitry and power management circuitry, among others. Although the disclosed communication lanes are described as being included in a single communication network, in other embodiments, a communication fabric having any suitable number of networks can be included. The circuitry of the system 700 can be implemented using any suitable combination of sequential and combinatorial logic circuits. The memory circuits can be implemented as suitable types of memory cells, such as SRAM, DRAM, and flash memory, among others.
[0086] In FIG. 1 to FIG. 4 disclosed, for example, a single communication network on a single IC die. Some systems can implement a multi-die solution, with each die including at least one respective communication network. In FIG. 11 an example of a multi-die embodiment is depicted.
[0087] Moving to FIG. 11, a block diagram illustrating an embodiment of a system including two IC dies, each including a respective multi-pass communication network. As shown, system 800 includes SoCs 801a and 801b. In various embodiments, SoCs 801a and 801b can be packaged separately and coupled together via a circuit board, coupled together within a common package, or coupled together via any other suitable manner. SoC 801a includes network switch circuitry 810a-810g and 812a-812e, agent circuitry 820a-820f, and memory circuitry 825a-825f. Communication network 805a includes communication passes 807a and 807b. SoC 801b includes network switch circuitry 810h-810k and 812f-812j, agent circuitry 820g-820j, and memory circuitry 825g-825k. Communication network 805b includes communication passes 807c and 807d. In a similar manner as described with respect to system 700 in, for example FIG. 12
[0088] As illustrated, SoC 801b includes agent circuitry 820h, which is a die-to-die interface circuit. Further, agent circuitry 820b in SoC 801a includes a die-to-die interface circuit configured to pass memory transactions between SoCs 801a and 801b. The die-to-die interfaces in agent circuitry 820b and 820h can enable SoCs 801a and 801b to communicate and process transactions such that an operating system executing on SoCs 801a and / or 801b can behave as if system 800 is a single IC.
[0089] For example, agent circuitry 820b is coupled to network switch circuitry 810b in communication pass 807a of SoC 801a. Agent circuitry 820f is coupled to network switch circuitry 812c in communication pass 807b and can be configured to initiate a particular memory transaction for memory circuitry 825j in SoC 801b. To pass the particular memory transaction to memory circuitry 825j, agent circuitry 820f can be configured to transmit the particular memory transaction to network switch circuitry 810c, which in turn can be configured to pass the particular memory transaction to network switch circuitry 810b via network switch circuitry 810c or 812b, both of which are coupled to network switch circuitry 810b.
[0090] As illustrated, network switching circuit 810b can be configured to pass specific memory transactions to proxy circuit 820b. Using a die-to-die interface, proxy circuit 820b passes specific memory transactions to proxy circuit 820h in SoC 801b. Proxy circuit 820h can be configured to pass specific memory transactions to network switching circuit 812i via network switching circuits 810i, 812f, and 812h. The specific memory transactions can then be delivered to memory circuit 825j via network switching circuit 812j.
[0091] Using multiple communication paths allows SoCs 801a and 801b to be coupled to each other via a single die-to-die interface on each SoC. Once a transaction has crossed the die-to-die interface to reach its destination SoC, using multiple communication paths helps avoid traffic congestion. Furthermore, for example, regarding... FIG. 12 As described in System 700, the proxy circuits 820a to 820j can be physically placed in the corresponding area of each SoC, close to the corresponding die-to-die interface. This placement helps reduce transaction latency from the source proxy circuit to the die-to-die interface.
[0092] It should be noted that FIG. 11 The system described is an example used to demonstrate the disclosed technology. System 800 has been simplified to show the components relevant to this demonstration. Although one communication network is shown for each SoC, other implementations may include a communication architecture with any suitable number of networks.
[0093] exist FIG. 11 The document discloses various implementation schemes for communication networks. A communication network can be a single network or a network configuration comprising multiple networks. Whether a single network or a network configuration, each network can be implemented using any suitable type of network topology, where a given network configuration includes networks with different structures. FIG. 12 and FIG. 11 Two examples of network topologies are shown.
[0094] Go to FIG. 11 A block diagram of an implementation scheme using a ring topology to couple multiple proxy circuits is shown. FIG. 1 to FIG. 12 In this example, the ring is formed by network switching circuits 914AA to 914AH. Proxy circuit 910A is coupled to network switching circuit 914AA; proxy circuit 910B is coupled to network switching circuit 914AB; and proxy circuit 910C is coupled to network switching circuit 914AE.
[0095] As illustrated, a“network switch circuit” or simply“network switch” is a circuit configured to receive communications on a network and forward the communications on the network in the direction of the destination of the communications. For example, a communication initiated by a processor can be sent to a memory controller that controls a memory mapped to a communication address. At each network switch, the communication can be sent forward to the memory controller. If the communication is a read, the memory controller can communicate data back to the source, and each network switch circuit can forward the data on the network to the source. In embodiments, the network can support multiple virtual lanes. The network switch circuit can employ resources (e.g., buffers) dedicated to each virtual lane such that communications on a virtual lane can remain logically independent. The network switch circuit can also employ arbitration circuitry to select among buffered communications to forward on the network. The virtual lanes can be lanes that are physically shared network but logically independent from the network (e.g., communications in one virtual lane do not block the progress of communications on another virtual lane).
[0096] In a ring topology, each network switch circuit 914AA-914AH can be connected to two other network switch circuits 914AA-914AH, and the switches form a ring such that any network switch circuit 914AA-914AH can reach any other network switch circuit in the ring by sending a communication in the direction of the other network switch on the ring. A given communication can pass through one or more intermediate network switch circuits in the ring to reach a target network switch circuit. When a given network switch circuit 914AA-914AH receives a communication from an adjacent network switch circuit 914AA-914AH on the ring, the given network switch circuit can inspect the communication to determine whether the agent circuit 910A-910C to which the given network switch circuit is coupled is the destination of the communication. If so, the given network switch circuit can terminate the communication and forward the communication to the agent. If not, the given network switch circuit can forward the communication to the next network switch circuit on the ring (e.g., other network switch circuit 914AA-914AH that is adjacent to the given network switch circuit and is not the adjacent network switch circuit from which the given network switch circuit received the communication). As used herein, an“adjacent network switch” to a given network switch circuit can be a network switch circuit to which the given network switch circuit can send a communication directly without the communication traveling through any intermediate network switch circuit.
[0097] FIG. 13 An example of a ring network topology is illustrated in FIG. 9. As illustrated, any pair of adjacent network switch circuits can communicate in both directions (as indicated by the arrows). However, in some embodiments, a ring network can allow communication in only one direction (e.g., only clockwise or only counterclockwise). Such embodiments can be used, for example, to simplify the design of each of the network switch circuits in the network switch.
[0098] Proceeding to FIG. 1 to FIG. 11 , a block diagram illustrating one embodiment of a network using a mesh topology to couple the agent circuits 1010A-P is shown. As FIG. 13 shown, the network 1000 can include network switch circuits 1014AA-1014AH. The network switch circuits 1014AA-1014AH are coupled to two or more other network switch circuits. For example, network switch circuit 1014AA is coupled to network switch circuits 1014AB and 1014AE; network switch circuit 1014AB is coupled to network switch circuits 1014AA, 1014AF, and 1014AC; and so on, as FIG. 13 illustrated. Thus, individual network switch circuits in the mesh network can be coupled to different numbers of other network switch circuits. Moreover, while the network 1000 has a relatively symmetrical structure, other mesh networks can be asymmetrical, for example, depending on the various traffic patterns expected to be prevalent on the network. At each network switch circuit 1014AA-1014AH, one or more attributes of a received communication can be used to determine the neighboring network switch circuits 1014AA-1014AH to which the receiving network switch circuit 1014AA-1014AH will send the communication (unless the agent circuit 1010A-P to which the receiving network switch circuit 1014AA-1014AH is coupled is the destination of the communication, in which case the receiving network switch circuit 1014AA-1014AH can terminate the communication on the network 1000 and provide it to the destination agent circuit 1010A-P). For example, in one embodiment, the network switch circuits 1014AA-1014AH can be programmed at system initialization to route communications based on various attributes.
[0099] In embodiments, communications can be routed based on the destination agent. The routing can be configured to transmit communications between a source agent and a destination agent in the mesh topology through a minimum number of network switch circuits ("shortest path") that can be supported. Alternatively, different communications from a given source agent to a given destination agent can take different paths through the mesh network. For example, latency sensitive communications can be sent on a shorter path, while less critical communications can take another path to avoid consuming bandwidth on the short path, where the other path, for example, can be lightly loaded during use. Additionally, the path can change between two particular network switch circuits for different communications at different times. For example, one or more intermediate network switch circuits in a first path used to send a first communication can experience heavy traffic when a second communication is transmitted at a later time. To avoid delays that can result from heavy traffic, the second communication can be routed via a second path that avoids the heavy traffic.
[0100] FIG. 13 Examples can be partially connected mesh networks: at least some communication can pass through one or more intermediate network switching circuits within the mesh network. Fully connected mesh networks can have connections from each network switching circuit to every other network switch, and therefore any communication can be sent without passing through any intermediate network switching circuits. In various implementations, any level of interconnectivity can be used.
[0101] exist FIG. 14 The description of the network interface circuit is included in the section on network switching circuitry. Communication between the network switching circuitry and the network interface can be implemented in various ways. FIG. 14 The diagram illustrates an example of how a network switching circuit can exchange transactions with a network interface.
[0102] Now move to FIG. 14 A block diagram illustrating an implementation of a system with proxy circuitry, a network interface, and a network switch is shown. As illustrated, system 1100 includes proxy circuitry 1120 coupled to a network interface (I / F) 1101, which is in turn coupled to network switching circuitry 1110a. Network switching circuitry 1110a is further coupled to network switching circuitries 1110b and 1110c. Network switching circuitries 1110a to 1110c are included in a communication network 1105. In some implementations, proxy circuitry 1120, network interface 1101, and network switching circuitries 1110a to 1110c may correspond to... The components shown are similarly named and numbered.
[0103] As illustrated, the proxy circuit 1120 can be configured to transmit data transactions to and from the communication network 1105 via the network interface 1101. As disclosed, the communication network 1105 includes network switching circuits 1110a to 1110c for transmitting data transactions between the proxy circuit 1120 and other proxy circuits (not shown) coupled to other network switching circuits in the communication network 1105.
[0104] As shown, the network interface 1101 includes a plurality of input channels (Rx CHs 1145a and 1145b) and one or more output channels (Tx CH 1143). The network interface 1101 can be configured to receive a first plurality of data transactions (data transactions 1140a and 1140b) from the network switch circuit 1110a via the Rx CHs 1145a and 1145b. In some embodiments, the network interface 1101 is further configured to concurrently receive the data transactions 1140a and 1140b from the network switch circuit 1110a. The network interface 1101 can also be configured to transmit a second plurality of data transactions (data transactions 1140c and 1140d) to the network switch circuit 1110a via the Tx CH 1143. As depicted, the network interface 1101 can be further configured to serially transmit the data transactions 1140c and 1140d to the network switch circuit 1110a. To concurrently receive the data transactions 1140a and 1140b (and subsequent data transactions), in some embodiments, the network interface 1101 can include respective transaction buffers coupled to the Rx CHs 1145a and 1145b. Additionally, the network interface 1101 can include an arbitration circuit (arbiter) 1147 to select a data transaction from a given one of the transaction buffers. The arbitration circuit 1147 can use any suitable arbitration scheme to select between the Rx CHs 1145a and 1145b, such as least recently used, round robin, credit, number of buffered transactions, etc. In some embodiments, the Tx CH 1143 can be coupled to respective buffering circuitry, while in other embodiments, the agent circuit 1120 can transmit data transactions to the Tx CH 1143 one at a time.
[0105] In some embodiments, the agent circuit 1120 can be configured to receive data transactions 1140 at a first data rate. For example, the agent circuit 1120 can be a graphics processing unit or a display capable of consuming data transactions at a first rate to support displaying video frames at a particular frame rate and resolution. In such embodiments, failing to receive data transactions at the first data rate can result in video playback stuttering or short burst of glitch. To transmit data transactions 1140 to the agent circuit 1120 at the first data rate, the network switch circuit can be configured to enter a first performance state to serially transmit data transactions 1140 to the agent circuit 1120 at the first data rate and enter a second performance state to concurrently transmit data transactions 1140 to the agent circuit 1120 at the first data rate. For example, to serially transfer data transactions 1140 from the network switch circuit 1110a to the network interface 1101 at the first rate, the communication network 1105 can operate in a performance mode capable of enabling data transfer at the first data rate. The clock frequency in such a performance mode must be high enough to support the serial data rate. However, concurrently transmitting data transactions can allow for slower frequencies (e.g., half the frequency of the first performance mode) to be used because two data transactions can be transferred in parallel. Thus, the second performance state can use less power than the first performance state.
[0106] To avoid congestion at the network switch circuit 1110a when the agent circuit 1120 is transmitting data transactions 1140c and 1140d, the network switch circuit 1110a can include a virtual lane queue (VQ) 1155a configured to receive data transactions 1140c and 1140d from the network interface 1101. The network switch circuit 1110a can also include a virtual lane queue (VQ) 1155b configured to receive a different plurality of data transactions from another network switch circuit (e.g., network switch circuit 1110b) in the communication network 1105. Thus, the network switch circuit 1110a can be capable of concurrently receiving data transactions from the network switch circuit 1110b and the network interface 1101 and placing them into the virtual lane queues 1155b and 1155a, respectively. An arbitration circuit (arbiter) 1157 can be configured to select a data transaction from either of the virtual lane queues 1155a or 1155b using any suitable arbitration algorithm, such as those disclosed above.
[0107] By using a first number of input lanes that is greater than a second number of output lanes, the network interface can be able to pull data transactions out of the communication network 1105 faster than would be the case if the data transactions were received serially. This can alleviate congestion in the communication network 1105, particularly in the network switch circuit 1110a, thereby increasing the bandwidth available for communicating data transactions. Moreover, if the data traffic in the communication network 1105 is not heavy, the communication network 1105 can be able to operate in a lower performance mode, thereby reducing power consumption while maintaining a desired data rate for communicating data transactions to the agent circuit 1120.
[0108] Note that The system 1100 is merely an example. The system 1100 has been simplified to illustrate elements that are relevant to this demonstration. For example, a single agent circuit coupled to a network interface is shown. In other implementations, multiple agent circuits can be coupled to a single network interface. Although one output lane and two input lanes are shown for the network interface 1101, in other implementations, any suitable number of input lanes and output lanes can be included. However, the number of input lanes can be greater than the number of output lanes.
[0109] In summary, various implementations of an apparatus are disclosed that can include one or more first agent circuits configured to communicate data transactions using an in-order protocol and one or more second agent circuits configured to communicate data transactions using a protocol without in-order enforcement. The apparatus can also include one or more input / output (I / O) interfaces coupled to respective ones of the first agent circuits and configured to enforce the in-order protocol. In addition, the apparatus can include a communication network that includes a plurality of network switch circuits. A particular network switch circuit of the plurality of network switch circuits is coupled to at least one other network switch circuit of the plurality of network switch circuits. The apparatus can also include a network interface circuit coupled to the second agent circuits, the I / O interfaces, and the particular network switch circuit. The network interface circuit can be configured to communicate data transactions between the second agent circuits and the particular network switch circuit and to communicate transactions between the I / O interfaces and the particular network switch circuit.
[0110] In another example, a particular first agent circuit can be configured to transmit data transactions to a respective I / O interface using a first protocol. A particular second agent circuit can be configured to transmit data transactions to the network interface circuit using a second protocol that is different from the first protocol. In another example, the respective I / O interface is configured to transmit data transactions received from the particular first agent circuit to the network interface circuit using the second protocol.
[0111] In one example, a particular first agent circuit can be configured to transmit a series of data transactions to a respective I / O interface in a first order. The respective I / O interface can be configured to transmit the series of data transactions to the network interface circuit in the first order. The network interface circuit can be configured to transmit the series of data transactions to the communication network in a second order different from the first order.
[0112] Another apparatus can include a particular integrated circuit including a plurality of agent circuits configured to initiate memory transactions. The particular integrated circuit can also include a plurality of memory circuits configured to respond to the memory transactions. A given memory circuit can be configured to process a given memory transaction based on an address accessed by the given memory transaction. The particular integrated circuit can also include a communication network including one or more first network switch circuits coupled to a first portion of the plurality of agent circuits and one or more second network switch circuits coupled to a second portion of the plurality of memory circuits. The communication network can be configured to communicate the memory transactions via one of two communication pathways. A first communication pathway can include a first suitable subset of the first and second network switch circuits, and a second communication pathway can include a second suitable subset of the first and second network switch circuits. The first suitable subset and the second suitable subset can be mutually exclusive. At least one of the first network switch circuits in the first communication pathway is coupled to a respective one of the first network switch circuits in the second communication pathway. The second network switch circuits in the first communication pathway are isolated from the network switch circuits in the second communication pathway.
[0113] In another example, a plurality of first network switch circuits in the first communication pathway and a plurality of first network switch circuits in the second communication pathway can be coupled to form a mesh subnetwork between the first communication pathway and the second communication pathway. In one example, a particular agent coupled to a particular first network switch circuit in the first communication pathway can be configured to initiate a particular memory transaction for a particular memory circuit coupled to a particular second network switch circuit in the second communication pathway.
[0114] In another example, a particular first network switch circuit can be configured to send a particular memory transaction to a different first network switch circuit in the second communication pathway. The different first network switch circuit can be configured to send the particular memory transaction to a particular second network switch circuit in the second communication pathway.
[0115] Another example of an apparatus can include a proxy circuit configured to communicate data transactions; and a communication network including a plurality of network switch circuits. A particular network switch circuit of the plurality of network switch circuits is coupled to at least one other network switch circuit of the plurality of network switch circuits. The apparatus can also include a network interface circuit coupled to the proxy circuit and the particular network switch circuit, and can include a plurality of input lanes and one or more output lanes. A first number of the input lanes can be greater than a second number of the one or more output lanes. The network interface circuit can be configured to receive a first plurality of data transactions from the particular network switch circuit via the plurality of input lanes; and transmit a second plurality of data transactions to the particular network switch circuit via the one or more output lanes.
[0116] In another example, the network interface circuit can be further configured to concurrently receive the first plurality of data transactions from the particular network switch circuit. In another example, the network interface circuit can be further configured to serially transmit the second plurality of data transactions to the particular network switch circuit.
[0117] In one example, the proxy circuit can be configured to receive data transactions at a first data rate. To transmit data transactions to the proxy circuit at the first data rate, the particular network switch circuit can be configured to enter a first performance state to serially transmit data transactions to the proxy circuit at the first data rate, and enter a second performance state to concurrently transmit data transactions to the proxy circuit at the first data rate. The second performance state can use less power than the first performance state.
[0118] The circuits and techniques described above can be implemented using a variety of methods. The following describes example methods associated with communicating transactions via a network interface. The circuits and techniques described above can be implemented using a variety of methods. The following describes example methods associated with communicating transactions via a network interface.
[0119] Reference will now be made to , which illustrates a flow diagram of an embodiment of a method for communicating data transactions to and from a network interface. The method 1200 can be used in conjunction with any of the systems disclosed herein, such as the system 1100 in . The method 1200 is described below using the system 1100 of as an example. References to elements in are included as non-limiting examples. As depicted, the method 1200 can be performed concurrently with the previously described methods 500 and 600.
[0120] Method 1200 begins at 1210, where the network interface circuit receives a first plurality of data transactions from a particular network switch circuit via a plurality of input channels. For example, network interface 1101 can concurrently receive data transactions 1140a and 1140b from network switch circuit 1110a and can place the received data transactions 1140a and 1140b into respective buffers coupled to each of Rx CHs 1145a and 1145b. Data transactions 1140a and 1140b can each be received at a first data rate per input channel.
[0121] At 1220, method 1200 continues with the network interface circuit communicating the first plurality of data transactions to an agent circuit coupled to the network interface circuit using a second data rate higher than the first data rate. For example, arbitration circuit 1147 can be used to select data transaction 1140a from Rx CH 1145a or data transaction 1140b from Rx CH 1145b. As described above, any suitable arbitration technique can be utilized. Arbitration circuit 1147 can then forward the selected data transaction to agent circuit 1120 at a second data rate higher than the first data rate. Communicating data transactions 1140 from network switch circuit 1110a to network interface 1101 using a lower data rate can allow communication network 1105 to operate in a reduced power state while maintaining a desired data rate for providing data transactions to agent circuit 1120.
[0122] Method 1200 continues at 1230, where the network interface circuit receives a second plurality of data transactions from the agent circuit using the second data rate. For example, agent circuit 1120 can generate data transactions 1140c and 1140d to be communicated to an agent circuit (not shown) also coupled to communication network 1105. Agent circuit 1120 can communicate data transactions 1140c and 1140d to network interface 1101. In some embodiments, network interface 1101 can receive data transactions 1140a and 1140b one at a time. In other embodiments, Tx CH 1143 can include or be coupled to a buffer that allows multiple data transactions to be received and queued for communication.
[0123] At 1240, method 1200 can continue with the network interface circuit communicating the second plurality of data transactions to the particular network switch circuit via one or more output channels. For example, network interface 1101 communicates data transactions 1140a and 1140b to network switch circuit 1110a via Tx CH 1143. Since network interface 1101 has a single output channel (Tx CH 1143) as shown, data transactions 1140a and 1140b are communicated serially. As described above, any suitable arbitration technique can be utilized to select which data transaction to communicate next. The first number of input channels (Rx CH 1145) is greater than the second number of one or more output channels (Tx CH 1143), as exemplified. In some embodiments, the agent circuit 1120 can not generate as many data transactions as it receives. For example, if the agent circuit 1120 is a display, the agent circuit 1120 can receive many large data transactions corresponding to a frame of image data to be displayed, while transmitting an appropriate amount of data transactions to, for example, reply to status requests, and / or indicate that too much or too little image data is being received. By including additional input channels, data transactions can be offloaded from the communication network 1105 more quickly, freeing bandwidth within the communication network 1105. Since the agent circuit 1120 can not generate as many data transactions as it consumes, additional output channels can not be implemented in some embodiments.
[0124] Note that the method 1200 includes blocks 1210-1240. The method 1200 can end in block 1240, or some or all of the blocks of the method can be repeated. For example, the method 1200 can repeat operations 1210 and 1220 to receive a large number of data transactions. The method 1200 can be performed concurrently with the methods 500 and 600. The method 1200 can also be performed concurrently with another instance of the method 1200. For example, a network interface circuit can be coupled to multiple agent circuits, and have multiple input channels for two or more of the agent circuits.
[0125] Circuits and methods are exemplified for systems, such as integrated circuits, that include various network interfaces and I / O interfaces for passing transactions between agent circuits in a communication network. Any embodiment of the disclosed systems can be included in one or more computer systems of various types of computer systems, such as desktop computers, laptop computers, smartphones, tablet computers, and wearable devices, among others. In some embodiments, the above-described circuits can be implemented on a system on a chip (SoC) or other type of integrated circuit, including multi-die packages. A block diagram exemplifying an embodiment of a system 1300 is exemplified in FIG. 13. In some embodiments, the system 1300 can include any disclosed embodiment of the systems disclosed herein, such as the systems 100, 400, 700, 800, and 1100 shown in various figures in
[0126] In the illustrated implementation, system 1300 includes at least one instance of a system on a chip (SoC) 1306, which can include multiple types of processor circuitry such as central processing units (CPUs), graphics processing units (GPUs), or others, communication fabric, and interfaces to memory and input / output devices. SoC 1306 can correspond to an instance of the SoCs disclosed herein. In various implementations, SoC 1306 is coupled to external memory circuitry 1302, peripherals 1304, and a power supply 1308.
[0127] A power supply 1308 is also provided, which supplies a supply voltage to SoC 1306 and one or more supply voltages to external memory circuitry 1302 and / or peripherals 1304. In various implementations, power supply 1308 represents a battery (e.g., a rechargeable battery in a smartphone, laptop, or tablet computer, or other device). In some implementations, more than one instance of SoC 1306 is included (and also more than one external memory circuitry 1302).
[0128] External memory circuitry 1302 is any type of memory such as dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate (DDR, DDR2, DDR3, etc.) SDRAM (including mobile versions of SDRAM such as mDDR3, etc., and / or low power versions of SDRAM such as LPDDR2, etc.), RAMBUS DRAM (RDRAM), static RAM (SRAM), etc. In some implementations, external memory circuitry 1302 can include non-volatile memory such as flash memory, ferroelectric random access memory (FRAM), or magnetoresistive RAM (MRAM). One or more memory devices can be coupled to a circuit board to form a memory module such as a single inline memory module (SIMM), dual inline memory module (DIMM), etc. Alternatively, these devices can be mounted with the SoC or integrated circuit in a chip stack configuration, package stack configuration, or multi-chip module configuration.
[0129] Peripherals 1304 include any desired circuitry, depending on the type of system 1300. For example, in one implementation, peripherals 1304 include devices for various types of wireless communication such as Wi-Fi, Bluetooth, cellular phone, global positioning system, etc. In some implementations, peripherals 1304 also include additional storage devices including RAM storage devices, solid state storage devices, or disk storage devices. Peripherals 1304 include user interface devices such as a display screen including a touch display screen or multi-touch display screen, a keyboard or other input device, a microphone, a speaker, etc.
[0130] As illustrated, the system 1300 is shown to have applications in a wide range of domains. For example, the system 1300 can be used as part of a chip, circuit, component, etc. of a desktop computer 1310, a laptop computer 1320, a tablet computer 1330, a cellular or mobile phone 1340, or a television 1350 (or a set-top box coupled to a television). A smart watch and health monitoring device 1360 are also illustrated. In some embodiments, the smart watch can include a variety of general-purpose computing related functionality. For example, the smart watch can provide access to email, cell phone service, a user's calendar, etc. In various embodiments, the health monitoring device can be a dedicated medical device or otherwise include dedicated health related functionality. In various embodiments, the smart watch described above can or can not include some or any health monitoring related functionality. Other wearable devices 1360 are also contemplated, such as devices worn around the neck, devices attached to hats or other headgear, devices capable of being implanted in a human body, glasses designed to provide an augmented and / or virtual reality experience, etc.
[0131] The system 1300 can also be used as part of a cloud-based service 1370. For example, the previously mentioned devices and / or other devices can access computing resources in the cloud (i.e., remotely located hardware and / or software resources). Still further, the system 1300 can be used in one or more devices of a home 1380 other than the previously mentioned devices. For example, home appliances can monitor and detect noteworthy conditions. Various devices in the home (e.g., a refrigerator, a cooling system, etc.) can monitor the state of the device and provide an alert to the homeowner (or, for example, a repair agency) in the event that a particular event is detected. Alternatively, a thermostat can monitor the temperature in the home and can automate adjustments to the heating / cooling system based on a history of reactions to various conditions by the homeowner. Applications of the system 1300 to various modes of transportation 1390 are also illustrated. For example, the system 1300 can be used in control and / or entertainment systems for airplanes, trains, buses, taxicabs, private cars, watercraft ranging from private boats to cruise ships, scooters (for rent or private), etc. In various cases, the system 1300 can be used to provide automated guidance (e.g., self-driving vehicles) and general system control, etc.
[0132] It is noted that various potential applications of the system 1300 can include various performance, cost, and power consumption requirements. Thus, scalable solutions that enable the use of one or more integrated circuits to provide suitable combinations of performance, cost, and power consumption can be beneficial. These and numerous other embodiments are possible and are contemplated. The illustrated devices and applications are merely exemplary and are not intended to be limiting. Other devices are possible and are contemplated.
[0133] As disclosed with respect to System 1300 can include one or more integrated circuits that are included within a personal computer, a smart phone, a tablet computer, or other type of computing device. Processes for designing and producing integrated circuits using design information are presented below in Designing and Producing Integrated Circuits
[0134] is a block diagram illustrating an example of a non-transitory computer-readable storage medium storing circuit design information in accordance with some embodiments. Embodiments of the system 1300 can be used in processes for designing and manufacturing integrated circuits, for example, including one or more instances of the systems 100, 400, 700, 800, and 1100 (or portions thereof) described above. In the illustrated embodiment, a semiconductor manufacturing system 1420 is configured to process design information 1415 stored on a non-transitory computer-readable storage medium 1410 and manufacture an integrated circuit 1430 based on the design information 1415.
[0135] The non-transitory computer-readable storage medium 1410 can include any of a variety of appropriate types of memory devices or storage devices. The non-transitory computer-readable storage medium 1410 can be an installation medium, such as a CD-ROM, floppy disks, or tape device; computer system memory or random access memory such as DRAM, DDR RAM, SRAM, EDO RAM, Rambus RAM, etc.; non-volatile memory such as flash memory, magnetic media, e.g., a hard disk drive or optical storage; registers, or other like type of storage elements, etc. The non-transitory computer-readable storage medium 1410 can also include other types of non-transitory memory or combinations thereof. The non-transitory computer-readable storage medium 1410 can include two or more memory media that can reside in different locations, such as on different computer systems that are connected over a network.
[0136] Design information 1415 can be specified using any of a variety of appropriate computer languages, including hardware description languages such as, but not limited to, VHDL, Verilog, SystemC, SystemVerilog, RHDL, M, MyHDL, etc. Design information 1415 can be capable of being used by a semiconductor fabrication system 1420 to fabricate at least a portion of integrated circuit 1430. For example, the format of design information 1415 can be recognized by at least one semiconductor fabrication system, such as semiconductor fabrication system 1420. In some embodiments, design information 1415 can include a netlist specifying elements of a cell library and their connectivity. One or more cell libraries used during logic synthesis of circuits included in integrated circuit 1430 can also be included in design information 1415. Such cell libraries can include information indicative of device or transistor level netlists, mask design data, and characterization data, etc. of cells included in the cell library.
[0137] In various embodiments, integrated circuit 1430 can include one or more custom macro cells, such as memory, analog or mixed-signal circuits, etc. In such cases, design information 1415 can include information related to the included macro cells. Such information can include, but is not limited to, schematic capture databases, mask design data, behavioral models, and device or transistor level netlists. As used herein, mask design data can be formatted according to Graphic Data System (GDSII) or any other suitable format.
[0138] Semiconductor fabrication system 1420 can include any of a variety of appropriate elements configured to fabricate integrated circuits. This can include, for example, elements for depositing semiconductor materials, removing materials, altering the shape of deposited materials, modifying materials (e.g., by doping materials or using ultraviolet treatment to modify the dielectric constant), etc. (e.g., on a wafer that can include masks). Semiconductor fabrication system 1420 can also be configured to perform various testing of fabricated circuits to enable correct operation.
[0139] In various embodiments, integrated circuit 1430 is configured to operate according to a circuit design specified by design information 1415, which can include performing any of the functionality described herein. For example, integrated circuit 1430 can include any of the various elements shown or described herein. Additionally, integrated circuit 1430 can be configured to perform the various functions described herein in connection with other components.
[0140] As used herein, the phrase “design information specifying a design of a circuit configured to...” does not imply that the circuit in question must be fabricated in order to satisfy that element. Rather, the phrase indicates that the design information describes a circuit that, when fabricated, will be configured to perform the indicated action or will include the specified components.
[0141] The disclosure includes reference to “an embodiment” or “embodiments” (e.g., “some embodiments” or “various embodiments”). An embodiment is a different specific implementation or example of the disclosed concepts. References to “an embodiment,” “one embodiment,” and “the particular embodiment” or similar terms, do not necessarily refer to the same embodiment. Numerous embodiments are contemplated, including those specifically disclosed, as well as modifications or alternative arrangements within the scope of the disclosure.
[0142] The disclosure can discuss potential advantages of the disclosed embodiments. Not all specific embodiments of the disclosure will necessarily achieve any or all of the potential advantages. Whether a particular specific embodiment achieves one or more of the potential advantages can depend on the particular implementation, the context in which the implementation is deployed, and the relative significance of each potential advantage to the overall objectives of the embodiment. In fact, there can be numerous reasons why a particular specific embodiment can not demonstrate some or all of the potential advantages disclosed herein. For example, a particular implementation can include other circuits, in addition to those disclosed herein, which negate or mitigate one or more of the disclosed advantages. Further, sub-optimal design performance of a particular implementation (e.g., implementation technology or tools) can also negate or mitigate the disclosed advantages. Even assuming a technically sound implementation, achievement of an advantage can depend on other factors, such as the circumstances surrounding the implementation of the embodiment. For example, inputs provided to a particular implementation can prevent the occurrence of one or more problems addressed in the disclosure, and as a result, the benefits of its solutions can not be realized. In view of the potential for factors outside the disclosure to affect the realization of any potential advantage, it is hereby expressly
[0143] Unless otherwise stated, embodiments are non-limiting. That is, disclosed embodiments are not intended to limit the scope of claims drafted based on the disclosure, even if only a single example is described in relation to a particular feature. Disclosed embodiments are intended to be illustrative rather than restrictive, without any requirement in the disclosure that any particular feature be included with any other particular feature. Accordingly, the present application is intended to permit claims covering the disclosed embodiments, as well as such alternatives, modifications, and equivalents, as would be apparent to one skilled in the art having the benefit of the present disclosure.
[0144] For example, features in the present application can be combined in any suitable manner. Accordingly, new claims can be formed during prosecution of the present application (or an application claiming priority from it) directed to any such combination of features. Specifically, dependent claims refer to features in the appended claims as appropriate, and dependent claims can be combined with features of other dependent claims if appropriate. Similarly, features from corresponding independent claims can be combined if appropriate.
[0145] Accordingly, while the appended claims can be drafted in the singular form for the purpose of each claim, the claims in the plural form are also intended to cover each of the dependent claims. The combination of features of the dependent claims in accordance with the disclosure can also be claimed in this application or another application. In short, combinations are not limited to those specifically recited in the appended claims.
[0146] It is also contemplated where appropriate, that a claim drafted in one format or statutory class (e.g., means plus function) is intended to cover corresponding claims in another format or statutory class (e.g., method).
[0147] As the present disclosure is a legal document, various terms and phrases can be subject to administrative and judicial interpretation. It is hereby given notice that the following paragraphs, as well as the definitions provided throughout the present disclosure, will be used to determine how claims drafted based on the present disclosure are interpreted.
[0148] Unless the context clearly dictates otherwise, a reference to an item in the singular (i.e., a noun or noun phrase preceded by “a”, “an” or “the”) is intended to mean “one or more”. Thus, a reference in a claim to “an item” does not exclude additional instances of the item. A plurality of items is a collection of two or more items.
[0149] The word “can” is used herein in the permissive sense (i.e., having the potential to, being able to), rather than the mandatory sense (i.e., must).
[0150] The terms “comprise” and “include” and variations thereof are open-ended, and intend to mean “including but not limited to”.
[0151] When the term "or" is used in the disclosure with respect to a list of options, it will be understood, unless otherwise provided in the context, that it is used in an inclusive sense, i.e., as meaning one or the other or both. Thus, the expression "x or y" is equivalent to "x or y, or both," and thus covers 1) x but not y, 2) y but not x, and 3) both x and y. On the other hand, phrases such as "one of x or y, but not both" make it clear that "or" is used in an exclusive sense.
[0152] The expression "w, x, y, or z, or any combination thereof" or "at least one of... w, x, y, and z" is intended to cover all of the possible combinations of the elements in the set. For example, given the set [w, x, y, z], these phrases cover any single element of the set (e.g., w, but not x, y, or z), any two elements (e.g., w and x, but not y or z), any three elements (e.g., w, x, and y, but not z), and all four elements. The phrase "at least one of... w, x, y, and z" thus refers to at least one element of the set [w, x, y, z], thereby covering all possible combinations of the elements in the list. This phrase is not to be construed as requiring the presence of at least one instance of w, at least one instance of x, at least one instance of y, and at least one instance of z.
[0153] In the present disclosure, various "labels" can precede a noun or noun phrase. Unless otherwise provided by the context, different labels used for a feature (e.g., "first circuit," "second circuit," "particular circuit," "given circuit," etc.) refer to different instances of the feature. Additionally, unless otherwise noted, the labels "first," "second," and "third" when applied to a feature do not imply any type of ordering (e.g., spatial, temporal, logical, etc.).
[0154] The phrase "based on" is used to describe one or more factors that affect a determination. This term does not foreclose the possibility that additional factors can affect a determination. That is, a determination can be solely based on specified factors or based on specified factors and other, unspecified factors. Consider the phrase "determine A based on B." This phrase specifies that B is a factor that is used to determine A. This phrase does not foreclose the possibility that the determination of A can also be based on some other factor, such as C. This phrase is also intended to cover an implementation where A is determined based only on B. As used herein, the phrase "based on" is synonymous with the phrase "based, at least in part, on."
[0155] The phrases "in response to" and "in response" describe one or more factors that trigger an effect. The phrases do not exclude the possibility that additional factors can influence or otherwise trigger the effect, either jointly or independently of the specified factors. That is, the effect can be in response to only those factors, or can be in response to the specified factors as well as other, unspecified factors. Consider the phrase "perform A in response to B." This phrase specifies that B is a factor that triggers performance of A, or that triggers a particular result of A. The phrase does not exclude the possibility that A can also be performed in response to some other factor, such as C. The phrase also does not exclude the possibility that A can be performed in response to both B and C, jointly. The phrase is also intended to cover an implementation in which A is performed in response to B alone. As used herein, the phrase "in response" is synonymous with the phrase "at least partially in response to." Similarly, the phrase "in response to" is synonymous with the phrase "at least partially in response to."
[0156] Within the present disclosure, different entities (which can be variously referred to as "units," "circuits," other components, etc.) can be described or claimed as "configured to" perform one or more tasks or operations. This manner of description is used herein to generally refer to the structure being so configured as to perform the task(s) in question. More specifically, such "configured to" language can refer to a structure being so configured so as to perform a task during operation. Thus, a structure can be considered to be "configured to" perform some task where the structure has been modified, or will be modified, to perform the task during operation. Accordingly, the "configured to" language used herein indicates that a structure has been configured (e.g., altered from a core or generic configuration) to perform the task in question during operation. Thus, "configured to" can be understood as describing a structure that is altered or adapted in some manner, to perform the task in question during operation. The structure can be "configured to" perform the task in question during operation by being programmed to perform the task, if the structure includes a processor and / or memory; in addition to being programmed to perform the task, the structure can be "configured to" perform the task in question during operation by including one or more other components that cause or contribute to the structure performing the task during operation, either alone or in combination with other components.
[0157] In some cases, various units / circuits / components can be described herein as performing a set of tasks or operations. It should be understood that these entities "configured to" perform those tasks / operations, even if not specifically indicated.
[0158] The term "configured to" is used herein to mean that an entity was purposefully designed or made to perform one or more tasks in a specified manner. For example, a processor can be "configured to" perform a task in a specified manner by executing one or more programs encoded on a storage medium.
[0159] For purposes of United States patent practice, the mere statement of a structure "configured to" perform one or more tasks in a claim is expressly not an attempt to invoke 35 U.S.C. § 112(f). If an applicant desires to invoke section 112(f) for any reason, it will use the "means for" construction in the claim.
[0160] Different “circuits” can be described in this disclosure. These circuits or “circuitry” constitute hardware that includes various types of circuit elements, such as combinatorial logic, clocked storage devices (e.g., flip-flops, registers, latches, etc.), finite state machines, memory (e.g., random access memory, embedded dynamic random access memory), programmable logic arrays, etc. The circuitry can be custom-designed, or taken from a standard library. In various implementations, the circuitry can include digital components, analog components, or a combination of both, as appropriate. Certain types of circuitry can be commonly referred to as “units” (e.g., decode units, arithmetic logic units (ALUs), functional units, memory management units (MMUs), etc.). Such units also refer to circuitry or circuitry.
[0161] Accordingly, the disclosed circuits / units / components and other elements illustrated in the drawings and described herein include hardware elements, such as those described in the preceding paragraph. In many cases, the internal arrangement of a hardware element within a particular circuit can be specified by describing the functionality of that circuit. For example, a particular “decode unit” can be described as performing the function of “processing the operation code of an instruction and routing the instruction to one or more of a plurality of functional units,” which means that the decode unit is “configured to” perform that function. To the skilled artisan in computer hardware, this functional specification is sufficient to imply a set of possible structures for the circuitry.
[0162] In various embodiments, circuits, units, and other elements can be defined by the functions or operations that they are configured to implement, as discussed in the preceding paragraph. The arrangement of relative to one another and the manner in which such circuits / units / components interact, as well as the way they are programmed, define the microarchitecture of the hardware, which is ultimately fabricated in integrated circuits or programmed into FPGAs to form the physical instantiation of the microarchitecture definition. The microarchitecture definition is therefore considered by those skilled in the art to be the structure from which many physical instantiations can be derived, all falling within the broader structure described by the microarchitecture definition. That is, a skilled artisan provided with the microarchitecture definition according to the present disclosure can implement the structure by coding the description of the circuits / units / components in a hardware description language (HDL) such as Verilog or VHDL, without undue experimentation and with the application of ordinary skill. The HDL description is often expressed in a manner that can appear functional. But to the skilled artisan, the HDL description is a way to transform the structure of the circuits, units, or components into next-level instantiation details. Such HDL descriptions can take the form of behavioral code (which is often not synthesizable), register-transfer language (RTL) code (which is often synthesizable as compared to behavioral code), or structural code (e.g., netlists that specify logic gates and their interconnections). The HDL description can be synthesized sequentially against a cell library designed for a given integrated circuit fabrication technology, and can be modified for timing, power, and other reasons to arrive at a final design database that is sent out to a foundry to generate masks and ultimately produce integrated circuits. Some hardware circuits or portions thereof can also be custom-designed in a schematic editor and captured into the integrated circuit design along with the synthesized circuits. The integrated circuits can include transistors and other circuit elements (e.g., passive elements such as capacitors, resistors, inductors, etc.), as well as interconnects between the transistors and circuit elements. Some embodiments can implement multiple integrated circuits coupled together to implement the hardware circuits, and / or can use discrete elements in some embodiments. Alternatively, the HDL design can be synthesized for a programmable logic array such as a field-programmable gate array (FPGA), and can be implemented in the FPGA. This decoupling between the design of a set of circuits and the subsequent lower-level instantiations of those circuits often results in scenarios in which the circuit or logic designers never specify a particular set of structures for the lower-level instantiations beyond what is described for what the circuits are configured to do, as that process is performed at a different stage of the circuit implementation process.
[0163] The fact that the same functionality of a circuit can be implemented using many different low-level combinations of circuit elements results in a large number of equivalent structures of that circuit. As noted, these low-level circuit implementations can vary according to variations in manufacturing technology, foundry chosen for fabricating the integrated circuit, cell library provided for a particular project, and the like. In many cases, the selection of these different implementations can be arbitrary by different design tools or methodologies.
[0164] Furthermore, for a given implementation, a single implementation of a particular functional specification of a circuit typically includes a large number of devices (e.g., millions of transistors). Thus, the sheer volume of this information makes it impractical to provide a complete recitation of the low-level structures used to implement a single implementation, much less the large number of equivalent possible implementations. To this end, the present disclosure describes the structure of a circuit using functional shorthand commonly used in the industry.
Claims
1. An apparatus, the apparatus comprising: Multiple proxy circuits, the multiple proxy circuits including: One or more first proxy circuits, the one or more first proxy circuits being configured to use an ordered protocol to deliver data transactions; One or more second proxy circuits, the one or more second proxy circuits being configured to use a protocol without mandatory ordering to deliver data transactions; One or more input / output (I / O) interfaces, said one or more input / output (I / O) interfaces being coupled to a corresponding first proxy circuit in the first proxy circuit and configured to enforce the ordered protocol; A communication network comprising a plurality of network switching circuits, wherein a particular network switching circuit is coupled to at least one other network switching circuit; and A network interface circuit, coupled to the second proxy circuit, the I / O interface, and the specific network switching circuit, and configured to: Transmitting data transactions between the second proxy circuit and the specific network switching circuit; and Data transactions are transferred between the I / O interface and the specific network switching circuit.
2. The apparatus of claim 1, wherein a specific first proxy circuit is configured to transmit data transactions to a corresponding I / O interface using a first protocol; and A specific second proxy circuit is configured to transmit data transactions to the network interface circuit using a second protocol different from the first protocol.
3. The apparatus of claim 2, wherein the corresponding I / O interface is configured to transmit data transactions received from the specific first proxy circuit to the network interface circuit using the second protocol.
4. The apparatus of claim 1, wherein a specific first proxy circuit is configured to transmit a series of data transactions to the corresponding I / O interface in a first sequence; The corresponding I / O interface is configured to transmit the series of data transactions to the network interface circuit in the first sequence; and The network interface circuit is configured to transmit the series of data transactions to the communication network in a second sequence, different from the first sequence.
5. The apparatus of claim 4, wherein the network interface circuitry is configured to transmit responses to the series of data transactions to the corresponding I / O interfaces in a third sequence, the third sequence corresponding to the order in which the responses are received from the communication network; and The corresponding I / O interface is configured to transmit the response to the series of data transactions to the specific first proxy circuit in the first sequence.
6. The apparatus of claim 5, wherein, in order to transmit a response to the series of data transactions to the specific first proxy circuit in the first sequence, the corresponding I / O interface is configured to buffer responses received from the network interface circuit in the third sequence.
7. The apparatus according to claim 1, further comprising: One or more third proxy circuits, the one or more third proxy circuits being configured to use the ordered protocol to deliver data transactions; One or more additional I / O interfaces, which are coupled to a corresponding third proxy circuit in the third proxy circuit and configured to enforce the ordered protocol; and Different network interface circuits are coupled to the additional I / O interface and to different network switching circuits among the plurality of network switching circuits.
8. The apparatus according to claim 1, further comprising: Multiple memory circuits; and First communication path and second communication path; The plurality of network switching circuits include one or more first network switching circuits coupled to a corresponding proxy circuit in a first portion of the plurality of proxy circuits, and one or more second network switching circuits coupled to a corresponding memory circuit in the plurality of memory circuits. The first communication path includes a first appropriate subset of the first network switching circuit and the second network switching circuit; and The second communication path includes a second appropriate subset of the first network switching circuit and the second network switching circuit, wherein the first appropriate subset and the second appropriate subset are mutually exclusive.
9. The apparatus of claim 8, wherein at least one first network switching circuit in the first network switching circuit of the first communication path is coupled to a corresponding first network switching circuit in the first network switching circuit of the second communication path; and The second network switching circuit in the first communication path is isolated from the network switching circuit in the second communication path.
10. The apparatus of claim 1, wherein the network interface circuit includes a plurality of input channels and one or more output channels, wherein a first number of input channels is greater than a second number of the one or more output channels, and wherein the network interface circuit is further configured to: Receive first plurality of data transactions from the specific network switching circuit via the plurality of input channels; and A second plurality of data transactions are transmitted to the specific network switching circuit via the one or more output channels.
11. A method, the method comprising: The first proxy circuit uses an ordered protocol to transmit the first set of data transactions to the input / output (I / O) interface circuit; The second proxy circuit transmits the second set of data transactions to the network interface circuit using a non-forced order protocol. The I / O interface circuit uses the ordered protocol to transmit the first set of data transactions to the network interface circuit. as well as The network interface circuit transmits the first set of data transactions and the second set of data transactions to the communication architecture in an order based on the availability of the corresponding destination.
12. The method according to claim 11, further comprising: The network interface circuit receives responses to the first set of data transactions and the second set of data transactions from the communication architecture in a receiving order based on the corresponding destination response time.
13. The method of claim 12, further comprising: The network interface circuit transmits the response to the first group of data transactions to the I / O interface circuit based on the receiving order; as well as The network interface circuit transmits the response to the second set of data transactions to the second proxy circuit based on the receiving order.
14. The method according to claim 13, further comprising: The I / O interface circuit uses the ordered protocol to transmit the response to the first group of data transactions to the first proxy circuit.
15. The method according to claim 11, further comprising: The network interface circuit receives a third set of data transactions from the communication architecture via multiple input channels, wherein the data transactions are received at a first data rate per input channel. The network interface circuit transmits the third set of data transactions to a specific proxy circuit at a second data rate higher than the first data rate. The network interface circuit receives a fourth set of data transactions from the specific proxy circuit using the second data rate. as well as The fourth set of data transactions is transmitted to the communication architecture via the network interface circuit through one or more output channels, wherein the first number of input channels is greater than the second number of the one or more output channels.
16. A system comprising: Multiple proxy circuits, the multiple proxy circuits including: A first proxy circuit is configured to transmit a first group of data transactions in the first order using an ordered protocol. The second proxy circuit is configured to use a protocol without mandatory ordering to transmit the second set of data transactions; An input / output (I / O) interface, which is coupled to the first proxy circuit and configured to: The first group of data transactions is received in the first order; and The first sequence is used to transmit the first group of data transactions; A communication network, comprising multiple network switching circuits; and A network interface circuit, and the network interface circuit is configured to: The first group of data transactions is received from the I / O interface in the first sequence; Transmit the first group of data transactions without mandatory ordering to a specific network switching circuit among the plurality of network switching circuits; Receive the second set of data transactions from the second proxy circuit; and The second set of data transactions is transmitted to the specific network switching circuit without mandatory ordering.
17. The system of claim 16, further comprising a plurality of memory circuits coupled to a first portion of the plurality of network switching circuits, wherein a portion of the plurality of proxy circuits is coupled to a second portion of the plurality of network switching circuits; and The communication network mentioned above includes: A first communication path, the first communication path comprising a first appropriate subset of the network switching circuits of the first part and the second part; and A second communication path, comprising a second appropriate subset of the network switching circuits of the first and second portions, wherein the first appropriate subset and the second appropriate subset are mutually exclusive.
18. The system of claim 17, wherein at least one network switching circuit in the network switching circuit of the second portion of the first communication path is coupled to a corresponding network switching circuit in the network switching circuit of the second portion of the second communication path; and The network switching circuit of the first part of the first communication path is isolated from the network switching circuits of the first part and the second part of the second communication path.
19. The system of claim 16, wherein the network interface circuit is configured to: Responses to the first set of data transactions and the second set of data transactions are received from the specific network switching circuit in the order of reception, the order of reception being based on the corresponding destination response time; as well as Based on the receiving order, the response to the first group of data transactions is transmitted to the I / O interface; as well as The response to the second set of data transactions is transmitted to the second proxy circuit based on the receiving order.
20. The system of claim 19, wherein the I / O interface is further configured to: The responses to the first set of data transactions are received from the network interface circuit in the order based on the receiving order; and The response to the first group of data transactions is transmitted to the first agent circuit in the first sequence.