Arbitration of portions of transactions via virtual channels associated with interconnects
The described arbitration system addresses latency issues in AMBA by arbitrating transaction portions per clock cycle and transmitting them through virtual channels, improving SoC interconnect efficiency and performance.
Patent Information
- Application Number
- JP2024039915
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-02-04
- Filing Date
- 2024-03-14
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2039-03-19
AI Technical Summary
The Advanced Microcontroller Bus Architecture (AMBA) standard for System on Chip (SoC) interconnects suffers from latency issues due to transaction-based arbitration, where the bus is controlled for the entire duration of a transaction, preventing cycle-by-cycle arbitration and leading to reduced efficiency and performance.
An arbitration system that arbitrates among portions of multiple transactions per clock cycle, selecting a winning portion and transmitting it through virtual channels associated with the interconnect, allowing for interleaved transmission of multiple transactions across multiple clock cycles.
This approach reduces waiting time and improves the efficiency and utilization of the interconnect by enabling cycle-by-cycle arbitration and interleaved transmission of transaction portions, enhancing SoC performance.
Smart Images

Figure 0007717210000005 
Figure 0007717210000006 
Figure 0007717210000007
Abstract
Description
Technical Field
[0001] Cross-reference to related applications This application claims priority based on U.S. Provisional Patent Application No. 62 / 650,589 (PR TIP001P), filed on Mar. 30, 2018, and U.S. Provisional Application No. 62 / 800,897 (PRT1P003P), filed on Feb. 4, 2019. Each of these priority claim basis applications is hereby incorporated by reference in its entirety for all purposes.
[0002] This application relates to arbitrating access to shared resources, and more particularly to arbitrating among portions of multiple transactions and transmitting a winning portion via one of a plurality of virtual channels associated with an interconnect on a per clock cycle basis.
Background Art
[0003] A system on chip (“SoC”) is an integrated circuit that includes multiple subsystems, and is often referred to as an intellectual property (“IP”) agent or core. An IP agent is typically a “reusable” block of circuitry designed to implement or perform a particular function. Developers of SoCs typically layout and interconnect their IP agents on a chip such that the multiple IP agents communicate with each other. By using IP agents, the time and cost of developing complex SoCs can be significantly reduced.
[0004] One problem faced by developers of SoCs is interconnecting the various IP agents on a chip such that they operate with each other. To address this problem, semiconductor The industry has been adapting interconnect standards.
[0005] One such standard was developed and popularized by ARM Ltd. in Cambridge, UK. It is the Advanced Microcontroller Bus Architecture (AMBA). AM BA is a widely used bus interconnect standard for connecting and managing functional IP agents on a SoC.
[0006] In AMBA, a transaction defines a request and requires a separate response transaction. In a write transaction, the source requests to write data to a remote destination. When the write operation is executed, the destination sends back a confirmation response transaction to the source. The write operation is considered complete only when the response transaction has been received by the source. In a read transaction, the source requests access to read a remote location. The read transaction is complete only when the response transaction (i.e., the accessed content) has been returned to the source.
[0007] In AMBA, arbitration processing is used to permit access to the interconnecting bus among multiple competing transactions. During a given arbitration cycle, one of the competing transactions is selected as the winner. The interconnecting bus is then controlled for the duration of the data portion of the winning transaction. The next arbitration cycle starts only after all the data for the current transaction has been completed. This process handles multiple outstanding transactions competing for access to the interconnect. It is continuously repeated on the condition that there is Yon.
[0008] One problem with the AMBA standard is latency. The bus interconnect is arbitrated transaction by transaction. Between any part of the transaction, regardless of whether the transaction is a read, write, or response, the bus is controlled for the entire transaction. Once a transaction starts, it cannot be interrupted. For example, if the data part of a transaction is 4 cycles long, all 4 cycles must complete before another transaction can access the bus. As a result, (1) transactions cannot be arbitrated cycle by cycle, and (2) all non-winning competing transactions must wait until the data part of the current transaction is complete. Both of these factors tend to reduce the efficiency of the interconnect and the overall performance of the SoC.
[0009]
Summary of the Invention
[0010] By repeatedly executing (1) to (3), a plurality of winning portions are each interleaved over a plurality of clock cycles and transmitted via a plurality of virtual channels. For each clock cycle, arbitration of a plurality of transaction portions via a plurality of virtual channels is performed, resulting in many advantages such as reduction of waiting time, and improvement of the efficiency and utilization rate of the interconnection. Due to these attributes, the arbitration system and method disclosed in this specification are very suitable for arbitrating access to the interconnection on a system-on-chip (SoC). BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The present application and its advantages will be best understood by reference to the following description in connection with the accompanying drawings.
[0012]
Figure 1
[0013]
Figure 2
[0014]
Figure 3A
[0015]
Figure 3B
[0016]
Figure 4
[0017]
Figure 5
[0018]
Figure 6
[0019]
Figure 7
[0020]
Figure 8
[0021]
Figure 9A
[0022]
Figure 9B
[0023]
Figure 10A
Figure 10B
[0024]
Figure 11A
Figure 11B
[0025] In the drawings, like reference numerals may be used to designate like structural elements. It should be understood that the depictions in the figures are schematic and not necessarily to scale. desired.
DETAILED DESCRIPTION OF THE INVENTION
[0026] Hereinafter, a detailed description of the present application will be given with reference to several non-exclusive embodiments illustrated in the accompanying drawings. In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, as will be apparent to those skilled in the art, the present disclosure can be practiced without some or all of these specific details. Also, in order to avoid obscuring the present disclosure unnecessarily, well-known processing steps and / or structures are not described in detail. Further, in order to avoid obscuring the present disclosure unnecessarily, detailed descriptions of well-known processing steps and / or structures are omitted.
[0027] Many of the integrated circuits currently under development are very complex. As a result, many chip designers have interconnected multiple subsystems or IP agents on a single piece of silicon using a system-on-chip or "SoC" approach. Consumer devices (e.g. consumer devices (e.g. , handheld, mobile phone, tablet computer, laptop and desktop Computers, media processing, etc.), virtual or augmented reality (e.g., robotics, autonomous Vehicles, aircraft, etc.), medical devices (e.g., imaging, etc.), industrial, home automation -tion, industrial (e.g., smart home appliances, home monitoring devices, etc.) and data centers Applications (e.g., network switches, connected storage devices, etc.), etc., various SoCs for various applications are currently available or being developed.
[0028] This application generally targets arbitration systems and methods for arbitrating access to shared resources. Such shared resources can be, for example, bus interconnects, memory resources, processing resources, or pretty much any other resource shared among multiple competing parties. For the sake of convenience of explanation, the shared resources detailed below are assumed to be the interconnects shared by multiple subsystems on a system-on-chip, i.e., "SoC". In an SoC, as will be detailed later, there are multiple subsystems that exchange traffic with each other in the form of transactions, and the shared resource is a physical interconnect, and various transactions or parts thereof are transmitted through a plurality of virtual channels associated with the shared interconnect, and one of a plurality of different arbitration schemes and / or priorities is used to arbitrate access to the shared interconnect for the transmission of transactions between sub-functions.
[0029] In an SoC, as will be detailed later, there are multiple subsystems that exchange traffic with each other in the form of transactions, and the shared resource is a physical interconnect, and various transactions or parts thereof are transmitted through a plurality of virtual channels associated with the shared interconnect, and one of a plurality of different arbitration schemes and / or priorities is used to arbitrate access to the shared interconnect for the transmission of transactions between sub-functions. is used to arbitrate access to the shared interconnect for the transmission of transactions between sub-functions. is used to arbitrate access to the shared interconnect for the transmission of transactions between sub-functions. is used to arbitrate access to the shared interconnect for the transmission of transactions between sub-functions.
[0030] Transaction class Within the above-described shared interconnect used in the SoC, there are at least three types or classes of transactions including Posted (P), Non-pos ted (NP), and Completion (C). Simple definitions for each are provided in Table 1 below.
Table 1
[0031] Posted transactions (such as writes) do not require a response transaction. When the source writes data to the specified destination, the transaction ends. Non -posted transactions (such as either reads or writes) require a response. However, the response branches off as a separate Completion transaction. In other words, for a read, the first transaction is used for the read operation, and a separate but related Completion transaction is used to read and return the content. For a Non-posted write, the first transaction is used for the write, while upon completion of the write, a second related Co mpletion transaction is requested for confirmation.
[0032] Transactions, regardless of type, can be represented by one or more packets. In some situations, a transaction can be represented by a single packet. In other situations, multiple packets may be required to represent the entire transaction.
[0033] A beat is the amount of data that can be transmitted per clock cycle over the shared interconnect. For example, if the shared interconnect is physically 128 bits wide, 128 bits can be transmitted per bit or clock cycle.
[0034] In some situations, a transaction may need to be split into multiple parts for transmission. Consider a transaction with a single packet having a payload of 512 bits (64 bytes). If the shared interconnect is only 128 bits wide (16 bytes), the transaction needs to be split into 4 parts (e.g., 4 × 128 = 512) and transmitted in 4 clock cycles or beats. On the other hand, if the transaction is only a single packet less than 128 bits wide, the entire transaction can be sent in 1 clock cycle or beat. If the same transaction happens to include additional packets, additional clock cycles or beats may be required.
[0035] Thus, the term "part" of a transaction is the amount of data that can be transferred over the shared interconnect during a given clock cycle or beat. The size of a part can vary depending on the physical width of the shared interconnect. For example, if the shared interconnect is physically 64 data bits wide, the maximum number of bits that can be transferred during any one cycle or beat is 64 bits. If a given transaction has a payload of 64 bits or less, the entire transaction can be sent over the shared interconnect in a single part. On the other hand, if the payload is larger, the packet must be sent over the shared interconnect in multiple parts. For a transaction with a payload of 128, 256, or 512 bits Each requires 2, 4, and 8 parts. Thus, the term "part" should be interpreted broadly to mean either part or all of a transaction that can be transmitted via a shared interconnect during any given clock cycle or beat. A stream is defined as a pairing of virtual channels and transaction classes. For example, if there are four virtual channels (e.g., VC0, VC1, VC2, and VC3 ), as well as three transaction classes (P, NP, C), there are up to 1 2 different possible streams. The various combinations of virtual channels and transaction classes are detailed in Table 2 below.
[0036] Stream It should be noted that the number of transaction classes described above is merely illustrative and should not be construed as limiting. Conversely, any number of virtual channels and / or transaction classes may be used. ) and three transaction classes (P, NP, C), there are up to 1 2 different possible streams. The various combinations of virtual channels and transaction classes are detailed in Table 2 below. Refer to Figure 1, which shows a block diagram of arbitration system 10. In a non-exclusive embodiment, the arbitration system is used to arbitrate access to shared interconnect 12 by a plurality of sub-functions 14 (i.e., IP1, IP2, and I
Table 2
[0037] P3) that attempt to send transactions to upstream sub-functions 14 (i.e., IP4, IP5, and IP6). P3) that attempt to send transactions to upstream sub-functions 14 (i.e., IP4, IP5, and IP6). It should be noted that the number of transaction classes described above is merely illustrative and should not be construed as limiting. Conversely, any number of virtual channels and / or transaction
[0038] Arbitration in virtual channels of shared interconnect Referring to Figure 1, a block diagram of arbitration system 10 is shown. In a non-exclusive embodiment, the arbitration system is used to arbitrate access to shared interconnect 12 by a plurality of sub-functions 14 (i.e., IP1, IP2, and I P3) that attempt to send transactions to upstream sub-functions 14 (i.e., IP4, IP5, and IP6). P3) that attempt to send transactions to upstream sub-functions 14 (i.e., IP4, IP5, and IP6). P3) that attempt to send transactions to upstream sub-functions 14 (i.e., IP4, IP5, and IP6). It is used to arbitrate access to shared interconnect 12 by a plurality of sub-functions 14 (i.e., IP1, IP2, and I
[0039] The shared interconnect 12 is a physical interconnect that is N data bits wide and includes M control bits. Also, the shared interconnect 12 is unidirectional, which means it handles traffic only in the direction from the sources (i.e., IP1, IP2, and IP3) to the destinations (i.e., IP4, IP5, and IP6).
[0040] In various alternatives, the number of N data bits may be any integer, but typically is a bit width that is a power of 2 (e.g., 21, 22, 23, 24, 25, 26, 27, 28, 29, etc.) or (2, 4, 6, 8, 16, 32, 64, 128, 2 56, etc.). In the most realistic applications, the number of N bits is either 32, 64, 128, 256, or 512. However, it should be understood that these widths are merely illustrative and should not be construed as limiting in any way.
[0041] The number of control bits M can also vary and can be any number.
[0042] One or more logical channels (not shown) (hereinafter referred to as "virtual channels" or "VCs") are associated with the shared interconnect 12. Each virtual channel is independent. Each virtual channel may be associated with multiple independent streams. The number of virtual channels can vary widely. For example, up to 32 or more virtual channels may be defined or associated with the shared interconnect 12.
[0043] In various alternative embodiments, each virtual channel may be assigned a different priority. A higher priority may be assigned to one or more virtual channels, while a lower priority may be assigned to one or more other virtual channels. The high-priority channels are given or arbitrated for access to the shared interconnect 12 with higher rights than the low-priority virtual channels. In another embodiment, the same priority may be given to each of the virtual channels, in which case, when granting or arbitrating access rights to the shared interconnect 12, no virtual channel is prioritized over another virtual channel. In yet another embodiment, the priority assigned to one or more of the virtual channels may change dynamically. For example, in a first set of situations, the same priority may be assigned to all virtual channels, but in a second set of situations, a higher priority may be assigned to a particular virtual channel than to other virtual channels. Thus, as the situation changes, the priority scheme used among the virtual channels can be changed to best suit the current operating conditions. A lower priority may be assigned to one or more other virtual channels. The high-priority channels are given or arbitrated for access to the shared interconnect 12 with higher rights than the low-priority virtual channels. In another embodiment, the same priority may be given to each of the virtual channels, in which case, when granting or arbitrating access rights to the shared interconnect 12, no virtual channel is prioritized over another virtual channel. In yet another embodiment, the priority assigned to one or more of the virtual channels may change dynamically. For example, in a first set of situations, the same priority may be assigned to all virtual channels, but in a second set of situations, a higher priority may be assigned to a particular virtual channel than to other virtual channels. Thus, as the situation changes, the priority scheme used among the virtual channels can be changed to best suit the current operating conditions. Each of subsystem 14 is typically a "reusable" circuit or logic block, generally referred to as an IP core or agent. Most IP agents are designed to perform a specific function, for example, a controller for peripheral devices such as an Ethernet port, a display driver, an SDRAM interface, a USB port, etc. Such IP agents generally provide the necessary subsystem functions within the overall design of a complex system provided on an integrated circuit (IC) such as an application specific integrated circuit (ASIC) or a field programmable gate array (FPGA). In another embodiment, the same priority may be given to each of the virtual channels, in which case, when granting or arbitrating access rights to the shared interconnect 12, no virtual channel is prioritized over another virtual channel. In yet another embodiment, the priority assigned to one or more of the virtual channels may change dynamically. For example, in a first set of situations, the same priority may be assigned to all virtual channels, but in a second set of situations, a higher priority may be assigned to a particular virtual channel than to other virtual channels. Thus, as the situation changes, the priority scheme used among the virtual channels can be changed to best suit the current operating conditions. Each of subsystem 14 is typically a "reusable" circuit or logic block, generally referred to as an IP core or agent. Most IP agents are designed to perform a specific function, for example, a controller for peripheral devices such as an Ethernet port, a display driver, an SDRAM interface, a USB port, etc. Such IP agents generally provide the necessary subsystem functions within the overall design of a complex system provided on an integrated circuit (IC) such as an application specific integrated circuit (ASIC) or a field programmable gate array (FPGA).
[0044] Each of subsystem 14 is typically a "reusable" circuit or logic block, generally referred to as an IP core or agent. Most IP agents are designed to perform a specific function, for example, a controller for peripheral devices such as an Ethernet port, a display driver, an SDRAM interface, a USB port, etc. Such IP agents generally provide the necessary subsystem functions within the overall design of a complex system provided on an integrated circuit (IC) such as an application specific integrated circuit (ASIC) or a field programmable gate array (FPGA). Most IP agents are designed to perform a specific function, for example, a controller for peripheral devices such as an Ethernet port, a display driver, an SDRAM interface, a USB port, etc. Such IP agents generally provide the necessary subsystem functions within the overall design of a complex system provided on an integrated circuit (IC) such as an application specific integrated circuit (ASIC) or a field programmable gate array (FPGA). Each of subsystem 14 is typically a "reusable" circuit or logic block, generally referred to as an IP core or agent. Most IP agents are designed to perform a specific function, for example, a controller for peripheral devices such as an Ethernet port, a display driver, an SDRAM interface, a USB port, etc. Such IP agents generally provide the necessary subsystem functions within the overall design of a complex system provided on an integrated circuit (IC) such as an application specific integrated circuit (ASIC) or a field programmable gate array (FPGA). Each of subsystem 14 is typically a "reusable" circuit or logic block, generally referred to as an IP core or agent. Most IP agents are designed to perform a specific function, for example, a controller for peripheral devices such as an Ethernet port, a display driver, an SDRAM interface, a USB port, etc. Such IP agents generally provide the necessary subsystem functions within the overall design of a complex system provided on an integrated circuit (IC) such as an application specific integrated circuit (ASIC) or a field programmable gate array (FPGA). It is used as a "routing block (component)". By using the available IP agent library, the chip designer can easily "bolt on" various logic functions in the design of more complex integrated circuits, thus reducing the design time and saving the development cost at the same time. The subsystem agent 14 has been described above in relation to the dedicated IP core, but it should be understood that this is not a requirement. On the contrary, the subsystem 14 may be a collection of IP functions connected to or shared by a single port 20. Therefore, the term "agent" should be broadly interpreted as any subsystem of IP connected to port 20, whether the subsystem performs a single function or multiple functions. By using the available IP agent library, the chip designer can easily "bolt on" various logic functions in the design of more complex integrated circuits, thus reducing the design time and saving the development cost at the same time. The subsystem agent 14 has been described above in relation to the dedicated IP core, but it should be understood that this is not a requirement. On the contrary, the subsystem 14 may be a collection of IP functions connected to or shared by a single port 20. Therefore, the term "agent" should be broadly interpreted as any subsystem of IP connected to port 20, whether the subsystem performs a single function or multiple functions. is connected to or shared by a single port 20. Therefore, the term "agent" should be broadly interpreted as any subsystem of IP connected to port 20, whether the subsystem performs a single function or multiple functions. A pair of switches 16 and 18 each provide access between the respective subsystem agent 14 and the shared interconnect 12 via a dedicated access port 20. In the exemplary embodiment of the figure, (1) Subsystem agents IP1, IP2, and IP3 are each connected to switch 16 via access ports Port0, Port1, and Port2. is connected to port 20, whether the subsystem performs a single function or multiple functions.
[0045] A pair of switches 16 and 18 each provide access between the respective subsystem agent 14 and the shared interconnect 12 via a dedicated access port 20. In the exemplary embodiment of the figure, (2) Subsystem agents IP4, IP5, and IP6 are each connected to switch 18 via Ports Port3, Port4, and Port5. (3) Further, access port 22 provides access to subsystem agents IP4, IP5, and IP6 to switch 16 as a whole via interconnect 12. (1) Subsystem agents IP1, IP2, and IP3 are each connected to switch 16 via access ports Port0, Port1, and Port2. (2) Subsystem agents IP4, IP5, and IP6 are each connected to switch 18 via Ports Port3, Port4, and Port5. (3) Further, access port 22 provides access to subsystem agents IP4, IP5, and IP6 to switch 16 as a whole via interconnect 12. switches 16 and 18 perform multiplexing and demultiplexing functions. Switch 16 (3) Further, access port 22 provides access to subsystem agents IP4, IP5, and IP6 to switch 16 as a whole via interconnect 12. switches 16 and 18 perform multiplexing and demultiplexing functions. Switch 16
[0046] performs multiplexing and demultiplexing functions. Switch 16 Generated by subsystem agents IP1, IP2, and / or IP3 Selects upstream traffic and sends the traffic downstream via shared interconnect 12. At switch 18, a demultiplexing operation is performed and the traffic is provided to the destination subsystem agent (i.e., either IP4, IP5, or IP6).
[0047] Each access port 20 has a unique port identifier (ID) and provides dedicated access for each subsystem agent 14 to either switch 16 or 18. For example subsystem agents IP1, IP2, and IP3 are each assigned to access ports Port0, Port1, and Port2, respectively. Similarly, subsystem agents IP4, IP5, and IP6 are each assigned to access ports Port3, Port4, and Port5, respectively.
[0048] In addition to providing ingress and egress points to / from switches 16, 18, the unique port ID 20 is used to address traffic between subsystem agents 14. Each port 20 has a specific amount of addressable space allocated within system memory 24.
[0049] In some non-exclusive embodiments, a "global" port identifier may be assigned to all or a portion of the access ports 20 in addition to the unique port ID. Transactions and other traffic are assigned to the address associated with the global port identifier It can be sent to all or part of the access ports. Therefore, by using a global identifier, transactions and other traffic can be widely transmitted or broadcast to all or part of the access port 20, and the need to individually address each access port 20 using a unique identifier can be eliminated.
[0050] The switch 16 further includes an arbitration element 26, an address resolution logic (AR L) 28, and an address resolution lookup table (LUT) 30.
[0051] During operation, the subsystem agents IP1, IP2, and IP3 generate transactions. Each time a transaction is generated, it is packetized by the transmitting subsystem 14, and then the packetized transaction is input into the local switch 16 via the corresponding port 20. For example, portions of the transactions generated by IP1, IP2, and IP3 are provided to the switch 16 via Port0, Port1, and Port2, respectively. Each port 20 includes a plurality of first-in first-out buffers (not shown) for each of the virtual channels associated with the interconnect channel 12. In a non-exclusive embodiment, there are four virtual channels. In that case, one for each virtual channel, and each port 20 includes four buffers. Again, it should be understood that the number of virtual channels and buffers included in the port 20 can vary and is not limited to four. Conversely, the number of virtual channels and buffers can be more or less than four.
[0052]
[0053] If a given transaction is represented in two (or more) parts, those parts are kept in the same buffer. For example, if interconnect 12 has a 128 data-bit width and the transaction is represented by a packet containing a 512-bit payload, the transaction needs to be split into four parts that are transmitted in four clock cycles or beats. On the other hand, if the transaction can be represented by a single packet with a 64-bit payload, a single part can be transmitted in one clock cycle or beat. By keeping all parts of a given transaction in the same buffer, the virtual channels remain logically independent. In other words, all traffic related to a given transaction is always sent on the same virtual channel as the stream and is not branched across multiple virtual channels.
[0054] The arbitration element 26 is responsible for arbitrating between the competing buffered parts of the transactions maintained by the various access ports 20. In a non-exclusive embodiment, if multiple competing transactions are available, the arbitration element 26 performs arbitration every clock cycle. The arbitration winner for each cycle generates the part of the transaction that is granted access to the interconnect 12 and transmitted over the interconnect 12 from one of subsystems IP1, IP2, and IP3.
[0055] When generating a transaction, the source subsystems IP1, IP2, and IP3 It usually knows the addresses within the address space for possible destination subsystem agents IP4, IP5, and IP6, but does not know the information (e.g., port IDs 20 and / or 22) necessary to route a transaction to the destination. In one embodiment, the local address resolution logic (ARL) 28 is used to resolve the known destination address to the necessary routing information. In other words, the source subsystem agent 14 may simply know that it wants to access a given address in the system memory 24. Thus, the ARL 28 is tasked with accessing the LUT 30 and performing an address lookup for port 20 / 22 along the delivery path to the ultimate destination corresponding to the specified address. When the port 20 / 22 is known, this information is inserted into the destination field in the packet of the transaction. As a result, the packet is delivered to port 20 / 22 along the delivery path. In principle, since the required delivery information is already known and included in the destination field of the packet, downstream nodes along the delivery path do not need to perform further lookups. In another type of transaction, called source-based routing (SBR) which will be described in detail later, the source P agent knows the destination port address. As a result, the lookup performed by the ARL 28 typically does not need to be executed. It usually knows the addresses within the address space for possible destination subsystem agents IP4, IP5, and IP6, but does not know the information (e.g., port IDs 20 and / or 22) necessary to route a transaction to the destination. In one embodiment, the local address resolution logic (ARL) 28 is used to resolve the known destination address to the necessary routing information. In other words, the source subsystem agent 14 may simply know that it wants to access a given address in the system memory 24. Thus, the ARL 28 is tasked with accessing the LUT 30 and performing an address lookup for port 20 / 22 along the delivery path to the ultimate destination corresponding to the specified address. When the port 20 / 22 is known, this information is inserted into the destination field in the packet of the transaction. As a result, the packet is delivered to port 20 / 22 along the delivery path. In principle, since the required delivery information is already known and included in the destination field of the packet, downstream nodes along the delivery path do not need to perform further lookups. In another type of transaction, called source-based routing (SBR) which will be described in detail later, the source P agent knows the destination port address. As a result, the lookup performed by the ARL 28 typically does not need to be executed. It usually knows the addresses within the address space for possible destination subsystem agents IP4, IP5, and IP6, but does not know the information (e.g., port IDs 20 and / or 22) necessary to route a transaction to the destination. In one embodiment, the local address resolution logic (ARL) 28 is used to resolve the known destination address to the necessary routing information. In other words, the source subsystem agent 14 may simply know that it wants to access a given address in the system memory 24. Thus, the ARL 28 is tasked with accessing the LUT 30 and performing an address lookup for port 20 / 22 along the delivery path to the ultimate destination corresponding to the specified address. When the port 20 / 22 is known, this information is inserted into the destination field in the packet of the transaction. As a result, the packet is delivered to port 20 / 22 along the delivery path.
[0056] In an alternative embodiment, not all nodes within the interconnection need the ARL 28 and the LUT 30. For nodes without these elements, the necessary routing information Transactions without routing information can be forwarded to the default node. The default node accesses the ARL 28 and the LUT 30, and then the necessary routing information can be inserted into the header of the transaction packet. The default node typically is upstream of nodes that do not have the ARL 28 and the LUT 30. However this is not necessarily the case. One or more default nodes can be located anywhere on the SoC Excluding the ARL 28 and the LUT 30 from some nodes can reduce the complexity of the nodes.
[0057] In addition to decoding the destination for the winning part of the transaction, the ARL 28 defines the order for the winning part of the transaction within each virtual channel, so it may be called an "ordering point". Each time an arbitration is resolved, whether or not the ARL 28 is used to perform an address port lookup, the winning part of the transaction is inserted into the first-in first-out queue provided to each virtual channel. Then, the winning part of the transaction waits for the order of transmission via the interconnect 12 within the buffer.
[0058] Also, the ARL 28 is used to define "upstream" and "downstream" traffic. In other words, any transaction generated by the IP agents 14 (i.e., IP1, IP2, and IP3) associated with the switch 16 is considered to be upstream of the ARL 28. All transactions after the ARL 28 (i.e., transmitted to IP4, IP5, and IP6) The cushion is regarded as downstream traffic.
[0059] The IP agents 14 associated with switch 16 (i.e., IP1, IP2 , and IP3) may communicate with each other, either directly or indirectly, to send transactions to each other. Through direct communication (often called source-based routing (SBR)), the IP agents 14 can send transactions to each other in a peer-to-peer model. In this model, the source IP agent knows the unique port ID of its peer IP agent 14, eliminating the need to use ARL28 to access the LUT30. Alternatively, transactions between the IP agents associated with switch 16 may be routed using ARL28. In this model, as described above, the source IP agent only knows the address of the destination IP agent 14 and does not know the information necessary for routing. Then, ARL28 accesses the LUT30 and is used to find the corresponding port ID, and thereafter, the port ID is inserted into the destination field of the transaction packet.
[0060] Packet format The IP agents 14 generate transactions and process them through virtual channels associated with the interconnect 12. Each transaction typically consists of one or more packets. Each packet typically has a fixed header size and format. In some examples, each packet may have a fixed-size payload. In another example, the packet payload may be of various sizes from large to small, and also There may be no payload at all.
[0061] Referring to FIG. 2, an example 32 of a packet is shown. Packet 32 includes a header 34 and a payload 36. In this particular embodiment, the header 34 is 16 bytes in size. This size is exemplary, and packets of larger sizes (e.g., more bytes count) or smaller sizes (e.g., fewer bytes count) may be used. It should be understood that not all of the headers 34 of the packets 32 necessarily have the same size. In an alternative embodiment, the size of the packet header in the SoC may be variable.
[0062] The header 34 includes a plurality of fields such as a destination identifier (DST_ID), a source identifier (SRC_ID), a payload size indicator (PLD_SZ), a reserved field (RSVD), a command field (CMD), a TAG field, a status (STS), a transaction ID field (TAG), an address or ADDR field, a USDR / compact payload field, a transaction class or TC field, a format FMT field, and a byte enable (BE) field. The various fields of the header 34 will be briefly described in Table 3 below.
Table 3
[0063] The payload 36 includes the content of the packet. The size of the payload may vary. In some examples, the payload may be large. In other examples, the pay The load may be small. In yet another example, if the content is very small, i.e., " compact", it can be carried within the USRD field of the header 34.
[0064] The type of transaction often indicates whether one or more packets used to represent the transaction have a payload. For example, for either a Posted or Non-Posted read, the packet specifies the location address to be accessed, but typically has no payload. However, the packets of the related Completion transaction contain a payload that includes the read content. For both Posted and Non-Posted write transactions, the packets contain a payload that includes the data to be written to the destination. In the case of a Non-Posted write, the packet of the Completion transaction typically does not define a payload. However, in some situations, the Completion transaction
[0065] defines a payload. It should be understood that the examples of packets and the above description cover many of the basic fields that may be included in a packet. Additional fields may be deleted or added.
[0066] Arbitration Referring to FIG. 3A, a Peripheral Component Interconnect (PCI) Shows arbitration logic executed by the arbitration element 26 A logic diagram is shown.
[0067] In PCI ranking, each port 20 has a separate buffer for each combination of virtual channel and transaction class (P, NP, and C). For example, if there are four virtual channels (VC0, VC01, VC2, and VC3), Port0 , Port1, and Port2 each have 12 first-in first-out buffers. In other words for each port 20, the buffers are provided for each combination of transaction class (P, NP, and C) and virtual channel (VC0, VC1, VC2, and VC30).
[0068] When each IP agent 14 (e.g., IP1, IP2, and IP3) generates a transaction, the resulting packets are respectively placed in the appropriate buffer within the corresponding port (e.g , port 0, port 1, and port 2) based on the transaction type. For example, Posted (P), Non-posted (NP), and Completion (C) transactions generated by IP1 are respectively placed in the Posted, No n-posted, and Completion buffers for the assigned virtual channel within port 0. Transactions generated by IP2 and IP3 are placed in a similar manner within port 1 and port 2 in the Posted, Non-posted, and , Completion buffers for the assigned virtual channel.
[0069] When a given transaction is represented by multiple packets, all packets of that transaction are inserted into the same buffer. As a result, all packets of the transaction are ultimately transmitted via the same virtual channel. In this policy, the virtual channels remain independent, which means that different virtual channels are not used for the transmission of multiple packets related to the same transaction. Within each port 20, packets can be assigned to a given virtual channel in many different ways. For example, the assignment may be random. Alternatively, the assignment may be based on the workload for each virtual channel and the amount of outstanding traffic. If one channel is very busy and other channels are not, port 20 often tries to balance the load and assigns newly generated transaction traffic to a virtual channel with low utilization. As a result, the routing efficiency is improved. In yet another alternative, transaction traffic may be assigned to a specific virtual channel based on urgency, security, or a combination of both. If a specific virtual channel is given a higher priority and / or security than other virtual channels, high-priority and / or secure traffic is assigned to the higher-priority virtual channel. In yet another embodiment, port 20 may be hard-coded, which means that port 20 has only one virtual channel and all traffic generated by port 20 is transmitted via that one virtual channel. When a given transaction is represented by multiple packets, all packets of that transaction are inserted into the same buffer. As a result, all packets of the transaction are ultimately transmitted via the same virtual channel. In this policy, the virtual channels remain independent, which means that different virtual channels are not used for the transmission of multiple packets related to the same transaction. Within each port 20, packets can be assigned to a given virtual channel in many different ways. For example, the assignment may be random. Alternatively, the assignment may be based on the workload for each virtual channel and the amount of outstanding traffic. If one channel is very busy and other channels are not, port 20 often tries to balance the load and assigns newly generated transaction traffic to a virtual channel with low utilization. As a result, the routing efficiency is improved.
[0070] Within each port 20, packets can be assigned to a given virtual channel in many different ways. For example, the assignment may be random. Alternatively, the assignment may be based on the workload for each virtual channel and the amount of outstanding traffic. If one channel is very busy and other channels are not, port 20 often tries to balance the load and assigns newly generated transaction traffic to a virtual channel with low utilization. As a result, the routing efficiency is improved. In yet another alternative, transaction traffic may be assigned to a specific virtual channel based on urgency, security, or a combination of both. If a specific virtual channel is given a higher priority and / or security than other virtual channels, high-priority and / or secure traffic is assigned to the higher-priority virtual channel. In yet another embodiment, port 20 may be hard-coded, which means that port 20 has only one virtual channel and all traffic generated by port 20 is transmitted via that one virtual channel. When a given transaction is represented by multiple packets, all packets of that transaction are inserted into the same buffer. As a result, all packets of the transaction are ultimately transmitted via the same virtual channel. In this policy, the virtual channels remain independent, which means that different virtual channels are not used for the transmission of multiple packets related to the same transaction. Within each port 20, packets can be assigned to a given virtual channel in many different ways. For example, the assignment may be random. Alternatively, the assignment may be based on the workload for each virtual channel and the amount of outstanding traffic. If one channel is very busy and other channels are not, port 20 often tries to balance the load and assigns newly generated transaction traffic to a virtual channel with low utilization. As a result, the routing efficiency is improved. In yet another alternative, transaction traffic may be assigned to a specific virtual channel based on urgency, security, or a combination of both. If a specific virtual channel is given a higher priority and / or security than other virtual channels, high-priority and / or secure traffic is assigned to the higher-priority virtual channel. In yet another embodiment, port 20 may be hard-coded, which means that port 20 has only one virtual channel and all traffic generated by port 20 is transmitted via that one virtual channel. When a given transaction is represented by multiple packets, all packets of that transaction are inserted into the same buffer. As a result, all packets of the transaction are ultimately transmitted via the same virtual channel. In this policy, the virtual channels remain independent, which means that different virtual channels are not used for the transmission of multiple packets related to the same transaction. Within each port 20, packets can be assigned to a given virtual channel in many different ways. For example, the assignment may be random. Alternatively, the assignment may be based on the workload for each virtual channel and the amount of outstanding traffic. If one channel is very busy and other channels are not, port 20 often tries to balance the load and assigns newly generated transaction traffic to a virtual channel with low utilization. As a result, the routing efficiency is improved. In yet another alternative, transaction traffic may be assigned to a specific virtual channel based on urgency, security, or a combination of both. If a specific virtual channel is given a higher priority and / or security than other virtual channels, high-priority and / or secure traffic is assigned to the higher-priority virtual channel. In yet another embodiment, port 20 may be hard-coded, which means that port 20 has only one virtual channel and all traffic generated by port 20 is transmitted via that one virtual channel. When a given transaction is represented by multiple packets, all packets of that transaction are inserted into the same buffer. As a result, all packets of the transaction are ultimately transmitted via the same virtual channel. In this policy, the virtual channels remain independent, which means that different virtual channels are not used for the transmission of multiple packets related to the same transaction. Within each port 20, packets can be assigned to a given virtual channel in many different ways. For example, the assignment may be random. Alternatively, the assignment may be based on the workload for each virtual channel and the amount of outstanding traffic. If one channel is very busy and other channels are not, port 20 often tries to balance the load and assigns newly generated transaction traffic to a virtual channel with low utilization. As a result, the routing efficiency is improved. In yet another alternative, transaction traffic may be assigned to a specific virtual channel based on urgency, security, or a combination of both. If a specific virtual channel is given a higher priority and / or security than other virtual channels, high-priority and / or secure traffic is assigned to the higher-priority virtual channel. In yet another embodiment, port 20 may be hard-coded, which means that port 20 has only one virtual channel and all traffic generated by port 20 is transmitted via that one virtual channel. When a given transaction is represented by multiple packets, all packets of that transaction are inserted into the same buffer. As a result, all packets of the transaction are ultimately transmitted via the same virtual channel. In this policy, the virtual channels remain independent, which means that different virtual channels are not used for the transmission of multiple packets related to the same transaction. Within each port 20, packets can be assigned to a given virtual channel in many different ways. For example, the assignment may be random. Alternatively, the assignment may be based on the workload for each virtual channel and the amount of outstanding traffic. If one channel is very busy and other channels are not, port 20 often tries to balance the load and assigns newly generated transaction traffic to a virtual channel with low utilization. As a result, the routing efficiency is improved. In yet another alternative, transaction traffic may be assigned to a specific virtual channel based on urgency, security, or a combination of both. If a specific virtual channel is given a higher priority and / or security than other virtual channels, high-priority and / or secure traffic is assigned to the higher-priority virtual channel. In yet another embodiment, port 20 may be hard-coded, which means that port 20 has only one virtual channel and all traffic generated by port 20 is transmitted via that one virtual channel.
[0071] In yet another embodiment, the virtual channel assignment may be performed by the source IP agent 1 4, either alone or in cooperation with the corresponding port 20. For example, the source IP agent 14 generates a control signal to the corresponding port 20 requiring that the packets of a given transaction be assigned to a specific virtual channel. The IP agent 14 can also make assignment decisions that are random, hard-coded, or balanced utilization across all virtual channels, based on security, urgency, etc., as described above.
[0072] In the selection of the arbitration winner, the arbitration element 26 executes a plurality of arbitration steps per cycle. These arbitration steps include the following. (1) A step of selecting a port, (2) A step of selecting a virtual channel, and (3) A step of selecting a transaction class. (2) A step of selecting a virtual channel, and (3) A step of selecting a transaction class.
[0073] The above order (1), (2), and (3) is not fixed. Conversely, the above three steps may be completed in any order. Regardless of the order used, a single arbitration winner is selected for each cycle. The winning transaction is then transmitted via the corresponding virtual channel associated with the interconnect 12. The winning transaction is then transmitted via the corresponding virtual channel associated with the interconnect 12.
[0074] For each arbitration (1), (2), and (3) performed by the arbitration element 26, a plurality of arbitration schemes or rule sets are used For each arbitration (1), (2), and (3) performed by the arbitration element 26, a plurality of arbitration schemes or rule sets are used Such arbitration schemes may include strict or absolute priority, four hypothetical Weighted virtual channels, each assigned a certain percentage of transaction traffic. Priority, or the routing by which transactions are assigned to virtual channels in a predetermined order. In further embodiments, other priority schemes may be used. The arbitration element 26 may also use different arbitration may dynamically switch between schemes from time to time, and / or (1), (2), and (3) the same or different arbitrations for each of the arbitrations. It should be understood that each scheme may be used.
[0075] In an optional embodiment, the unconsidered arbitration results during a given arbitration cycle are The availability of the destination port 20 defined by the transaction of the transaction is taken into account. The buffer in destination port 20 is available to process a given transaction. If it does not have the resources, the corresponding virtual channel will not be available. Transactions do not compete for arbitration, but rather wait until the target resource is available. Meanwhile, if the target resource is available, the request waits until the next arbitration cycle. If so, the corresponding transaction is arbitrated and the access to the interconnect 12 is granted. Competing for Seth.
[0076] The availability of the destination port 20 is determined by the multiple arbitration steps (1), (2) described above. ), and (3) may be checked at different times. For example, availability checks The check is performed before the arbitration cycle (i.e., steps (1), (2), and It can be executed before the completion of any of (3). As a result, only the transactions that define the available destination resources are considered during subsequent arbitration. Alternatively, the arbitration check can be executed between any of the three arbitration steps (1), (2), and (3), regardless of the order in which the arbitration process is executed. There are advantages and disadvantages to performing the destination resource availability check early or late during the arbitration process. Executing the check early can potentially eliminate parts of the transaction that may conflict if their destinations are not available. However, knowing the availability early can generate a significant amount of overhead on system resources. As a result, depending on the situation, it may be more practical to perform the availability check later during a given arbitration cycle. For the arbitration process, including the selection of transaction classes, multiple rules are defined to arbitrate between the conflicting parts of N, NP, and C transactions. These rules include the following. For Posted (P) transactions,
[0077] - The Posted transaction part cannot overtake another Posted transaction part. - The Posted transaction part must be able to overtake the Non-posted transaction part to avoid deadlocks.
[0078] - The Posted transaction part must not overtake Completion if both are in the strong order mode. In other words, in the strong mode, the transaction must be executed strictly according to the rules, and the rules cannot be relaxed. - The Posted request is allowed to overtake Completion if any transaction part has its Relaxed Order (RO) bit set, but overtaking is not mandatory. In the relaxed order, generally the rules are observed, but exceptions may be allowed. - For Non-posted (NP) transactions, - The Non-posted transaction part must not overtake the Posted transaction part. - The Non-posted transaction part must not overtake another Non-posted transaction part. - The Non-posted transaction part must not overtake Completion if both are in the strong order mode. - The Non-posted transaction part is allowed to overtake Completion if any transaction part has its RO bit set, but it is not mandatory. - For Completion (C) transactions, - Completion must not overtake the Posted transaction part if both are in the strong order mode. - Completion is allowed to overtake any transaction part that has its RO bit set, but it is not mandatory. - The Posted transaction part must not overtake Completion if both are in the strong order mode. In other words, in the strong mode, the transaction must be executed strictly according to the rules, and the rules cannot be relaxed. - The Posted request is allowed to overtake Completion if any transaction part has its Relaxed Order (RO) bit set, but overtaking is not mandatory. In the relaxed order, generally the rules are observed, but exceptions may be allowed. - For Non-posted (NP) transactions, - The Non-posted transaction part must not overtake the Posted transaction part. - The Non-posted transaction part must not overtake another Non-posted transaction part. - The Non-posted transaction part must not overtake Completion if both are in the strong order mode. - The Non-posted transaction part is allowed to overtake Completion if any transaction part has its RO bit set, but it is not mandatory. - For Completion (C) transactions, - Completion must not overtake the Posted transaction part if both are in the strong order mode. - Completion is allowed to overtake any transaction part that has its RO bit set, but it is not mandatory. - The Posted transaction part must not overtake Completion if both are in the strong order mode. In other words, in the strong mode, the transaction must be executed strictly according to the rules, and the rules cannot be relaxed. - The Posted request is allowed to overtake Completion if any transaction part has its Relaxed Order (RO) bit set, but overtaking is not mandatory. In the relaxed order, generally the rules are observed, but exceptions may be allowed. When doing so, it is permitted to overtake the Posted transaction part, but it is mandatory It is not. - Completion must not overtake the Non-posted transaction part when both are in strong order mode. - Completion is permitted to overtake the Non-posted transaction part when any transaction part has its RO bit set but it is not mandatory. - Completion is not permitted to overtake another Completion.
[0079] Table 4 below provides an overview of the PCI ordering rules. In boxes without the options (a) and (b), the strict ordering rules do not have to be followed. In boxes of the table with the options (a) and (b), depending on whether the RO bit is reset or set, either the strict order (a) rule or the relaxed order (b) rule may be applied respectively. In various alternative embodiments, the RO bit may be set or reset globally or individually at the packet level.
Table 4
[0080] Arbitration element 26 selects the final winning transaction part by performing arbitration on the respective contending ports 20, virtual channels, and transaction classes without a specific order. The winning part per cycle accesses the shared interconnect 12 and is transmitted via the corresponding virtual channel.
[0081] Referring to FIG. 3B, a logic diagram showing the arbitration logic executed by the arbitration element 26 in device ranking is shown. The arbitration process and, presumably, the consideration of available destination resources are basically the same as those described above, except for two differences.
[0082] First, in device ranking, (a) Non- posted read or write transactions for which responses are required for all requests, and (b) Completion transactions that specify the requested responses, include only two transaction classes. Since there are only two transaction classes, there are only two buffers for each virtual channel at each port 20. For example, if there are four virtual channels (VC0, VC 1, VC2, and VC3), each port 20 (e.g., Port0, Por t1, and Port2) has a total of eight buffers.
[0083] Second, the rules for selecting device-ranked transactions are also different from PCI ranking. In device ranking, there are no strict rules applied to the selection of one class over another. Conversely, any transaction class can be arbitrarily selected. However, in a general way, typically, a favorable Completion transaction is requested so as to release resources that may not be available until the Completion transaction is resolved.
[0084] In other respects, the arbitration process for device ranking is basically the same as above. It is the same as described above. In other words, for each arbitration cycle, arbitration steps (1), (2), and (3) are executed in an arbitrary specific order to select an arbitration winner. When transaction class arbitration is executed, the device order is used rather than the PCI order rule. Further, the availability of the destination resource and / or virtual channel may be considered either before or during any of arbitration steps (1), (2), and (3). (2), and (3).
[0085] Operation flow chart As described above, the above arbitration scheme may be used to share accesses to any shared resource and is not limited to only the use of a shared interconnect. Such other shared resources may include ARL28, processing resources, memory resources (such as LUT30), or almost any other type of resource that is shared among multiple parties competing for access. Such other shared resources may include ARL28, processing resources, memory resources (such as LUT30), or almost any other type of resource that is shared among multiple parties competing for access. or almost any other type of resource that is shared among multiple parties competing for access.
[0086] Referring to FIG. 4, a flowchart 40 showing the operational steps for arbitrating access to a shared resource is shown. Referring to FIG. 4, a flowchart 40 showing the operational steps for arbitrating access to a shared resource is shown.
[0087] In step 42, various source subsystem agents 14 generate transactions. The transaction can be any of three classes including Posted (P), Non-posted (NP), and Completion (C). and Completion (C).
[0088] In step 44, the transactions generated by the source subsystem agents 14 Each cushion is packetized. As described above, the packetization of a given transaction can result in one or more packets. The packets can have various sizes, some packets having large payloads and other packets having small payloads or no payload at all. In the situation where a transaction is represented by a single packet having a data payload 36 that is smaller than the width of the interconnect 12, the transaction can be represented by a single portion. In the situation where a transaction is represented by multiple packets or a single packet having a data payload 36 that is larger than the access width of the shared resource, multiple portions are required to represent the transaction. In step 46, the packetized portions of the transactions generated by each of the subsystem agents 14 are input to the local switch 16 via the corresponding ports 20. Within the port 20, the packets of each transaction are assigned to virtual channels. As described above, the assignment can be random, hard-coded, or based on balanced utilization, security, urgency, etc. across all virtual channels. In step 48, the packetized portions of the transactions generated by each of the subsystem agents 14 are each stored in appropriate first-in first-out buffers by both transaction classes and the virtual channels (e.g., VC0, VC1, VC2, and VC3) assigned to them. As described above, the virtual channels
[0089]
[0090] A NER may be assigned by one of many different priority schemes, such as strict or absolute priority, round robin, weighted priority, least recently serviced. (least recently serviced). In the case where a given transaction has multiple parts, each part is stored within the same buffer. As a result, multiple parts of a given transaction are transmitted over the same virtual channel associated with interconnect 12. When a transaction part is injected, the corresponding counter for tracking the number of content items within each buffer is decremented. If a particular buffer becomes full, its counter is decremented to zero, which means the buffer can no longer accept further content.
[0091] In steps 50, 52, and 54, first, second, and third level arbitration is performed. As described above, the selection of port 20, virtual channel, and transaction class may be performed in any order.
[0092] Element 56 may be used to maintain the rules used for performing first, second, and third level arbitration. In each case, element 56 is used as necessary to resolve each of the arbitration levels. For example, element 56 may maintain PCI and / or device ordering rules. Element 56 may comprise rules for implementing several priority schemes (such as strict or absolute priority, weighted priority, round robin, etc.) and logic or intelligence for determining which to use in a given arbitration cycle.
[0093] In step 58, the winner of the arbitration is determined. In step 60, the winning portion is placed within a buffer used to access the shared resource, and the counter associated with the buffer is decremented.
[0094] In step 62, the buffer associated with the winning portion is incremented since the winning portion is no longer within the buffer.
[0095] In step 64, the winning portion accesses the shared resource. When the access is complete, the buffer for the shared resource is incremented.
[0096] Steps 42 - 64 are each continuously repeated during successive clock cycles. As different winning portions, each accesses the shared resource.
[0097] Interleaving - Example 1 Transactions can be transmitted via the interconnect 12 in one of several modes.
[0098] In one mode called the "header in-line" mode, the header 34 of the transaction packet 32 is always transmitted first, in separate parts or beats, before the payload 36. The header in-line mode may or may not waste bits available on the interconnect 12 depending on the relative sizes of the header 34 and / or payload 36 with respect to the number of data bits N of the interconnect 12. For example, for an interconnect 12 with a 512 - bit width (N = 512) and a 128 - bit header and an Consider a packet having a 256-bit payload and, in this scenario, 12 8-bit headers are transmitted in the first part or beat, and the remaining 384 bits of the interconnect 12 bandwidth are not utilized. In the second part or beat, the 256-bit payload is transmitted, and the remaining 256 bits of the interconnect 12 are not utilized. In this example, a significant portion of the interconnect bandwidth is not utilized during the two beats. On the other hand, if most of the transaction packets are larger than the interconnect, the degree of wasted bandwidth is reduced or eliminated. For example, for headers and / or payloads that are 384 or 512 bits, the amount of waste is significantly reduced (e.g., 384 bits)
[0099] or eliminated entirely (e.g., 512 bits). In another mode called "header on side-band", the header 34 of the packet is transmitted "on the side" of the data, which means that control bits M are utilized while the payload 36 of the packet 32 is being transmitted over the N data bits of the interconnect 12. In the header on side-band mode, the number of bits or size of the payload 36 of the packet 32 determines the number of beats required to transmit the packet over a given interconnect 12. For example, for packets 32 having payloads 36 of 64, 128, 256, or 512 bits, and an interconnect 12 having 128 data bits (N = 128), the packets require 1, 1, 2, and 4 beats, respectively. In each transmission of a beat, the header information is transmitted along with or "on the side" of the payload data over the N data bits of the interconnect 12 using
[0100] In yet another mode, the header 34 of packet 32 is transmitted in the same way as the payload, but there is no requirement that the header 34 and the payload 36 must be transmitted in separate parts or beats. If packet 32 has a 128-bit header 34 and a 128-bit payload 36, the total size is 256 bits (128 + 128). When the N data bits of the interconnect 12 are 64, 128, 256, and 512-bit wide, the 256-bit packet is transmitted in 4, 2, 1, and 1 beats respectively. In another example, packet 32 has a 128-bit header and a 256-bit payload 36, i.e., a total packet size of 384 bits (128 + 256). For the same interconnect 12 with N data bits of 64, 128, 256, or 512 width, the packets are transmitted in 6, 3, 2, or 1 beats respectively. This mode is always at least as efficient as the header in-line mode described above. In yet another mode, the header 34 of packet 32 is transmitted in the same way as the payload, but there is no requirement that the header 34 and the payload 36 must be transmitted in separate parts or beats. If packet 32 has a 128-bit header 34 and a 128-bit payload 36, the total size is 256 bits (128 + 128). When the N data bits of the interconnect 12 are 64, 128, 256, and 512-bit wide, the 256-bit packet is transmitted in 4, 2, 1, and 1 beats respectively. In another example, packet 32 has a 128-bit header and a 256-bit payload 36, i.e., a total packet size of 384 bits (128 + 256). For the same interconnect 12 with N data bits of 64, 128, 256, or 512 width, the packets are transmitted in 6, 3, 2, or 1 beats respectively. This mode is always at least as efficient as the header in-line mode described above. In yet another mode, the header 34 of packet 32 is transmitted in the same way as the payload, but there is no requirement that the header 34 and the payload 36 must be transmitted in separate parts or beats. If packet 32 has a 128-bit header 34 and a 128-bit payload 36, the total size is 256 bits (128 + 128). When the N data bits of the interconnect 12 are 64, 128, 256, and 512-bit wide, the 256-bit packet is transmitted in 4, 2, 1, and 1 beats respectively. In another example, packet 32 has a 128-bit header and a 256-bit payload 36, i.e., a total packet size of 384 bits (128 + 256). For the same interconnect 12 with N data bits of 64, 128, 256, or 512 width, the packets are transmitted in 6, 3, 2, or 1 beats respectively. This mode is always at least as efficient as the header in-line mode described above. In yet another mode, the header 34 of packet 32 is transmitted in the same way as the payload, but there is no requirement that the header 34 and the payload 36 must be transmitted in separate parts or beats. If packet 32 has a 128-bit header 34 and a 128-bit payload 36, the total size is 256 bits (128 + 128). When the N data bits of the interconnect 12 are 64, 128, 256, and 512-bit wide, the 256-bit packet is transmitted in 4, 2, 1, and 1 beats respectively. In another example, packet 32 has a 128-bit header and a 256-bit payload 36, i.e., a total packet size of 384 bits (128 + 256). For the same interconnect 12 with N data bits of 64, 128, 256, or 512 width, the packets are transmitted in 6, 3, 2, or 1 beats respectively. This mode is always at least as efficient as the header in-line mode described above. In yet another mode, the header 34 of packet 32 is transmitted in the same way as the payload, but there is no requirement that the header 34 and the payload 36 must be transmitted in separate parts or beats. If packet 32 has a 128-bit header 34 and a 128-bit payload 36, the total size is 256 bits (128 + 128). When the N data bits of the interconnect 12 are 64, 128, 256, and 512-bit wide, the 256-bit packet is transmitted in 4, 2, 1, and 1 beats respectively. In another example, packet 32 has a 128-bit header and a 256-bit payload 36, i.e., a total packet size of 384 bits (128 + 256). For the same interconnect 12 with N data bits of 64, 128, 256, or 512 width, the packets are transmitted in 6, 3, 2, or 1 beats respectively. This mode is always at least as efficient as the header in-line mode described above. In yet another mode, the header 34 of packet 32 is transmitted in the same way as the payload, but there is no requirement that the header 34 and the payload 36 must be transmitted in separate parts or beats. If packet 32 has a 128-bit header 34 and a 128-bit payload 36, the total size is 256 bits (128 + 128). When the N data bits of the interconnect 12 are 64, 128, 256, and 512-bit wide, the 256-bit packet is transmitted in 4, 2, 1, and 1 beats respectively. In another example, packet 32 has a 12-eight-bit header and a 256-bit payload 36, i.e., a total packet size of 384 bits (128 + 256). For the same interconnect 12 with N data bits of 64, 128, 256, or 512 width, the packets are transmitted in 6, 3, 2, or 1 beats respectively. This mode is always at least as efficient as the header in-line mode described above. In yet another mode, the header 34 of packet 32 is transmitted in the same way as the payload, but there is no requirement that the header 34 and the payload 36 must be transmitted in separate parts or beats. If packet 32 has a 128-bit header 34 and a 128-bit payload 36, the total size is 256 bits (128 + 128). When the N data bits of the interconnect 12 are 64, 128, 256, and 512-bit wide, the 256-bit packet is transmitted in 4, 2, 1, and 1 beats respectively. In another example, packet 32 has a 128-bit header and a 256-bit payload 36, i.e., a total packet size of 384 bits (128 + 256). For the same interconnect 12 with N data bits of 64, 128, 256, or 512 width, the packets are transmitted in 6, 3, 2, or 1 beats respectively. This mode is always at least as efficient as the header in-line mode described above. In yet another mode, the header 34 of packet 32 is transmitted in the same way as the payload, but there is no requirement that the header 34 and the payload 36 must be transmitted in separate parts or beats. If packet 32 has a 128-bit header 34 and a 128-bit payload 36, the total size is 256 bits (128 + 128). When the N data bits of the interconnect 12 are 64, 128, 256, and 512-bit wide, the 256-bit packet is transmitted in 4, 2, 1, and 1 beats respectively. In another example, packet 32 has a 128-bit header and a 256-bit payload 36, i.e., a total packet size of 384 bits (128 + 256). For the same interconnect 12 with N data bits of 64, 128, 256, or 512 width, the packets are transmitted in 6, 3, 2, or 1 beats respectively. This mode is always at least as efficient as the header in-line mode described above. In yet another mode, the header 34 of packet 32 is transmitted in the same way as the payload, but there is no requirement that the header 34 and the payload 36 must be transmitted in separate parts or beats. If packet 32 has a 128-bit header 34 and a 128-bit payload 36, the total size is 256 bits (128 + 128). When the N data bits of the interconnect 12 are 64, 128, 256, and 512-bit wide, the 256-bit packet is transmitted in 4, 2, 1, and 1 beats respectively. In another example, packet 32 has a 128-bit header and a 256-bit payload 36, i.e., a total packet size of 384 bits (128 + 256). For the same interconnect 12 with N data bits of 64, 128, 256, or 512 width, the packets are transmitted in 6, 3, 2, or 1 beats respectively. This mode is always at least as efficient as the header in-line mode described above. In yet another mode, the header 34 of packet 32 is transmitted in the same way as the payload, but there is no requirement that the header 34 and the payload 36 must be transmitted in separate parts or beats. If packet 32 has a 128-bit header 34 and a 128-bit payload 36, the total size is 256 bits (128 + 128). When the N data bits of the interconnect 12 are 64, 128, 256, and 512-bit wide, the 256-bit packet is transmitted in 4, 2, 1, and 1 beats respectively. In another example, packet 32 has a 128-bit header and a 256-bit payload 36, i.e., a total packet size of 384 bits (128 + 256). For the same interconnect 12 with N data bits of 64, 128, 256, or 512 width, the packets are transmitted in 6, 3, 2, or 1 beats respectively. This mode is always at least as efficient as the header in-line mode described above. This mode is always at least as efficient as the header in-line mode described above.
[0101] Referring to Figure 5, a first example of interleaving portions of different transactions on multiple virtual channels is illustrated. In this example, for simplicity, only two transactions are shown. The two transactions are competing for access to a shared interconnect 12 that is 128 data bits wide (N = 128) in this example. Details of the two transactions include the following. Referring to Figure 5, a first example of interleaving portions of different transactions on multiple virtual channels is illustrated. In this example, for simplicity, only two transactions are shown. The two transactions are competing for access to a shared interconnect 12 that is 128 data bits wide (N = 128) in this example. Details of the two transactions include the following. Referring to Figure 5, a first example of interleaving portions of different transactions on multiple virtual channels is illustrated. In this example, for simplicity, only two transactions are shown. The two transactions are competing for access to a shared interconnect 12 that is 128 data bits wide (N = 128) in this example. Details of the two transactions include the following. Referring to Figure 5, a first example of interleaving portions of different transactions on multiple virtual channels is illustrated. In this example, for simplicity, only two transactions are shown. The two transactions are competing for access to a shared interconnect 12 that is 128 data bits wide (N = 128) in this example. Details of the two transactions include the following. Referring to Figure 5, a first example of interleaving portions of different transactions on multiple virtual channels is illustrated. In this example, for simplicity, only two transactions are shown. The two transactions are competing for access to a shared interconnect 12 that is 128 data bits wide (N = 128) in this example. Details of the two transactions include the following. (1) Transaction 1 (T1): Generated at time T1 and assigned to virtual channel VC2. The size of T1 is 4 bits, and those bits are T1A, T1B, (1) Transaction 1 (T1): Generated at time T1 and assigned to virtual channel VC2. The size of T1 is 4 bits, and those bits are T1A, T1B, Specified as T1C and T1D. (2) Transaction 2 (T2): Generated at time T2 (after time T1) and assigned to virtual channel VC0. The size of T2 is a single part or beat. In this example, an absolute or strict priority is assigned to the VCO. Over multiple cycles, parts of the two transactions T1 and T2 are transmitted over the shared interconnect as follows, as shown in Figure 5.
[0102] In this example, an absolute or strict priority is assigned to the VCO. Over multiple cycles, parts of the two transactions T1 and T2 are transmitted over the shared interconnect as follows, as shown in Figure 5. Cycle 1: Since the beat T1A of T1 is the only available transaction, it is transmitted on VC2. Cycle 2: The beats T1B of T1 and the single part of T2 compete for access to interconnect 12. Since the VCO has strict priority, T2 wins automatically. Therefore, the beat of T2 is transmitted on VC0. Cycle 3: Since there is no competing transaction, the beat T1B of T1 is transmitted on VC2. Cycle 4: Since there is no competing transaction, the beat T1C of T1 is transmitted on VC2. Cycle 5: Since there is no competing transaction, the beat T1D of T1 is transmitted on VC2. This example shows the following: (1) In a virtual channel with absolute priority, access to the shared interconnect 12 is immediately granted whenever traffic becomes available, regardless of whether other traffic was waiting first, and (2) the winning parts or beats of different transactions are transmitted on different virtual channels associated with interconnect 12. Cycle 1: Since the beat T1A of T1 is the only available transaction, it is transmitted on VC2. Cycle 2: The beats T1B of T1 and the single part of T2 compete for access to interconnect 12. Since the VCO has strict priority, T2 wins automatically. Therefore, the beat of T2 is transmitted on VC0. Cycle 3: Since there is no competing transaction, the beat T1B of T1 is transmitted on VC2. Cycle 4: Since there is no competing transaction, the beat T1C of T1 is transmitted on VC2. Cycle 5: Since there is no competing transaction, the beat T1D of T1 is transmitted on VC2. Cycle 5: Since there is no competing transaction, the beat T1D of T1 is transmitted on VC2.
[0103] This example shows the following: (1) In a virtual channel with absolute priority, access to the shared interconnect 12 is immediately granted whenever traffic becomes available, regardless of whether other traffic was waiting first, and (2) the winning parts or beats of different transactions are transmitted on different virtual channels associated with interconnect 12. In a virtual channel with absolute priority, access to the shared interconnect 12 is immediately granted whenever traffic becomes available, regardless of whether other traffic was waiting first. And (2) the winning parts or beats of different transactions are transmitted on different virtual channels associated with interconnect 12. In a virtual channel with absolute priority, access to the shared interconnect 12 is immediately granted whenever traffic becomes available, regardless of whether other traffic was waiting first. to be interleaved and transmitted by the NEL. In this example, the virtual channel VCO is given an absolute priority. In an absolute or strict priority scheme, it should be understood that any of the virtual channels may be assigned the highest priority.
[0104] Interleaving - Example 2 Referring to FIG. 6, a second example of interleaving of portions of different transactions on multiple virtual channels is illustrated.
[0105] In this example, the priority scheme for access to the interconnect 12 is weighted, which means that the VCO is given access rights with a probability of (40%), and VC1 - VC3 are each given access rights with a probability of (20%). Also, the interconnect is 12 8-bit wide.
[0106] Furthermore, in this example, there are four competing transactions T1, T2, T3, and T4. - T1 is assigned to VC0 and includes four parts or beats T1A, T1B, T1C, and T1D. - T2 is assigned to VC1 and includes two parts or beats T2A and T2B. - T3 is assigned to VC2 and includes two parts or beats T3A and T3B. - T4 is assigned to VC3 and includes two parts or beats T4A and T4B.
[0107] In this example, the priority scheme is weighted. As a result, each virtual channel wins according to its weight ratio. In other words, during 10 cycles, VC0 wins 4 times, and VC1, VC2, and VC3 each win 2 times. For example, as shown in FIG. 6, - The four parts or beats of T1, T1A, T1B, T1C, and T1D are transmitted through the VCO in 4 cycles (40%) out of 10 cycles (i.e., cycles 1, 4, 7, and 10). - The two parts or beats of T2, T2A and T2B, are transmitted through VC1 in 2 cycles (20%) out of 10 cycles (i.e., cycles 2 and 6). - The two parts or beats of T3, T3A and T3B, are transmitted through VC2 in 2 cycles (20%) out of 10 cycles (i.e., cycles 5 and 9). - The two parts or beats of T4, T4A and T4B, are transmitted through VC3 in 2 cycles (20%) out of 10 cycles (i.e., cycles 3 and 8). - The two parts or beats of T2, T2A and T2B, are transmitted through VC1 in 2 cycles (20%) out of 10 cycles (i.e., cycles 2 and 6). - The two parts or beats of T3, T3A and T3B, are transmitted through VC2 in 2 cycles (20%) out of 10 cycles (i.e., cycles 5 and 9). - The two parts or beats of T4, T4A and T4B, are transmitted through VC3 in 2 cycles (20%) out of 10 cycles (i.e., cycles 3 and 8). - The two parts or beats of T4, T4A and T4B, are transmitted through VC3 in 2 cycles (20%) out of 10 cycles (i.e., cycles 3 and 8). - The two parts or beats of T4, T4A and T4B, are transmitted through VC3 in 2 cycles (20%) out of 10 cycles (i.e., cycles 3 and 8).
[0108] Thus, this example shows the following: (1) A weighted priority scheme in which each virtual channel is given access rights to the interconnect 12 based on a predetermined ratio, and (2) Another example in which the winning portions of different transactions are interleaved and transmitted on different virtual channels associated with the interconnect 12. In this example of weighting, it should be understood that there is sufficient traffic to allocate portions of the transaction to the various virtual channels according to the weighting ratio. On the other hand, if the amount of traffic is insufficient, the weighting ratio may or may not be strictly enforceable. For example, if there is a large amount of traffic on virtual channel VC3 and limited traffic on the other virtual channels VC0, VC1, and VC2 In this example of weighting, it should be understood that there is sufficient traffic to allocate portions of the transaction to the various virtual channels according to the weighting ratio. On the other hand, if the amount of traffic is insufficient, the weighting ratio may or may not be strictly enforceable. For example, if there is a large amount of traffic on virtual channel VC3 and limited traffic on the other virtual channels VC0, VC1, and VC2 In this example of weighting, it should be understood that there is sufficient traffic to allocate portions of the transaction to the various virtual channels according to the weighting ratio. On the other hand, if the amount of traffic is insufficient, the weighting ratio may or may not be strictly enforceable. For example, if there is a large amount of traffic on virtual channel VC3 and limited traffic on the other virtual channels VC0, VC1, and VC2
[0109] In this example of weighting, it should be understood that there is sufficient traffic to allocate portions of the transaction to the various virtual channels according to the weighting ratio. On the other hand, if the amount of traffic is insufficient, the weighting ratio may or may not be strictly enforceable. For example, if there is a large amount of traffic on virtual channel VC3 and limited traffic on the other virtual channels VC0, VC1, and VC2 In this example of weighting, it should be understood that there is sufficient traffic to allocate portions of the transaction to the various virtual channels according to the weighting ratio. On the other hand, if the amount of traffic is insufficient, the weighting ratio may or may not be strictly enforceable. For example, if there is a large amount of traffic on virtual channel VC3 and limited traffic on the other virtual channels VC0, VC1, and VC2 In this example of weighting, it should be understood that there is sufficient traffic to allocate portions of the transaction to the various virtual channels according to the weighting ratio. On the other hand, if the amount of traffic is insufficient, the weighting ratio may or may not be strictly enforceable. For example, if there is a large amount of traffic on virtual channel VC3 and limited traffic on the other virtual channels VC0, VC1, and VC2 In this example of weighting, it should be understood that there is sufficient traffic to allocate portions of the transaction to the various virtual channels according to the weighting ratio. On the other hand, if the amount of traffic is insufficient, the weighting ratio may or may not be strictly enforceable. For example, if there is a large amount of traffic on virtual channel VC3 and limited traffic on the other virtual channels VC0, VC1, and VC2 In this example of weighting, it should be understood that there is sufficient traffic to allocate portions of the transaction to the various virtual channels according to the weighting ratio. On the other hand, if the amount of traffic is insufficient, the weighting ratio may or may not be strictly enforceable. For example, if there is a large amount of traffic on virtual channel VC3 and limited traffic on the other virtual channels VC0, VC1, and VC2 If not at all, VC3 would carry all or most of the traffic if the weighting ratio were strictly enforced. However, as a result, not all clock cycles or beats can transmit parts of the transaction, so interconnect 12 cannot be fully utilized. On the other hand, if the weighting ratio is not strictly enforced, it is possible to reassign transaction traffic to increase the utilization of the interconnect (for example, the traffic is transmitted in a larger number of cycles or beats). However, as a result, not all clock cycles or beats can transmit parts of the transaction, so interconnect 12 cannot be fully utilized. On the other hand, if the weighting ratio is not strictly enforced, it is possible to reassign transaction traffic to increase the utilization of the interconnect (for example, the traffic is transmitted in a larger number of cycles or beats). However, as a result, not all clock cycles or beats can transmit parts of the transaction, so interconnect 12 cannot be fully utilized. On the other hand, if the weighting ratio is not strictly enforced, it is possible to reassign transaction traffic to increase the utilization of the interconnect (for example, the traffic is transmitted in a larger number of cycles or beats). However, as a result, not all clock cycles or beats can transmit parts of the transaction, so interconnect 12 cannot be fully utilized.
[0110] The two examples above are applicable regardless of which of the above-described transmission modes is utilized. When transactions are split into parts or beats, they can be interleaved and transmitted over shared interconnect 12 using any of the arbitration schemes defined herein. The two examples above are applicable regardless of which of the above-described transmission modes is utilized. When transactions are split into parts or beats, they can be interleaved and transmitted over shared interconnect 12 using any of the arbitration schemes defined herein. The two examples above are applicable regardless of which of the above-described transmission modes is utilized. When transactions are split into parts or beats, they can be interleaved and transmitted over shared interconnect 12 using any of the arbitration schemes defined herein. The two examples above are applicable regardless of which of the above-described transmission modes is utilized. When transactions are split into parts or beats, they can be interleaved and transmitted over shared interconnect 12 using any of the arbitration schemes defined herein.
[0111] The arbitration schemes described above are merely a few examples. In other examples, low jitter, weighting, strict, round robin, or almost any other arbitration scheme may be used. Therefore, the arbitration schemes listed or described herein are illustrative and should not be construed as limiting in any way. The arbitration schemes described above are merely a few examples. In other examples, low jitter, weighting, strict, round robin, or almost any other arbitration scheme may be used. Therefore, the arbitration schemes listed or described herein are illustrative and should not be construed as limiting in any way. The arbitration schemes described above are merely a few examples. In other examples, low jitter, weighting, strict, round robin, or almost any other arbitration scheme may be used. Therefore, the arbitration schemes listed or described herein are illustrative and should not be construed as limiting in any way. The arbitration schemes described above are merely a few examples. In other examples, low jitter, weighting, strict, round robin, or almost any other arbitration scheme may be used. Therefore, the arbitration schemes listed or described herein are illustrative and should not be construed as limiting in any way.
[0112] Multiple simultaneous arbitrations Up to this point, for simplicity, only a single arbitration has been described. However, it should be understood that in a realistic application (such as on an SoC), multiple arbitrations can occur simultaneously. Up to this point, for simplicity, only a single arbitration has been described. However, it should be understood that in a realistic application (such as on an SoC), multiple arbitrations can occur simultaneously. Up to this point, for simplicity, only a single arbitration has been described. However, it should be understood that in a realistic application (such as on an SoC), multiple arbitrations can occur simultaneously.
[0113] Referring to FIG. 7, traffic is processed in both directions between switches 16 and 18. A block diagram of two shared interconnects 12 and 12Z for this is shown. As described above The switch 16 directs transaction traffic from the source sub - function 14 (i.e., IP1, IP2, and IP3) to the destination sub - function 14 (i.e., IP4, IP5, and IP6) via the shared interconnect 12. To handle reverse - direction transaction traffic, the switch 18 includes an arbitration element 26Z and optionally an ARL 28Z. During operation, the elements 26Z and ARL 28Z operate complementarily to the above - described operation, which means that the transaction traffic generated by the source IP agent 14 (i.e., IP4, IP5, and IP6) is arbitrated and sent via the shared interconnect 12Z to the destination IP agent (i.e., IP1, IP2, and IP3). Alternatively, arbitration may be performed without ARL 28Z, which means that the arbitration simply makes a decision between the competing ports 20 (e.g., Port3, Port3 or Port5), and the portion of the transaction associated with the winning port is transmitted on the interconnect 12 regardless of the ultimate destination of that portion. Since the elements 12Z, 26Z, and 28Z have already been described, a detailed description is not provided here for simplicity.
[0114] The SoC may have multiple levels of sub - functions 14 and multiple shared interconnects 12. For each, using the above - described arbitration scheme, arbitration between transactions transmitted via the interconnect 12 between various sub - functions can be performed simultaneously.
[0115] Interconnect fabric Referring to FIG. 8, an example 100 of a SoC is shown. The SoC 100 includes a plurality of IP agents 14 (IP1, IP2, IP3,... IPN). Each IP agent 14 is connected to one of several nodes 102. Shared interconnects 12, 1 2Z are in the reverse direction and are provided between various nodes 102. In this configuration, trans actions can flow bidirectionally between each pair of nodes 102, for example, as described above with respect to FIG. 7.
[0116] In a non-exclusive embodiment, each node 102 includes various switches 16, 18, an access port 20 for connecting to the local IP agent 14, an access port 22 for connecting to the shared interconnects 12, 12Z, an arbitration element 26, an optional ARL 28, and an optional LUT 30. In an alternative embodiment, the node may not include the arbitration element 26 and / or the ARL 28. For nodes 102 without these elements, transactions without the necessary routing information can be transferred to the default node as described above. Since each of these elements has been described above with respect to FIG. 1, a detailed description is not provided here for simplicity. Collectively, the various nodes 102 and the bidirectional interconnects 12, 12Z define an interconnect fabric 106 for the SoC 100. The interconnect fabric 106 shown in the figure is relatively simple for simplicity. In an actual embodiment, on the SoC 100
[0117] is more complex. The interconnect fabric includes hundreds or thousands of IP agents 14 and multiple levels of nodes 102, all of which are interconnected by a large number of interconnects 12, 12Z, and it should be understood that this can be very complex.
[0118] Broadcast, multicast, and anycast In some applications, such as machine learning or artificial intelligence, transactions generated by one IP agent 14 are typically broadcast widely to multiple IP agents 14 on the SoC 100. Transactions broadcast widely to multiple IP agents 14 can be implemented by broadcast, multicast, or anycast. On a given SoC 100, broadcast, multicast, and / or anycast can each be implemented independently or together. A brief definition of each of these types of transactions is provided below. · Broadcast is a transaction that is sent to all IP agents on the SoC 100. For example, in the SoC 100 shown in FIG. 8, as a result of a broadcast sent by IP1, IP2 through IPN each receive the transaction. · Multicast is a transaction that is sent to two or more (potentially all) of the IP agents on the SoC. For example, if IP1 generates a multicast transaction that specifies IP5, IP7, and IP9, these agents 14 receive the transaction, but the remaining IP agents 14 on the SoC 100 do not. If the multicast is sent to all IP agents 14, it is basically is the same as a broadcast. · A read response multicast is a variation of the multicast transaction described above. In a read response multicast, a single IP agent 14 may read the content of the memory location. Rather than only the starting IP agent 14 believing the content, a number of destination IP agents 14 receive the content. The IP agents 14 that receive the read result may range from two or more IP agents 14 to all of the IP agents 14 on the SoC10 0. · An anycast is a transaction generated by an IP agent 14. However, the sending IP agent 14 does not specify the target IP agent 14 at all. Instead, the interconnect fabric 106 (i.e., one or more within the node 102) determines the receiving IP agent 14. For example, when IP1 generates an anycast transaction, one or more nodes within the node 102 determine which of the other agent IPs 2~IPN will receive the transaction. In various embodiments of the anycast transaction, one, multiple, or all of the IP agents 14 on the SoC may receive the anycast transaction. · An anycast is a transaction generated by an IP agent 14. However, the sending IP agent 14 does not specify the target IP agent 14 at all. Instead, the interconnect fabric 106 (i.e., one or more within the node 102) determines the receiving IP agent 14. For example, when IP1 generates an anycast transaction, one or more nodes within the node 102 determine which of the other agent IPs 2~IPN will receive the transaction. In various embodiments of the anycast transaction, one, multiple, or all of the IP agents 14 on the SoC may receive the anycast transaction. However, the sending IP agent 14 does not specify the target IP agent 14 at all. Instead, the interconnect fabric 106 (i.e., one or more within the node 102) determines the receiving IP agent 14. For example, when IP1 generates an anycast transaction, one or more nodes within the node 102 determine which of the other agent IPs 2~IPN will receive the transaction. In various embodiments of the anycast transaction, one, multiple, or all of the IP agents 14 on the SoC may receive the anycast transaction. However, the sending IP agent 14 does not specify the target IP agent 14 at all. Instead, the interconnect fabric 106 (i.e., one or more within the node 102) determines the receiving IP agent 14. For example, when IP1 generates an anycast transaction, one or more nodes within the node 102 determine which of the other agent IPs 2~IPN will receive the transaction. In various embodiments of the anycast transaction, one, multiple, or all of the IP agents 14 on the SoC may receive the anycast transaction. However, the sending IP agent 14 does not specify the target IP agent 14 at all. Instead, the interconnect fabric 106 (i.e., one or more within the node 102) determines the receiving IP agent 14. For example, when IP1 generates an anycast transaction, one or more nodes within the node 102 determine which of the other agent IPs 2~IPN will receive the transaction. In various embodiments of the anycast transaction, one, multiple, or all of the IP agents 14 on the SoC may receive the anycast transaction. However, the sending IP agent 14 does not specify the target IP agent 14 at all. Instead, the interconnect fabric 106 (i.e., one or more within the node 102) determines the receiving IP agent 14. For example, when IP1 generates an anycast transaction, one or more nodes within the node 102 determine which of the other agent IPs 2~IPN will receive the transaction. In various embodiments of the anycast transaction, one, multiple, or all of the IP agents 14 on the SoC may receive the anycast transaction. However, the sending IP agent 14 does not specify the target IP agent 14 at all. Instead, the interconnect fabric 106 (i.e., one or more within the node 102) determines the receiving IP agent 14. For example, when IP1 generates an anycast transaction, one or more nodes within the node 102 determine which of the other agent IPs 2~IPN will receive the transaction. In various embodiments of the anycast transaction, one, multiple, or all of the IP agents 14 on the SoC may receive the anycast transaction. However, the sending IP agent 14 does not specify the target IP agent 14 at all. Instead, the interconnect fabric 106 (i.e., one or more within the node 102) determines the receiving IP agent 14. For example, when IP1 generates an anycast transaction, one or more nodes within the node 102 determine which of the other agent IPs 2~IPN will receive the transaction. In various embodiments of the anycast transaction, one, multiple, or all of the IP agents 14 on the SoC may receive the anycast transaction.
[0119] A given transaction may be initiated as a broadcast, a multicast (including a read response multicast), or an anycast in several ways. For simplicity, hereinafter, these transactions will be collectively referred to as "BMA" transactions, which are broadcasts, multicasts (read response multicasts ), or anycasts. For simplicity, hereinafter, these transactions will be collectively referred to as "BMA" transactions, which are broadcasts, multicasts (read response multicasts ), or anycasts. (including To) or means an anycast transaction.
[0120] In one embodiment, the IP agent 14 may initiate a BMA transaction using a coded command inserted into the command field CMD of the header 34 of the packet 32 representing the transaction. The coded command causes the interconnect fabric 106 of the SoC10 to recognize or understand that the transaction is a BMA transaction rather than a normal transaction that specifies one source and one destination IP agent 14. For example, a unique combination of bits may define a given transaction as any of a broadcast, multicast, read response multicast, or anycast, respectively. 0 to recognize or understand that the transaction is a BMA transaction rather than a normal transaction that specifies one source and one destination IP agent 14. For example, a unique combination of bits may define a given transaction as any of a broadcast, multicast, read response multicast, or anycast, respectively. to recognize or understand that the transaction is a BMA transaction rather than a normal transaction that specifies one source and one destination IP agent 14. For example, a unique combination of bits may define a given transaction as any of a broadcast, multicast, read response multicast, or anycast, respectively. to recognize or understand that the transaction is a BMA transaction rather than a normal transaction that specifies one source and one destination IP agent 14. For example, a unique combination of bits may define a given transaction as any of a broadcast, multicast, read response multicast, or anycast, respectively. to recognize or understand that the transaction is a BMA transaction rather than a normal transaction that specifies one source and one destination IP agent 14. For example, a unique combination of bits may define a given transaction as any of a broadcast, multicast, read response multicast, or anycast, respectively. to recognize or understand that the transaction is a BMA transaction rather than a normal transaction that specifies one source and one destination IP agent 14. For example, a unique combination of bits may define a given transaction as any of a broadcast, multicast, read response multicast, or anycast, respectively.
[0121] In another embodiment, the BMA transaction may be implemented by issuing a read or write transaction having a BMA address defined in the ADDR field of the header 34 of the packet 32 representing the transaction. The BMA address is specified within the system of the SoC100 as indicating one of a broadcast, multicast, or anycast transaction. As a result, the BMA address is recognized by the interconnect fabric 106 and the transaction is treated as a broadcast, multicast, or anycast. address is recognized by the interconnect fabric 106 and the transaction is treated as a broadcast, multicast, or anycast. address is recognized by the interconnect fabric 106 and the transaction is treated as a broadcast, multicast, or anycast. address is recognized by the interconnect fabric 106 and the transaction is treated as a broadcast, multicast, or anycast. address is recognized by the interconnect fabric 106 and the transaction is treated as a broadcast, multicast, or anycast.
[0122] In yet another embodiment, both the command and the BMA address are used to specify a broadcast, multicast, or anycast transaction. to specify a broadcast, multicast, or anycast transaction. It can be used.
[0123] Anycast typically involves a source IP agent 14 wanting to send a transaction to multiple destinations, but not being aware of elements that can help select one or more preferred or ideal destination IP agents 14. For example, a source IP agent 14 may attempt to send a transaction to multiple IP agents each implementing an accelerator function. Designating the transaction as anycast causes one or more of the nodes 102 to participate in the selection of the destination IP agent 14. In various embodiments, the selection criteria can vary widely and can be based on congestion (busy IP agent versus idle state i.e., non-busy IP agent), random selection function, hardwired logic function, hash function, longest time unused function, power consumption considerations, or any other decision function or criterion. Thus, the responsibility for selecting the destination IP agent 14 is shifted to the node 102, which may have more information to make a better routing decision than the source agent 14. but not being aware of elements that can help select one or more preferred or ideal destination IP agents 14. For example, a source IP agent 14 may attempt to send a transaction to multiple IP agents each implementing an accelerator function. Designating the transaction as anycast causes one or more of the nodes 102 to participate in the selection of the destination IP agent 14. In various embodiments, the selection criteria can vary widely and can be based on congestion (busy IP agent versus idle state i.e., non-busy IP agent), random selection function, hardwired logic function, hash function, longest time unused function, power consumption considerations, or any other decision function or criterion. Thus, the responsibility for selecting the destination IP agent 14 is shifted to the node 102, which may have more information to make a better routing decision than the source agent 14. In various embodiments, the selection criteria can vary widely and can be based on congestion (busy IP agent versus idle state i.e., non-busy IP agent), random selection function, hardwired logic function, hash function, longest time unused function, power consumption considerations, or any other decision function or criterion. Thus, the responsibility for selecting the destination IP agent 14 is shifted to the node 102, which may have more information to make a better routing decision than the source agent 14. random selection function, hardwired logic function, hash function, longest time unused function, power consumption considerations, or any other decision function or criterion. Thus, the responsibility for selecting the destination IP agent 14 is shifted to the node 102, which may have more information to make a better routing decision than the source agent 14. Thus, the responsibility for selecting the destination IP agent 14 is shifted to the node 102, which may have more information to make a better routing decision than the source agent 14. Thus, the responsibility for selecting the destination IP agent 14 is shifted to the node 102, which may have more information to make a better routing decision than the source agent 14. It can be used.
[0124] Referring to FIG. 9A, FIG. 90 showing the logic 90 of node 102 for supporting BMA addressing is shown. The logic 90 includes a LUT 30, an Interconnect Fabric ID (IFID) table 124, and an optional physical link selector 126. The optional physical link selector 126 is used when there are two (or more) overlapping physical resources sharing a single logical identifier (such as in a trunking situation), which will be described in detail later. Referring to FIG. 9A, FIG. 90 showing the logic 90 of node 102 for supporting BMA addressing is shown. The logic 90 includes a LUT 30, an Interconnect Fabric ID (IFID) table 124, and an optional physical link selector 126. The optional physical link selector 126 is used when there are two (or more) overlapping physical resources sharing a single logical identifier (such as in a trunking situation), which will be described in detail later. The optional physical link selector 126 is used when there are two (or more) overlapping physical resources sharing a single logical identifier (such as in a trunking situation), which will be described in detail later. It can be used.
[0125] The IFID table includes, for each IP agent 14, (a) a corresponding logical IP ID for logically identifying each IP agent 14 within the SoC 100, and (b) either port 20 if the corresponding IP agent 14 is local to node 102 or (c an access port 22 to an appropriate interconnect 12 or 12Z that connects to the next node 102 along the delivery path to the corresponding IP agent 14. With this configuration, each node can access the identity of the physical ports 20 and / or 22 necessary to deliver transactions to each IP agent 14 within the SoC 100. That is, each node can deliver transactions to each IP agent 14 within the SoC 100.
[0126] The IFID table 124 for each node 102 within the fabric 106 is relative ( i.e., unique). In other words, each IFID table 124 includes only a list of the ports 20 and / or 22 necessary to deliver transactions to (1) its local IP agent 14 or (2) another node 102 along the delivery path to another IP agent 14 within the SoC that is not locally connected to node 120 via one of the shared interconnects 12, 12Z. With this configuration, each node 102 either (1) delivers transactions to its local IP agent 14 specified as the destination or (2) forwards transactions via the interconnects 12, 12Z to another node 102. At the next node, the above process is repeated. By delivering or forwarding transactions at each node 102, ultimately, a given transaction reaches its destination. By delivering or forwarding transactions locally at each node 102, ultimately, a given The transaction is delivered to all of the designated destination IP agents 14 within the interconnect fabric 106 for the SoC 100.
[0127] The LUT 30 is a first portion 120 used for routing conventional transactions (i.e., transactions sent to a single destination IP agent 14). When a conventional transaction is generated, the source IP agent 14 specifies a destination address in the system memory 24 within the ADDR field of the packet header representing the transaction. The transaction is then provided to the local node 102 for routing. In response, the ARL 28 accesses the first portion 120 of the LUT 30 to find the logical IP ID corresponding to the destination address. The IFID table 124 is then accessed to specify either (a) port 20 if the destination IP agent 14 is local to node 102 or (b) the access port 22 to the appropriate interconnect 12 or 12Z leading to the next node 102 along the delivery path to the destination IP agent 14. The IP ID is placed in the DST field of the header 34 of the packet 32 before being sent along the
[0128] appropriate port 20 or 22. For broadcast, multicast, or anycast transactions, the second portion 122 of the LUT 30 includes a plurality of BMA addresses (e.g., BMA 1 to BMA N, where N is any number that can be optionally or (1) One or more unique IP IDs (e.g., IP4 and IP7 for BMA address 1 , IP5, IP12, and IP24 for BMA address 2). (2) A unique code (e.g., code 1 and code 2 for BMA addresses 10 and 11). (3) A bit vector (e.g., the bit vector for BMA addresses 20 and 21 ).
[0129] Each code uniquely identifies a different set of destination IP agents. For example, the first code can be used to specify the first set of destination IP agents (e.g., IP1, IP13, and IP21), and the second code can be used to specify another set of destination agents (e.g., IP4, IP9, and IP17).
[0130] In a bit vector, each bit position corresponds to IP agent 14 on SoC100. Depending on whether a given bit position is set or reset, the corresponding IP agent 14 is designated as a destination or not, respectively. As an example, a bit vector of (101011...1) indicates that the corresponding IP agents 14 (IP1, IP3, IP5, IP6, and IPN) are set and the rest are reset. For example, a bit vector of (101011...1) indicates that the corresponding IP agents 14 (IP1, IP3, IP5, IP6, and IPN) are set and the rest are reset.
[0131] In each of the above embodiments, one or more logical IP IDs are identified as the destination IP agents for a given transaction. The IFID table 124 is used to convert the logical identifier IP ID values into the physical access ports 20 and / or 22 necessary to route the transaction to those destinations. In the case of BMA addresses, correctly To determine that physical access ports 20 and / or 22 are required, a deliberate code or bit vector may be used instead of the IP ID value.
[0132] Both the code and the bit vector can be used to specify a number of destination IP agents 14. The bit vector may perhaps be limited by the width of the destination field DST in the header 34 of the packet 32 representing the transaction. For example, if the destination field DST is 32, 64, 128, or 258 bits wide, the maximum number of IP agents 14 is limited to 32, 64, 128, and 256 respectively. If the number of IP agents 14 on a given SoC happens to exceed the number of possible IP agents that can be specified by the width of the destination field DST, other fields in the header 34 may be used in some cases, or the DST field may be extended. However, in a very complex SoC 100, the number of IP agents 14 may exceed the number of available bits
[0133] that can actually be used within the bit vector. Since any number of destination IP agents may be specified in the code, this problem is avoided.
[0134] Source-based routing (SBR) Source-based routing (SBR) differs from conventional routing in the following respects.(1) The source IP agent 14 has some knowledge or instruction to give to the interconnect fabric 106 at the time of issuing a transaction. For example, the source IP agent 14 knows the IP ID of the destination IP agent 14 to which it wants to send a transaction. (2) The source IP agent 14 has no interest in and / or does not know the address in the system memory 24 that is normally provided in the ADDR field of the packet header 34 of the transaction packet 32. (3) The node 102 in the interconnect fabric 106 knows to do something different from converting the address in the ADDR field of the packet header 34 of the packet into a single IP ID for a single destination IP agent.
[0135] Both broadcast and multicast are examples of transactions that may but are not necessarily SBR transactions. When the source issues a broadcast or multicast in the header 34 of the transaction packet 32 by specifying either (a) a broadcast and / or multicast code and (b) the destination IP agent 14, the transaction is considered source-based because the source IP agent specifies the destination IP agent. On the other hand, when the source starts a broadcast or multicast transaction using a BMA address without any specific knowledge of the destination, the transaction is considered non-source-based. An anycast transaction is not considered source-based because it does not specify the destination IP agent 14.
[0136] Hashing In hashing, a hash function is used to define a destination or a route to the destination. In some embodiments, the hash function may carefully define multiple destinations and / or multiple routes to multiple destinations.
[0137] Referring to FIG. 9B, FIG. 140 is shown that illustrates the use of a hash function to perform routing decisions. In this embodiment, a hash value 142 is provided within any number of fields of a header 34 of a packet 32 that represents a transaction. For example, a subset of any possible combination of address bits, commands, source agent IP, or information or data included in the header 34 may be used to define the hash value. Somewhere within the corresponding local node 102 or on the SoC 100, a hash function 144 is applied to the hash value 142. In response to the hash function 144, routing decisions may be made. For example, one or more IP IDs of the destination agent 14 may be defined. By providing different hash values, different routing decisions may be defined. It should be understood that hashing may be used for many other purposes within the SoC. One such use is the utilization of a hash function for tranking. In tranking, there are two (or more) duplicate physical resources that share a single logical identifier. As will be described in detail later, the hash function may be used to select from among the duplicate physical resources.
[0138] Optimization of transaction traffic Some applications of the SoC, such as machine learning, artificial intelligence, and data centers, may be transaction intensive. These types of applications tend to rely on broadcasting, multicasting, and anycasting, which can further increase transaction traffic. Broadcast transactions can significantly increase the amount of traffic sent through the interconnect fabric 106. To reduce bottlenecks, many procedures for reducing transaction traffic have been proposed. Such procedures include: (1) expanding transactions and integrating responses at nodes 102 of the interconnect fabric 106 of the SoC 100; (2) interleaving two or more transactions within a stream in a team defined by a combination of paired virtual channel transaction classes; (3) and "trunking" two or more physical links between IP agents sharing a common logical link or two or more identical IP agents sharing a common logical address. Broadcast, multicast, read response multicast, and anycast can each significantly increase the amount of transaction traffic among IP agents 14 on the SoC 100.
[0139] Broadcast transactions can significantly increase the amount of traffic sent through the interconnect fabric 106. To reduce bottlenecks, many procedures for reducing transaction traffic have been proposed. Such procedures include: (1) expanding transactions and integrating responses at nodes 102 of the interconnect fabric 106 of the SoC 100; (2) interleaving two or more transactions within a stream in a team defined by a combination of paired virtual channel transaction classes; (3) and "trunking" two or more physical links between IP agents sharing a common logical link or two or more identical IP agents sharing a common logical address.
[0140] To reduce bottlenecks, many procedures for reducing transaction traffic have been proposed. Such procedures include: (1) expanding transactions and integrating responses at nodes 102 of the interconnect fabric 106 of the SoC 100; (2) interleaving two or more transactions within a stream in a team defined by a combination of paired virtual channel transaction classes; (3) and "trunking" two or more physical links between IP agents sharing a common logical link or two or more identical IP agents sharing a common logical address. To reduce bottlenecks, many procedures for reducing transaction traffic have been proposed. Such procedures include: (1) expanding transactions and integrating responses at nodes 102 of the interconnect fabric 106 of the SoC 100; (2) interleaving two or more transactions within a stream in a team defined by a combination of paired virtual channel transaction classes; (3) and "trunking" two or more physical links between IP agents sharing a common logical link or two or more identical IP agents sharing a common logical address. To reduce bottlenecks, many procedures for reducing transaction traffic have been proposed. Such procedures include: (1) expanding transactions and integrating responses at nodes 102 of the interconnect fabric 106 of the SoC 100; (2) interleaving two or more transactions within a stream in a team defined by a combination of paired virtual channel transaction classes; (3) and "trunking" two or more physical links between IP agents sharing a common logical link or two or more identical IP agents sharing a common logical address. To reduce bottlenecks, many procedures for reducing transaction traffic have been proposed. Such procedures include: (1) expanding transactions and integrating responses at nodes 102 of the interconnect fabric 106 of the SoC 100; (2) interleaving two or more transactions within a stream in a team defined by a combination of paired virtual channel transaction classes; (3) and "trunking" two or more physical links between IP agents sharing a common logical link or two or more identical IP agents sharing a common logical address. To reduce bottlenecks, many procedures for reducing transaction traffic have been proposed. Such procedures include: (1) expanding transactions and integrating responses at nodes 102 of the interconnect fabric 106 of the SoC 100; (2) interleaving two or more transactions within a stream in a team defined by a combination of paired virtual channel transaction classes; (3) and "trunking" two or more physical links between IP agents sharing a common logical link or two or more identical IP agents sharing a common logical address. To reduce bottlenecks, many procedures for reducing transaction traffic have been proposed. Such procedures include: (1) expanding transactions and integrating responses at nodes 102 of the interconnect fabric 106 of the SoC 100; (2) interleaving two or more transactions within a stream in a team defined by a combination of paired virtual channel transaction classes; (3) and "trunking" two or more physical links between IP agents sharing a common logical link or two or more identical IP agents sharing a common logical address. To reduce bottlenecks, many procedures for reducing transaction traffic have been proposed. Such procedures include: (1) expanding transactions and integrating responses at nodes 102 of the interconnect fabric 106 of the SoC 100; (2) interleaving two or more transactions within a stream in a team defined by a combination of paired virtual channel transaction classes; (3) and "trunking" two or more physical links between IP agents sharing a common logical link or two or more identical IP agents sharing a common logical address. To reduce bottlenecks, many procedures for reducing transaction traffic have been proposed. Such procedures include: (1) expanding transactions and integrating responses at nodes 102 of the interconnect fabric 106 of the SoC 100; (2) interleaving two or more transactions within a stream in a team defined by a combination of paired virtual channel transaction classes; (3) and "trunking" two or more physical links between IP agents sharing a common logical link or two or more identical IP agents sharing a common logical address.
[0141] Expansion of transactions and integration of responses Broadcast, multicast, read response multicast, and anycast can each significantly increase the amount of transaction traffic among IP agents 14 on the SoC 100. Broadcast, multicast, read response multicast, and anycast can each significantly increase the amount of transaction traffic among IP agents 14 on the SoC 100. Broadcast, multicast, read response multicast, and anycast can each significantly increase the amount of transaction traffic among IP agents 14 on the SoC 100.
[0142] The SoC 100 has 25 IP agents, and one of the IP agents is a broadcast key. When generating a cast transaction, up to 24 individual transactions are typically sent via the interconnect fabric to other IP agents 14. A Non-posted (NP) transaction requests a response in the form of a Completion (C) transaction. If the transaction broadcast to 24 IP agents is Non-posted, another 24 Completion (C) transactions are similarly generated. As shown in this simple example, broadcasts can rapidly increase the amount of traffic transmitted on the interconnect fabric 106. Non-posted(NP) transactions request a response in the form of a Completion(C) transaction. Transactions broadcast to 24 IP agents. If the transaction is Non-posted, another 24 Completion( C) transactions are similarly generated. As shown in this simple example, broadcasts can rapidly increase the amount of traffic transmitted on the interconnect fabric 106. Multicast and anycast transactions can also rapidly increase the amount of traffic. In each of these transaction types, multiple receivers may be specified, which means that multiple transactions are sent and, presumably, multiple completion response transactions are received via the interconnect fabric 106. In a read response multicast transaction as well, the read content can be sent to multiple destination IP agents 14. As a result, the transaction volume can increase significantly with these types of transactions. 。
[0143] Multicast and anycast transactions can also rapidly increase the amount of traffic. In each of these transaction types, multiple receivers may be specified, which means that multiple transactions are sent and, presumably, multiple completion response transactions are received via the interconnect fabric 106. In a read response multicast transaction as well, the read content can be sent to multiple destination IP agents 14. As a result, the transaction volume can increase significantly with these types of transactions. In each of these transaction types, multiple receivers may be specified, which means that multiple transactions are sent and, presumably, multiple completion response transactions are received via the interconnect fabric 106. In a read response multicast transaction as well, the read content can be sent to multiple destination IP agents 14. As a result, the transaction volume can increase significantly with these types of transactions. In a read response multicast transaction as well, the read content can be sent to multiple destination IP agents 14. As a result, the transaction volume can increase significantly with these types of transactions. As a result, the transaction volume can increase significantly with these types of transactions. transactions.
[0144] To operate the interconnect fabric 106 more effectively, techniques are used to expand and integrate transactions at node 102 to reduce the amount of traffic. Referring to FIGS. 10A and 10B, a diagram illustrating an example of a SoC is shown.
[0145] Referring to FIGS. 10A and 10B, a diagram illustrating an example of a SoC is shown. In the example, the SoC includes five interconnected nodes 102A - 102E and an interconnect fabric 106 with ten IP agents 14 (IP1 - IP10).
[0146] Referring to FIG. 10A, IP1 broadcasts a Non - posted write transaction to the other IP agents IP2 - IP10. Using expansion, only a single transaction is sent on each shared interconnect 12. At each downstream node 102B - 102E, the node (1) provides the transaction to any local IP agent 14 and (2) forwards the transaction to any upstream node 102. Thus, in this example: · Node 102B provides the transaction to IP2 and forwards a single instantiation of the transaction to nodes 102C and 102D respectively. · At node 102C, the transaction is provided to local agents IP3, IP4, and IP5. · Node 102D provides the transaction to IP7. Further, node 102D also forwards a single instantiation of the transaction to node 102E. · At node 102E, the transaction is provided to IP8, IP9, and IP10.
[0147] In the above example, only a single transaction is sent on each shared interconnect 12 regardless of the number of IP agents 1 downstream of the transmitting - side IP agent 14.
[0148] Referring to FIG. 10B, the integration of response transactions will be described. The broadcast The transaction was a Non-posted write, so each destination agent IP2 to IP10 need to return the completed transaction. In the integration, each node 1 02B to 102E integrates the completed transaction received from its local IP agent 14, and then sends a single completed transaction upstream towards node 102A. In other words, · Node 102E integrates the completed transactions received from IP8 to IP10 and returns a single completed transaction to node 102D. · At node 102D, the completed transactions received from IP7 and node 102E are integrated, and one completed transaction is returned to node 102B. · Similarly, node 102C returns a single integrated transaction for IP3, IP4, and IP5. · Finally, node 102B integrates the completed transactions received from nodes 102C, 102D, and IP2 and returns a single completed transaction to nodes 102A and IP1. · The above example illustrates the efficiency of expansion and integration. Without expansion, nine separate transactions (one for each of agents IP2 to IP10) would need to be transmitted via the interconnection fabric 106. However, by using expansion, the number of transactions transmitted
[0149] via the various shared interconnections 12 is reduced to four. A total of nine completed transactions are also integrated into four. Occasionally, errors can occur, and the completed transaction may be incorrect at the receiving side IP agent
[0150] from time to time, and the completed transaction may be incorrect at the receiving It is not generated by one or more of IP2 to IP10. Errors can be handled in many different ways. For example, only successful completions can be integrated, while error responses can be combined and / or sent separately. In yet another alternative, both successful and error completions may be integrated, but each is flagged to indicate either a success response or a failure response. Although the above description is provided in the context of broadcast, it should be understood that the expansion and integration of transactions may be implemented in multicasting, read-response multicasting, and / or unicasting. In transaction-intensive application examples such as machine learning, artificial intelligence, and data centers where broadcast, multicast, and unicast transactions are common, if expansion and integration can be achieved, the amount of transaction traffic on the interconnect fabric 106 can be significantly reduced, eliminating or reducing bottlenecks and improving system efficiency and performance. The interconnect fabric 106 of the SoC 100 typically has a single physical link in each direction between (a) the IP agent 14 and the local node 102 and (b) between multiple nodes 102. If there is only a single link, there is a one-to-one correspondence between the physical link and the access port 20 or 22 for that physical link.
[0151]
[0152]
[0153] Trunking
[0153] Trunking The interconnect fabric 106 of the SoC 100 typically has a single physical link in each direction between (a) the IP agent 14 and the local node 102 and (b) between multiple nodes 102. If there is only a single link, there is a one-to-one correspondence between the physical link and the access port 20 or 22 for that physical link. between the IP agent 14 and the local node 102, and (b) between multiple nodes 102. If there is only a single link, there is a one-to-one correspondence between the physical link and the access port 20 or 22 for that physical link. If there is only a single link, there is a one-to-one correspondence between the physical link and the access port 20 or 22 for that physical link. between the physical link and the access port 20 or 22 for that physical link. Similarly, in most interconnect fabrics 106, there is also a one-to-one correspondence between the physical IP agent 14 and the logical IP ID used to access that IP agent 14. Similarly, in most interconnect fabrics 106, there is also a one-to-one correspondence between the physical IP agent 14 and the logical IP ID used to access that IP agent 14. a one-to-one correspondence.
[0154] In high-performance applications, it may be beneficial to use a technique called trancing. In trancing, there are two (or more) duplicate physical resources that share a single logical identifier. By duplicating physical resources, bottlenecks can be avoided, and the efficiency and performance of the system can be improved. For example, if one physical resource is busy, powered off, or unavailable, one of the duplicate resources can be used. For example, if one physical resource is busy, powered off, or unavailable, one of the duplicate resources can be used. Trancing can also improve reliability. If one physical resource (e.g., an interconnect or an IP agent) fails, becomes unavailable, or cannot be used for some reason, other physical resources can be used. If one physical resource (e.g., an interconnect or an IP agent) fails, becomes unavailable, or cannot be used for some reason, other physical resources can be used. By addressing duplicate resources using the same logical identifier, the advantages of duplicate physical resources can be realized without changing the logical addressing system used on the SoC100. By addressing duplicate resources using the same logical identifier, the advantages of duplicate physical resources can be realized without changing the logical addressing system used on the SoC100. However, it is a challenge to select and track which of the duplicate physical resources to use. However, it is a challenge to select and track which of the duplicate physical resources to use.
[0155] Referring to FIG. 11A, an interconnect fabric 106 of the SoC100 including some examples of trancing is shown. In this example, the interconnect fabric 106 includes three nodes 102A, 102B, and 102C. Node 102A includes two IP agents 14 iand 142. Node 102B includes two IP agents 143 and 144. Node 102C includes one IP agent 14s . The interconnect fabric 106 includes the following trancing examples. · A pair of physical "trunk" lines between node 102A and IP agent 142 . · Physical "trunk" interconnects 12 (1) and 12 (2) in the same direction from node 102A to 102C. · Pairs of the same IP agent 14s.
[0156] In each of these examples, there is no one-to-one correspondence between the logical identifier and the physical resource. Conversely, since there are two available physical resources, it is necessary to select which physical resource to use.
[0157] Referring to FIG. 11B, a diagram showing an optional physical link selector 126 (of FIG. 9A) is shown. As described above, in trancing, there are two ( or more) duplicate physical resources that share a single logical identifier. The IFID table 124, whenever it identifies a logical IP ID having duplicate physical resources in any trancing situation, uses an optional physical link selector 126 to make the selection. The physical link selector 126 may make the selection using one or more decision factors such as the availability (or lack thereof) of the physical resource, congestion, load balancing, hash functions, random selection, selection of the longest unused time, power conditions, etc. For example, if a certain physical resource is busy, congested, and / or unavailable, another physical resource is selected. Alternatively, and / or unavailable, another physical resource is selected. Alternatively, If a resource is powered off to reduce power consumption, another resource may be selected. Regardless of how it is done, the selection leads to the identification of physical port 20 or 22 used to access the selected physical resource. In a non-exclusive embodiment, it is preferable that the selection of the physical resource be utilized until the operation is completed. A series of related transactions are sent between the source IP agent 14 and a duplicate pair of destination IP agents (e.g., the two IP agents IP5 in FIG. 11A). All transactions are sent to the same destination IP agent until the operation is completed. Otherwise, data corruption or other problems may occur. When the duplicate physical resources are interconnected, typically, a similar approach is preferable.
[0158] In a transaction that requests a response (such as a read), it is preferable that both the read request transaction and the result response be sent through the same interconnection. Furthermore, to avoid packet corruption of the transaction, it is preferable that the entire transaction and the packets of the transaction be routed to the same destination along the same path. Packets require several bits to pass through a link and may possibly interleave with other virtual channels, so it is important not to change the port or link until the end of the packet. Otherwise, the multiple bits of the packet may become out of order as the parts of the packet move through the system, thereby corrupting the information. Although it is usually desirable to route the response through the same path as the request, it is not essential. (such as a read), it is preferable that both the read request transaction and the result response be sent through the same interconnection. Furthermore, to avoid packet corruption of the transaction, it is preferable that the entire transaction and the packets of the transaction be routed to the same destination along the same path. Packets require several bits to pass through a link and may possibly interleave with other virtual channels, so it is important not to change the port or link until the end of the packet. Otherwise, the multiple bits of the packet may become out of order as the parts of the packet move through the system, thereby corrupting the information. Although it is usually desirable to route the response through the same path as the request, it is not essential. Otherwise, the multiple bits of the packet may become out of order as the parts of the packet move through the system, thereby corrupting the information. Although it is usually desirable to route the response through the same path as the request, it is not essential.
[0159] In some non - exclusive embodiments, the ordering information transmitted with each beat is used to re - order the beats of a transaction received out of order or from different resources. It may be advantageous to provide this functionality to the destination IP agent. For example, control bit M may be used to identify a unique "beat count number" for each beat of a packet. Then, the beats of the packet can be assembled by the destination IP agent in the correct numerical order using the unique beat count number transmitted with each beat. By providing the beat count number transmitted with each beat, many of the above - mentioned problems regarding corruption can be solved. For example, control bit M may be used to identify a unique "beat count number" for each beat of a packet. Then, the beats of the packet can be assembled by the destination IP agent in the correct numerical order using the unique beat count number transmitted with each beat. By providing the beat count number transmitted with each beat, many of the above - mentioned problems regarding corruption can be solved. For example, control bit M may be used to identify a unique "beat count number" for each beat of a packet. Then, the beats of the packet can be assembled by the destination IP agent in the correct numerical order using the unique beat count number transmitted with each beat. By providing the beat count number transmitted with each beat, many of the above - mentioned problems regarding corruption can be solved. By providing the beat count number transmitted with each beat, many of the above - mentioned problems regarding corruption can be solved. As described above, a stream is defined as a pair of a virtual channel and a transaction class. If there are four virtual channels (e.g., VC0, VC1, VC2, and VC3), and three transaction classes (P, NP, C), there are up to twelve different possible streams. The twelve streams are completely independent. Since the streams are independent, they can be interleaved on shared resources such as interconnect wires 12 and 12Z. In each arbitration step, a stream of a virtual channel is selected, and the corresponding port 22 is locked for the rest of that transaction. Another virtual channel can be selected to interleave on the same port before the completion of the transmission of the transaction, but another stream of the same virtual channel cannot be selected until the transaction is completed.
[0160] In-stream interleaving As described above, a stream is defined as a pair of a virtual channel and a transaction class. If there are four virtual channels (e.g., VC0, VC1, VC2, and VC3), and three transaction classes (P, NP, C), there are up to twelve different possible streams. The twelve streams are completely independent. Since the streams are independent, they can be interleaved on shared resources such as interconnect wires 12 and 12Z. In each arbitration step, a stream of a virtual channel is selected, and the corresponding port 22 is locked for the rest of that transaction. Another virtual channel can be selected to interleave on the same port before the completion of the transmission of the transaction, but another stream of the same virtual channel cannot be selected until the transaction is completed. In each arbitration step, a stream of a virtual channel is selected, and the corresponding port 22 is locked for the rest of that transaction. Another virtual channel can be selected to interleave on the same port before the completion of the transmission of the transaction, but another stream of the same virtual channel cannot be selected until the transaction is completed. Another virtual channel can be selected to interleave on the same port before the completion of the transmission of the transaction, but another stream of the same virtual channel cannot be selected until the transaction is completed. Another virtual channel can be selected to interleave on the same port before the completion of the transmission of the transaction, but another stream of the same virtual channel cannot be selected until the transaction is completed.
[0161] In-stream interleaving means that, under the condition that two transactions are independent of each other, it is the interleaving of two or more transactions sharing the same stream. Examples of independent transactions include: (1) two different IP agents 14 generating transactions sharing the same stream, and (2) the same IP agent 14 generating two transactions sharing the same stream, but marking the generating IP agent 14 as independent for the two transactions. Marking a transaction as independent means that the transaction can be reordered and delivered using interleaving. According to in-stream interleaving, the above-mentioned constraint of locking the stream of the virtual channel to the port until the transaction is completed can be relaxed or eliminated. According to in-stream interleaving, (1) two or more independent transactions can be interleaved on the stream, and (2) different streams related to the same virtual channel can also be interleaved. In-stream interleaving is the interleaving of two or more transactions sharing the same stream under the condition that the two transactions are independent of each other. Examples of independent transactions include: (1) two different IP agents 14 generating transactions that share the same stream, and (2) the same IP agent 14 generating two transactions that share the same stream, but marking the generating IP agent 14 as independent for the two transactions. This includes generating transactions that share the same stream by two different IP agents 14, and generating two transactions that share the same stream by the same IP agent 14, but marking the generating IP agent 14 as independent for the two transactions. Generating two transactions that share the same stream by the same IP agent 14, but marking the generating IP agent 14 as independent for the two transactions. Marking the generating IP agent 14 as independent for the two transactions means that the transactions can be reordered and delivered using interleaving. Marking a transaction as independent means that the transaction can be reordered and delivered using interleaving. Marking a transaction as independent means that the transaction can be reordered and delivered using interleaving. According to in-stream interleaving, the above-mentioned constraint of locking the stream of the virtual channel to the port until the transaction is completed can be relaxed or eliminated. According to in-stream interleaving, the above-mentioned constraint of locking the stream of the virtual channel to the port until the transaction is completed can be relaxed or eliminated. According to in-stream interleaving, (1) two or more independent transactions can be interleaved on the stream, and (2) different streams related to the same virtual channel can also be interleaved. According to in-stream interleaving, (1) two or more independent transactions can be interleaved on the stream, and (2) different streams related to the same virtual channel can also be interleaved. According to in-stream interleaving, (1) two or more independent transactions can be interleaved on the stream, and (2) different streams related to the same virtual channel can also be interleaved.
[0162] In in-stream interleaving, additional information is required to indicate two (or more) independent transactions that can be interleaved on the same stream. In various embodiments, this can be achieved in many different ways. In one embodiment, the header 34 of the packets of the independent transactions is assigned a unique transaction identifier, i.e., an ID. By using the unique transaction identifier, each beat of each transaction is flagged as independent. In in-stream interleaving, additional information is required to indicate two (or more) independent transactions that can be interleaved on the same stream. In various embodiments, this can be achieved in many different ways. In one embodiment, the header 34 of the packets of the independent transactions is assigned a unique transaction identifier, i.e., an ID. By using the unique transaction identifier, each beat of each transaction is flagged as independent. By using the unique transaction identifier, each beat of each transaction is flagged as independent. By using a unique transaction ID for the cushion, various nodes 10 2 tracks the bits of multiple independent transactions interleaved on the same stream. Track.
[0163] For a given pair of interleaved transactions, the bits specifying the virtual channel and the transaction class are the same, but the bits representing the transaction ID for each are different. ID bits are different.
[0164] Therefore, the additional transaction ID information included in the control bit M allows the source and both the destination IP agent 14 and the interconnect fabric 106 to recognize or distinguish one transaction from another when interleaved on the same stream. -m and the interconnect fabric 106 to recognize or distinguish one transaction from another when interleaved on the same stream. It is possible to distinguish.
[0165] Synchronous delivery vs. asynchronous delivery In broadcasting, multicasting, read response multicasting , and anycasting, multiple instantiations of the same transaction may be transmitted on the interconnect fabric 106. If each of the target destinations and the paths thereto are available, each destination IP agent 14 delays on the network for the normal latency and eventually receives the transaction. On the other hand, if either the path or the destination is unavailable (e.g., the resource buffer is full), one or more available destinations may receive the transaction before the unavailable destination. Such different arrival times in such situations raise the possibility of two different embodiments. If each of the target destinations and the paths thereto are available, each destination IP agent 14 delays on the network for the normal latency and eventually receives the transaction. Delays for the normal latency on the network and eventually receives the transaction. On the other hand, if either the path or the destination is unavailable (e.g., the resource buffer is full), one or more available destinations may receive the transaction before the unavailable destination. Such different arrival times in such situations raise the possibility of two different embodiments. available destinations may receive the transaction before the unavailable destination. Such different arrival times in such situations raise the possibility of two different embodiments. Such different arrival times in such situations raise the possibility of two different embodiments.
[0166] In the first synchronous or "blocking" embodiment, an effort is made to ensure that each destination receives the transaction almost simultaneously. In other words, the delivery of transactions to available resources is delayed, that is, "blocked", until the unavailable resources become available. As a result, the receipt of transactions by each of the designated recipients is synchronized. This embodiment may be used in application examples where it is important for the receiving side to receive the transactions almost simultaneously.
[0167] In the second asynchronous or non-blocking embodiment, no blocking effort is made to delay the delivery of transactions to available destinations. Instead, each instantiation of the transaction is delivered based on availability, which means that available resources receive the transaction immediately while unavailable resources receive the transaction when they become available. As a result, the delivery is asynchronous, that is, it can occur at different times. The advantage of this approach is that the available destination IP agent 14 can process the transaction immediately and is not blocked waiting to synchronize with other IP agents. As a result, delays are avoided.
[0168] Although only some embodiments have been described in detail, it should be understood that the present application can be implemented in many other forms without departing from the spirit and scope provided here. Therefore, these embodiments are considered to be illustrative and not limiting, and are not limited to the details shown herein, and can be modified within the scope of the appended It may be done.
Claims
Claim 1 A system-on-chip (SoC) comprising: An interconnect fabric; and A plurality of IP agents interconnected by the interconnect fabric, the plurality of IP agents being configured to be sources and destinations of transaction traffic transmitted via the interconnect fabric between the IP agents. A first IP agent is configured to generate and transmit a transaction that is one of a broadcast, multicast, read-response multicast, or anycast type of transaction. The interconnect fabric is configured to make a routing decision as to how to route and deliver the transaction via the interconnect fabric to one or more destination IP agents among the plurality of IP agents on the SoC. The transaction comprises one of a plurality of uniquely encoded commands, each of the uniquely encoded commands indicating to the interconnect fabric that the transaction is respectively any one of the broadcast, multicast, read-response multicast, or anycast type of transaction. An SoC. Claim 2 The SoC according to claim 1, wherein the interconnect fabric is configured to make the routing decision so as to reduce transaction traffic transmitted via the interconnect fabric by integrating and transmitting one instantiation of the transaction via a shared interconnect between two nodes of the interconnect fabric, the shared interconnect leading to two or more destination IP agents along a delivery path. An SoC. Claim 3 The SoC according to claim 1 or 2, wherein the interconnect fabric is configured to make the routing decision to reduce transaction traffic transmitted through the interconnect fabric by integrating a plurality of received response transactions generated in response to the transaction and transferring one instantiation of the response transaction via a shared interconnect among the nodes of the interconnect fabric, and the shared interconnect reaches the destination of the response transaction along a response delivery path, SoC.
4. The SoC according to any one of claims 1 to 3, wherein the interconnect fabric is further configured to make the routing decision to interleave the transmission of a plurality of transactions on a plurality of streams associated with the shared interconnect, (a) each of the plurality of streams is defined by a unique combination of a virtual channel and a transaction type, (b) one or more portions of a given transaction are transmitted via the same stream, SoC.
5. The SoC according to any one of claims 1 to 4, wherein the interconnect fabric is further configured to make the routing decision to route two or more independent transactions via the same stream among the plurality of streams associated with the shared interconnect, and each of the plurality of streams is defined by a unique combination of a virtual channel and a transaction type, SoC.
6. The SoC according to any one of claims 1 to 5, wherein the interconnect fabric comprises duplicate physical resources that share a common logical identifier, and the routing decision comprises (a) selecting one of the duplicate physical resources; and (b) routing the transaction using the selected one of the duplicate physical resources, SoC.
7. The SoC according to claim 6, wherein the duplicate physical resources comprise (a) duplicate IP agents; (b) duplicate shared interconnects between two nodes of the interconnect fabric; (c) a shared line between an IP agent and a node of the interconnect fabric, and includes one of them, SoC.
8. The SoC according to claim 6, wherein the routing decision is (a) the availability of the duplicate physical resources, (b) the relative mixing state between the duplicate physical resources, (c) the load balancing between the duplicate physical resources, (d) the random selection between the duplicate physical resources, (e) the selection of the longest unused time between the duplicate physical resources, (f) the relative power consumption between the duplicate physical resources, (g) the use of a hash function for making a selection between the duplicate physical resources, or (h) the SoC based on one of any combinations of (a) to (g).
9. The SoC according to any one of claims 1 to 8, wherein the routing decision on how to route and deliver the transaction via the interconnect fabric includes synchronizing the delivery of the transaction to two or more destination IP agents.
10. The SoC according to any one of claims 1 to 9, wherein the routing decision on how to route and deliver the transaction via the interconnect fabric includes allowing the delivery of the transaction to two or more destination IP agents to occur asynchronously.
11. The SoC according to any one of claims 1 to 10, wherein the interconnect fabric includes a plurality of nodes, and at least one of the plurality of nodes includes a lookup table for resolving an address specified in the transaction into a logical identifier of the one or more destination IP agents.
12. The SoC according to any one of claims 1 to 11, wherein the interconnect fabric includes a plurality of nodes, and each of the plurality of nodes includes a table for converting a logical identifier of the one or more destination IP agents into one or more port identifiers used for routing the transaction.
13. The SoC according to any one of claims 1 to 12, wherein the transaction has a unique address, and the unique address indicates to the interconnect fabric that the transaction is any one of the broadcast, multicast, read response multicast, or anycast type of transaction, respectively.
14. The SoC according to any one of claims 1 to 13, wherein the transaction has a unique address, and the unique address is resolved by the interconnect fabric to one of: (a) the logical identifier of the one or more destination IP agents, (b) a unique code that designates the logical identifier of the one or more destination IP agents, and (c) a vector of bits representing the logical identifier of the one or more destination IP agents.
15. The SoC according to any one of claims 1 to 14, wherein the transaction is a broadcast, and the interconnect fabric makes a routing decision to route and deliver the transaction to each of the plurality of IP agents on the SoC.
16. The SoC according to any one of claims 1 to 14, wherein the transaction is a multicast, and the interconnect fabric makes a routing decision to route and deliver the transaction to two or more of the plurality of IP agents on the SoC.
17. The SoC according to any one of claims 1 to 14, wherein the transaction is a read response multicast transaction, the transaction is delivered to one destination IP agent, but the interconnect fabric routes and delivers response transactions to two or more of the plurality of IP agents on the SoC.
18. The SoC according to any one of claims 1 to 14, wherein the transaction is an anycast, and the interconnect fabric selects the one or more destination IP agents of the transaction.
19. An SoC according to any one of claims 1 to 18, wherein the interconnect fabric is further configured to make the routing decision so as to route two or more independent transactions via the same stream, and each of the two or more independent transactions is assigned a unique transaction identifier that enables the interconnect fabric to track each beat of the two or more independent transactions such that the two or more independent transactions can be routed via the same stream.
20. An SoC according to claim 19, wherein the two or more independent transactions routed via the same stream share common control information specifying the same stream, but each of the transaction identifiers of the two or more independent transactions is unique.
21. A system-on-chip (SoC) comprising: an interconnect fabric; a plurality of IP agents interconnected by the interconnect fabric, the plurality of IP agents being configured to be sources and destinations of transaction traffic transmitted via the interconnect fabric among the IP agents; a first IP agent configured to generate and transmit a transaction that is one of broadcast, multicast, read response multicast, or anycast type transactions; the interconnect fabric is configured to make a routing decision as to how to route and deliver the transaction via the interconnect fabric to one or more destination IP agents among the plurality of IP agents on the SoC; The interconnect fabric is further configured to make the routing decision so as to route two or more independent transactions through the same stream, and each of the two or more independent transactions is assigned a unique transaction identifier that enables the interconnect fabric to track each beat of the two or more independent transactions such that the two or more independent transactions can be routed through the same stream, SOC.
Citation Information
Patent Citations
Providing A Sideband Message Interface For System On A Chip (SoC)
US20130138858A1
Method for data throughput improvement in open core protocol based interconnection networks using dynamically selectable redundant shared link physical paths
US20130268710A1
Supporting multicast in noc interconnect
US20150043575A1
Semiconductor integrated circuit and its testing method
WO2008126471A1
Network-on-chip, network routing method, and system
WO2010137572A1