Transmission of address translation type packets

By adopting the time division multiplexing technology of virtual channels in the queue of the computing system, the problems of bandwidth limitation and queue delay in traditional communication structures are solved, and efficient data transmission and performance improvement are achieved.

CN117716679BActive Publication Date: 2025-05-20ATI TECHNOLOGIES ULC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202280051797.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-06-24
Filing Date
2022-07-12
Publication Date
2025-05-20
Estimated Expiration
2042-07-12

AI Technical Summary

Technical Problem

In computing systems, traditional communication structures are limited in bandwidth when connecting independent processing nodes, resulting in inefficient data transmission, especially the priority sorting and latency problems in the middle of queues, resulting in performance degradation.

Method used

By introducing time division multiplexing technology of virtual channels into the queue, multiple virtual channels are created using resources on existing physical channels, thereby improving the routing efficiency of requests and responses to shared resources.

Benefits of technology

It realizes that the bandwidth and efficiency of data transmission are improved without increasing the number of physical channels, reduce queue delay and priority sorting problems, and improve system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117716679B_ABST
    Figure CN117716679B_ABST
Patent Text Reader

Abstract

The present invention provides a device, system and method for routing requests and responses targeted at shared resources. A queue in a communication structure is located in a path between a requestor and a shared resource. In some embodiments, the shared resource is a shared address translation cache stored in an endpoint. The physical channel between the queue and the shared resource supports multiple virtual channels. The queue allocates at least one entry to each virtual channel in a group of virtual channels, wherein the group includes a virtual channel for each address translation request type of a single requestor from a plurality of requestors. When at least one entry for a given requestor is de-allocated, even if an empty entry is the only available entry for the queue, the queue only allocates the entry using a request from an allocated virtual channel.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND OF THE INVENTION

[0001] Description of the Related Art

[0002] In a computing system, some types of applications are able to better utilize the capabilities of parallel processing and shared memory than others. Examples of such applications include machine learning applications, entertainment and real-time applications, and some commercial, scientific, medical, and other applications. Although some processor architectures include more than one processing unit (e.g., CPU, GPU, multimedia engine, etc.) or processing core, in some cases, one or two additional processing units or cores coupled to memory may not provide a sufficient level of parallelism to provide the required level of performance.

[0003] In addition to read and write access commands and corresponding data, coherence probes, interrupts, and other communication messages are also transmitted in the system via a communication fabric (or fabrics). Examples of the interconnects in a fabric are a bus architecture, a crossbar-based architecture, a network-on-chip (NoC) communication subsystem, a communication channel between dies, a silicon interposer for side-by-side stacked chips, a through-silicon via (TSV) for vertically stacking dedicated dies on top of a processor die, and so on.

[0004] In many cases, a fabric has multiple physical channels, each physical channel supporting a relatively wide packet. When transmitting data within a single fabric, the fabric reduces latency because a relatively large number of physical lines are available. However, when connecting independent dies together via a fabric, and when connecting independent processing nodes (each with its own fabric) together, data is transmitted over a significantly smaller number of physical lines, which limits the available bandwidth. In some cases, the data rate at which data is transmitted over the link physical lines is a multiple of the data rate of the physical lines on the die. However, when communicating between dies and between nodes, the bandwidth is still significantly reduced.

[0005] In addition to the above inefficiencies in transmitting data, the intermediate queues in a communication fabric have the potential to become full or to prioritize the entries to be posted based on epochs. These entries store packets including requests or corresponding responses. When the queue posts lower-priority packets, higher-priority packets may wait. When one or more queues delay high-priority requests on the first path from a requester to a shared resource, and additionally, one or more queues delay high-priority responses corresponding to requests on the second path from the shared resource to the requester, performance is affected.

[0006] In view of the above, there is a need for effective methods and systems for efficiently routing requests and responses targeted at a shared resource. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Figure 1 It is a schematic diagram of a queue storing at least address translation requests.

[0008] Figure 2 It is a schematic diagram of one embodiment of a packet transmitter that processes at least address translation requests.

[0009] Figure 3 It is a schematic diagram of a computing system that transmits at least address translation requests.

[0010] Figure 4 It is a schematic diagram of one embodiment of a method for efficiently allocating entries of a queue storing at least address translation requests.

[0011] Figure 5 It is a schematic diagram of one embodiment of a method for efficiently issuing from entries of a queue storing at least address translation requests.

[0012] Although the present invention may have various modifications and alternative forms, specific embodiments are shown by way of example in the drawings and are described in detail herein. However, it should be understood that the drawings and the detailed description thereof are not intended to limit the present invention to the particular forms disclosed, but on the contrary, the present invention is to cover all modifications, equivalents, and alternatives falling within the scope of the present invention as defined by the appended claims. Detailed Description

[0013] In the following description, numerous specific details are set forth to provide a thorough understanding of the present invention. However, those of ordinary skill in the art should recognize that the present invention may be practiced without these specific details. In some instances, well-known circuits, structures, and techniques are not shown in detail to avoid obscuring the present invention. Additionally, it should be understood that, for simplicity and clarity of illustration, the elements shown in the figures are not necessarily drawn to scale. For example, the dimensions of some elements are enlarged relative to other elements.

[0014] Apparatuses, systems, and methods for efficiently routing requests and responses targeting shared resources are envisioned. In various embodiments, a computing system includes a shared resource accessed by a plurality of requesters via a communication fabric. In various embodiments, the shared resource is a copy of at least a portion of the virtual-to-physical address translation of a page table. In some cases, the destination is a memory controller storing the shared page table. In other cases, the destination is an I / O peripheral device (or peripheral) storing a copy of a portion of the shared page table. A queue in the communication fabric is located in the path between the source requesting access to the shared page table and the destination including the shared page table. The queue stores requests in its entries, which request service from the destination.

[0015] The queue includes a unidirectional channel that transfers data from the queue to a destination. The data transferred includes one or more requests issued by the queue selection. The unidirectional channel is referred to as a "physical channel" and includes a predetermined number of physical lines between the queue and the destination. Thus, the physical channel has limitations on the physical resources that support it, such as transmitter and receiver circuits and storage elements. The queue supports time-division multiplexing on an existing physical channel rather than increasing the throughput of multiple sources requesting access to the destination by increasing the number of physical channels. The time-division multiplexing supported by the queue creates multiple "virtual channels" between the queue and the destination. Thus, data transfer on a single physical channel is for multiple virtual channels without increasing physical resources. In some cases, virtual channels are assigned to specific sources. In other cases, virtual channels are assigned to specific request types. In other cases, virtual channels are assigned to specific request types from specific sources.

[0016] The circuitry of the queue between the source and the destination includes an arbitration unit and a control unit. The control unit of the queue maintains allocated and unallocated entries of the queue that store requests from multiple sources. In one embodiment, the control unit assigns at least one entry to each virtual channel in a group of virtual channels, where the group includes virtual channels for each address translation request type from a single source among the multiple sources. An example of an address translation request is a read request that requests retrieval of one or more virtual-to-physical address translation copies from the destination. Another example of an address translation request is an invalidate request that requests invalidation of one or more stored virtual-to-physical address translations stored at the destination.

[0017] The arbitration unit selects one or more requests from the allocated entries for publication based on a selection criterion such as one or more attributes. When at least one or more of the allocated entries in a given virtual channel are empty, the circuitry allocates entries using the requests for the given virtual channel stored in the unallocated entries. However, if the control unit determines that there are no unallocated entries allocated for a given virtual channel, the control unit leaves the one or more allocated entries empty. For example, even if the one or more allocated entries are the only available entries in the queue, the control unit does not allocate the one or more allocated entries to other virtual channels.

[0018] Reference Figure 1, a schematic block diagram of an embodiment of a queue 100 that stores at least address translation requests is shown. In the illustrated embodiment, queue 100 includes queue entries 110 (or entries 110), an arbitration unit 130, and a control unit 140. Additionally, two keys are provided in the lower right corner. The key provides a mapping that describes the device type associated with the source, as well as a mapping of virtual channels to the source and request types. Entries 110 store various request types from one or more sources that are sent to one or more destinations. Entries 110 include allocated entries 112, 114, and 116, and unallocated entry 118. The arbitration unit 130 includes circuitry that selects, based on selection criteria such as one or more attributes, one or more requests to be issued to the destination from the allocated entries 112, 114, and 116. The one or more issued requests are sent to the destination on a physical channel 132. The entries of the issued requests are later allocated based on virtual channels by requests stored in unallocated entry 118. The control unit 140 includes circuitry that maintains the allocation of entries 110 and determines how many entries are provided to allocated entries 112, 114, and 116, and unallocated entry 118.

[0019] Figure 1 The lower right corner of includes two keys. The first key includes a mapping from the source to the device type. For example, "Src 1" or source 1 is the first central processing unit (CPU) of a computing system. Similarly, "Src 2" or source 2 is the second central processing unit (CPU) of the computing system, and so on. Although a specific number and type of devices are shown in the first key, in other embodiments, the computing system using queue 100 includes another number and type of devices. The second key includes a mapping of virtual channels (VCs) to the source and request types. For example, "VC 1" or virtual channel 1 is allocated to memory write request types from sources 1 and 2, where the sources are CPU 1 and CPU 2. It should be noted that although a virtual channel is allocated to memory write requests and memory read requests from multiple sources, a single virtual channel is allocated to a combination of a single source and an address translation request type. Thus, when a source generates an address translation request type, no source shares a virtual channel. As shown, each virtual channel has a single allocated entry among allocated entries 112, 114, and 116. In other embodiments, a virtual channel has one or more entries among allocated entries 112, 114, and 116.

[0020] Entry 110 is implemented using one of a variety of random access memories (RAMs), content addressable memories (CAMs), multiple registers or flip-flop circuits, or other circuits. In some embodiments, entry 110 stores various transactions from multiple sources targeting the same destination. Queue 100 receives transactions from multiple sources via a communication fabric. The communication fabric supports the transmission of requests, responses, and other types of messages between sources and destinations. Examples of sources are a central processing unit (CPU), a multimedia engine that processes one or more of audio and video data, an application specific integrated circuit (ASIC), a graphics processing unit (GPU), one of various input / output (I / O) peripherals, and so on. Sources are also referred to as requesters and "clients" that are capable of generating requests and messages to be serviced by destinations. Examples of destinations include a memory controller and examples of sources when a received request requests the source to perform a service. Destinations are also referred to as "endpoints" that are devices that service requests received from clients targeting the device.

[0021] Examples of transactions stored in entry 110 are memory read requests, memory write requests, memory snoop requests, token or credit messages, and address translation requests. Other examples of request types are also included in other embodiments. One example of an address translation request is a read request to retrieve a copy of one or more virtual-to-physical address translations from a destination. Another example of an address translation request is an invalidate request to invalidate one or more stored virtual-to-physical address translations stored at a destination. Although the two keys at the lower right corner provide information for memory write requests, memory read requests, and address translation requests, in other embodiments, other types of requests are included in the allocation of virtual channels.

[0022] The operating system allocates virtual address space to software processes and divides the address space into blocks of a specific size. Each block is a "page" of the address space. Virtual pages are mapped to frames of physical memory, and the mapping of virtual addresses to physical addresses tracks the storage location of virtual pages in physical memory. These mappings are stored in a page table, and the external system memory stores the page table. A copy of at least a portion of the page table is stored in the address translation cache of a destination. The address translation cache may also be referred to as a translation lookaside buffer (TLB). Examples of destinations that store a copy of at least a portion of the page table include a memory controller for system memory and one or more of endpoints such as I / O peripherals (or peripherals).

[0023] Queue 100 includes physical channel 132 that transfers one or more requests selected for publication from entry 110 to a destination. The data transferred includes one or more requests selected for publication by the queue. Physical channel 132 includes a predetermined number of physical lines between queue 100 and the destination. Thus, physical channel 132 has limitations on the physical resources to support it, such as transmitter and receiver circuitry and storage elements. Control unit 140 supports time division multiplexing on the existing physical channel 132 instead of increasing the number of physical channels 132 to increase throughput. The time division multiplexing supported by control unit 140 creates multiple "virtual channels" between queue 100 and the destination.

[0024] In some cases, control unit 140 assigns virtual channels to a specific source. In other cases, control unit 140 assigns virtual channels to a specific request type. For example, the control unit maintains assigned entry 112 for memory write requests. Additionally, the control unit maintains assigned entry 114 for memory read requests. Further, the control unit maintains assigned entry 116 for address translation requests. Although three groups of assigned entries are shown, each with a specific number of entries, another number of groups of assigned entries is also feasible and contemplated. Each group of assigned entries may also have another number of entries different from the number shown.

[0025] Although assigned entry 116 is assigned to the address translation type, it should be noted that in the illustrated embodiment, a single entry is assigned to a specific virtual channel. Here, the virtual channel is assigned to an address translation request from a specific source. In one example, virtual channel (VC) 8 is assigned to an address translation request from a display engine, VC 9 is assigned to an address translation request from a first type of peripheral device, and VC 10 is assigned to an address translation request from a second type of peripheral device, and so on. Although virtual channel identifiers 7 - 10 are used here, in other embodiments, other identifiers are used. Thus, control unit 140 assigns at least one entry to each virtual channel in a group of virtual channels, where the group includes virtual channels for each address translation request type from a single source among multiple sources. In contrast, VC 1 is assigned to memory write requests from two or more sources such as CPU 1 and CPU 2.

[0026] As shown, entry 110 includes a plurality of fields 120 - 128 for storing information corresponding to received requests. Although a specific number and type of fields are shown in entry 110, in other embodiments, a different number of fields and other types of fields are stored in entry 110. As shown, entry 110 stores metadata such as status field 120, which stores an indication of whether at least a valid (V) bit of the corresponding entry is utilized for the request. Field 122 stores a virtual channel (VC) identifier (ID). The virtual channel ID is dynamically assigned during the runtime of the application. During the allocation phase, separate virtual channel IDs are assigned to address translation request types from a single source out of multiple sources. Field 124 stores a source or client identifier (ID) shown as "Src". The source ID identifies the source that generates and sends the request.

[0027] Entry 110 also includes a field 126 that stores an arbitration value indicated by "Arb". The arbitration value is based on selection criteria such as one or more attributes. Examples of such attributes are the priority level of the received request, quality of service (QoS) parameters, source identifier (ID), application ID or type (such as a real - time application), virtual channel ID, bandwidth requirements or latency tolerance requirements, indication of an epoch, and so on. In some embodiments, these values are stored in respective fields of entry 110, and an arbitration unit 130 receives one or more of these values and determines the final attribute values of the entries in allocated entries 112, 114, and 116. In other embodiments, a control unit 140 receives one or more of these values and determines the final attribute values to be stored in entry 110. The control unit 140 updates these arbitration values based on at least the epoch.

[0028] Field 128 of entry 110 includes request types such as memory write requests, memory read requests, and address translation requests. Examples of other types of information stored in other fields (not shown) of queue entry 110 are the target address corresponding to the request, an indication of the data size to be read or written, an indication of the epoch of the corresponding request, destination ID, the ID of the previous hop within the communication structure before queue 100 received the corresponding request, software process ID, application ID, an indication of data type such as real - time data or non - real - time data, and so on.

[0029] Based at least on the final values determined according to the selection criteria, the arbitration unit 130 selects one or more requests from the allocated entries 112, 114, and 116 to issue to the destination via the physical channel 132. In some embodiments, when the arbitration unit 130 selects two or more requests from the allocated entries 112, 114, and 116 for issuance, the arbitration unit 130 selects at least one request from the allocated entry 116 during each arbitration phase. Each arbitration phase can be one or more clock cycles or pipeline stages, depending on the specific implementation. In other words, each time a request is issued from the queue entry 110, the arbitration unit 130 selects at least one request from the allocated entry 116. Thus, the arbitration unit 130 selects at least one address translation request from the virtual channel group of the allocated entry 116 for issuance during each arbitration phase, but the selected address translation request has a lower arbitration value than one or more requests in the allocated entries 112 and 114. In one example, the arbitration unit 130 selects two requests for issuance during each arbitration phase. Instead of selecting two memory read requests of entry 114 with virtual channels 5 and 6 having arbitration values 10 and 9, the arbitration unit 130 selects a memory read request with VC 5 of entry 114 and an arbitration value of 10 and an address translation request with VC 9 of entry 116 and an arbitration value of 6. In this way, the arbitration unit 130 ensures that the address translation request is sent to the destination as quickly as possible.

[0030] When one or more of the entries in the allocated entries 116 are empty (V = 0) or deallocated, the control unit 140 allocates these entries using the address translation requests from the corresponding virtual channels stored in the unallocated entries 118. The entries in the allocated entries 112 and 114 are allocated in a manner similar to the requests of the type assigned to the entries. However, if the control unit 140 determines that there are no entries in the unallocated entries 118 allocated for the corresponding virtual channel, the control unit 140 leaves one or more of the entries in the allocated entries 116 empty. For example, even if one or more of the allocated entries are the only available entries in the entries 110, the control unit 140 does not allocate the one or more allocated entries in the entries 116 for other virtual channels. For example, if the entry in the allocated entries 116 for the address translation type for VC = 1 and Arb = 8 is deallocated because it is selected for issue, and there is no entry in the entries 118 with a request for VC = 1, the entry in the allocated entries 116 assigned to VC = 1 remains deallocated despite the requests stored in the unallocated entries 118 requiring entries in the allocated entries 112, 114, and 116. Another example is the entry for VC = 7 in the entries 116, which remains deallocated when the unallocated entries 118 do not have a request for VC = 7, despite the requests stored in the unallocated entries 118 requiring entries in the allocated entries 112, 114, and 116.

[0031] In some embodiments, the control unit 140 allocates received requests in an ordered and sequential manner starting from the allocated entries in the entries 112, 114, and 116 based on the virtual channel. In such embodiments, the control unit 140 keeps the earliest request corresponding to a particular virtual channel in the corresponding entry in the allocated entries 112, 114, and 116. From the perspective of a particular virtual channel, the queue 100 appears to provide a first-in-first-out (FIFO) data store. It should be noted that the queue 100 is in the path between the source requesting access to the shared address translation cache and the destination including the shared address translation cache. In some embodiments, at least the queue 100 and the destination support an interconnection communication protocol, and the protocol includes specifications for routing address translation requests. In some embodiments, the supported specification is the address translation service (ATS) specification of the PCIe (Peripheral Component Interconnect Express) interconnection communication protocol. The ATS specification supports remote caching (storage) of address translations at the endpoints. In other embodiments, another specification and another interconnection communication protocol are supported.

[0032] It should also be noted that another queue similar to queue 100 is used to store responses corresponding to at least address translation requests. The types of responses include completion acknowledgments indicating whether an address translation read request is approved, acknowledgments indicating whether an invalid request is approved, and response data such as one or more copies of the requested virtual-to-physical address translation. In various embodiments, this other queue is managed in a manner similar to queue 100, and responses stored in queue entries are processed in a similar manner.

[0033] Go to Figure 2 , a schematic block diagram of one embodiment of a structured packet transmitter 200 is shown. In the illustrated embodiment, the structured packet transmitter 200 includes queues 210 and 230, each queue for storing a corresponding type of packet. In some embodiments, the packets are flow control units (“microtiles”). A microtile is a subset of a larger packet. Microtiles typically carry data and control information, such as header and trailer information for the larger packet. Although the data to be transmitted is described as packets routed in a network, in other embodiments, the data to be transmitted is a bit stream or byte stream in a point-to-point interconnect. In various embodiments, queues 210 and 230 store control packets to be sent over a structured link. The corresponding data packets (such as the larger packets corresponding to the microtiles) are stored in other queues. In other embodiments, queues 210 and 230 store data packets, and the corresponding control packets are stored in other queues.

[0034] When the structured packet transmitter 200 is placed in the data stream from the source that generates the request to the destination that services the request, examples of the types of control packets stored in queues 210 and 230 are memory read request types, memory write request types, probe (snoop) message types, token or credit types, address translation read access types, and address translation invalidation types. When the structured packet transmitter 200 is placed in the data stream from the destination that services the request to the source that waits for the request to be serviced, examples of the types of control packets stored in queues 210 and 230 are read response types, write response types, probe (snoop) response types, address translation read access response types, and address translation invalidation response types. Other examples of packet types are also included in other embodiments.

[0035] In some embodiments, queue 210 stores packets of "Type 1" which are of a control request type. Queue 230 stores packets of "Type N", which is an address translation request type in one embodiment. In other embodiments, "Type 1" and "Type N" correspond to different virtual channels rather than request types. As previously described, an example of an address translation request is a read request to retrieve one or more copies of virtual-to-physical address translations from a destination. Another example of an address translation request is an invalidate request to invalidate one or more stored virtual-to-physical address translations stored at a destination. The queues between queue 210 and 230 store packets of "Type 2" to "Type N-1", which include other control response types or other different virtual channels, depending on the particular implementation. Thus, although only two queues are shown in Figure 2 , the structured packet transmitter 200 includes any number of queues. Although queues 210 and 230 are shown as separate queues, in other embodiments, the entries of queues 210 and 230 are maintained in a single queue.

[0036] Queues 210 and 230 are implemented using one of a variety of random access memories (RAMs), content addressable memories (CAMs), multiple registers or flip-flop circuits, or other circuitry. When the structured packet transmitter 200 receives a new packet, the control unit 220 uses hardware such as circuitry to determine which entries of queue 210 to allocate. When a packet is allocated to and posted from queue 210, the control unit 220 also updates the credit or token assigned to the source of the packet. For example, the control unit 220 determines the minimum number of clock cycles (or periods) between receiving new packets in order to avoid data conflicts in the entries of queue 210 when the entries become full or within a threshold number of full entries. The control unit 240 has a similar function, but the manner of accessing data in queue 210 may be different from the manner of accessing data in queue 230 due to the type of packets stored in queue 230.

[0037] In various embodiments, queue 230, control unit 240, and queue arbiter 242 have similar functions as previously described for ( Figure 1 of) queue 110, control unit 130, and arbitration unit 120. However, in addition to the virtual channel assigned to the address translation type, ( Figure 1Queue 110 also stores requests for virtual channels. Here, queue 230 stores requests for virtual channels that are only assigned to the address translation type. The control unit 240 assigns at least one entry to each virtual channel in a group of virtual channels, where the group includes virtual channels for each address translation request type from a single source among multiple sources. In some embodiments, the structured packet transmitter 200 is used at an intermediate position within the communication structure, while ( Figure 1 queue 100 serves as the last queue before the destination.

[0038] The queue arbiter 222 uses circuitry to select packets stored in the entries of queue 210 for transmission over the structured link. In some embodiments, the queue arbiter 222 determines the priority level of packets from the assigned entries based on one or more attributes. As previously described, these attributes are one or more of the following: the priority level of the received request, quality of service (QoS) parameters, source identifier (ID), application ID or type (such as a real-time application), virtual channel ID, bandwidth requirement or latency tolerance requirement, indication of an epoch, and so on. When the structured link is available, one or more candidate packets 224 are transmitted over the structured link. Similarly, the queue arbiter 242 selects one or more candidate packets 244 from queue 230 for transmission over the structured link. In some embodiments, the queue arbiters 222 - 242 select candidate packets 224 - 244 from queues 210 - 230 every clock cycle. In other embodiments, packets are selected after the previously selected candidate packets 230 - 234 have been inserted into the link packet and transmitted over the structured link.

[0039] As previously described, queue 230 stores "type N" packets, which are packet types corresponding to address translation requests or packet types corresponding to address translation responses (depending on the direction of the data flow upstream or downstream in the communication structure). For example, a requester with permission to access a specific page table generates a TLB miss, and the requester has sent an address translation request to a memory controller or other endpoint that controls access to a copy of at least a portion of the specific page table. In various implementations, the address translation request will initiate a page table walk. In other implementations, the address translation request will access a specific TLB that stores a copy of the requested address translation from a specific page table. Queue 230 is an intermediate queue on the path from the requester to the memory controller or other endpoint. Alternatively, the memory controller or other endpoint sends a corresponding address translation response to the requester, and queue 230 is an intermediate queue on the path from the memory controller or other endpoint to the requester. The received "type N" packets are stored in one of the entries 252 - 266 of queue 230.

[0040] In some embodiments, the address translation is stored in a shared resource, such as a memory controller or other endpoint that stores a shared page table. In some embodiments, the address translation request is a request based on the Address Translation Service (ATS) specification of the PCIe (Peripheral Component Interconnect Express) interconnect communication protocol. The ATS specification supports remote caching (storage) of address translations on endpoints. Queue 230 and support circuitry, such as control unit 240 and queue arbiter 242, reduce the latency in servicing address translation requests for requesters accessing the shared page table. For example, control unit 240 holds entries 252 - 254 as allocated entries 250, while control unit 240 holds entries 262 - 266 as unallocated entries 260. In various embodiments, each requester having access to a particular page table has at least one allocated entry in allocated entries 250. Each address translation request from a particular requester among multiple requesters is assigned a particular virtual channel. For a particular virtual channel, when each allocated entry in at least one of the allocated entries 250 is allocated and the received packet corresponds to that particular requester, control unit 240 selects an available entry in unallocated entries 260 for allocation.

[0041] Queue arbiter 242 selects one or more packets from the packets stored in allocated entries 250 for issuance. In one embodiment, if the packets in allocated entries 250 exceed a time period threshold, queue arbiter 242 selects the packet. Otherwise, queue arbiter 242 selects packets from allocated entries 250 for issuance based on one or more attributes as described above. Additionally, queue arbiter 242 is capable of selecting packets from allocated entries 250 based on a least recently selected algorithm, a round-robin algorithm, or other algorithms. When the allocated entries in entry 250 for a particular requester are empty (deallocated), control unit 240 allocates these entries in entry 250 with packets from unallocated entries 260 corresponding to the particular requester. Thus, in some embodiments, control unit 240 holds the earliest packet from a given virtual channel in one of allocated entries 250. For example, control unit 240 services packets from a particular virtual channel in an ordered manner.

[0042] However, if the control unit 240 determines that there are no unallocated entries of the entry 260 allocated for a specific virtual channel, the control unit 240 keeps the allocated one or more entries of the entry 250 empty for the specific virtual channel. For example, even if these allocated one or more entries are the only available entries in the queue 230, the control unit 240 does not utilize packets from other virtual channels to allocate these empty entries of the entry 250. Thus, although the current specific virtual channel has no allocated entry in the queue 230, at least one entry in the entry 250 is still available for the specific virtual channel. Thus, no virtual channel having access to the shared page table has an address translation packet (request or response) that would be blocked at the queue 230 due to no available entry in the entry 252 - 266. Instead, it is guaranteed that each virtual channel has at least one available entry in the allocated entry 250.

[0043] Now turning to Figure 3 , a schematic block diagram of a particular implementation of a computing system 300 that transmits at least address translation requests is shown. As shown, the computing system 300 includes a communication fabric 310 between each of a memory controller 340, peripherals 380 and 390, and a plurality of clients. The memory controller 340 is operative to interact with a memory 350. Examples of the plurality of clients are a central processing unit (CPU) 360, a graphics processing unit (GPU) 362, a hub 364, and peripherals 380 and 390. The hub 364 is operative to communicate with a multimedia engine 368. In some particular implementations, one or more hubs are operative to interact with a multimedia player (i.e., the hub 364 for the multimedia engine 368), a display unit, or other units. In such cases, the hub is a client in the computing system 300. In some particular implementations, one or more of the peripherals 380 and 390 use a hub. Each hub further includes control circuitry and storage elements for handling data transmission according to various communication protocols. Although five clients 360, 362, 364, 380, and 390 are shown, in other particular implementations, the computing system 300 includes any number of clients and other types of clients, such as a display unit, one or more other input / output (I / O) peripherals, and the like.

[0044] In some specific implementations, computing system 300 is a system-on-chip (SoC), where each of the depicted components is integrated on a single semiconductor die. In other specific implementations, these components are separate dies in a system-in-package (SiP) or a multi-chip module (MCM). In various specific implementations, CPU 360, GPU 362, multimedia engine 366, and peripherals 380 and 390 are used in smartphones, tablets, game consoles, smartwatches, desktop computers, virtual reality headsets, or other devices. CPU 360, GPU 362, multimedia engine 366, and peripherals 380 and 390 are examples of clients capable of generating network-on-chip data to be transmitted. Examples of network data include memory access requests, memory access response data, memory access acknowledgments, probes and probe responses, address translation requests, address translation responses, address translation invalidation requests, and other network messages between clients. This network data is placed in network packets (or packets). Each packet includes a specific type of network data. For example, one packet includes one or more requests, another packet includes one or more responses, and so on. These packets include headers with metadata that includes multiple identifiers for at least identifying the source, destination, virtual channel, packet type, data size of response data, priority level, application generating the message, and so on.

[0045] To effectively route packets, in various specific implementations, communication fabric 310 uses a routing network 320 that includes network switches. In various specific implementations, one or more of fabric 310 and routing network 320 include status and control registers for storing control parameters. In some specific implementations, fabric 310 includes hardware such as circuits for supporting communication, data transfer, and network protocols for routing packets over one or more buses. Fabric 310 includes circuits for supporting address formats, interface signals, and the use of synchronous / asynchronous clock domains. In some specific implementations, the network switches of fabric 310 are network-on-chip (NoC) switches. In one specific implementation, routing network 320 uses multiple network switches in a point-to-point (P2P) ring topology. In other specific implementations, routing network 320 uses network switches with programmable routing tables in a mesh topology. In still other specific implementations, routing network 320 uses network switches in a combination of topologies. In some specific implementations, routing network 320 includes one or more buses to reduce the number of lines in computing system 300. For example, one or more of interfaces 330 - 332 send read responses and write responses on a single bus within routing network 320.

[0046] Each of the CPU 360, GPU 362, multimedia engine 366, and peripherals 380 and 390 can act as both a source and a destination. A source generates requests for destinations to service. A destination services the requests and sends any responses back to the corresponding source. As previously mentioned, the CPU 360, GPU 362, multimedia engine 366, and peripherals 380 and 390 are referred to as clients, but these components are also endpoints. As previously mentioned, an endpoint is a device that acts as a destination for servicing device-targeted requests.

[0047] In various embodiments, one or more of the fabric 310, routing network 320, interfaces 312, 314, 316, 330, 332, and 334, and memory controller 340 use intermediate queues, such as queues 370 - 373, to store packets transmitted between a source and a destination. Although only the routing network 320 is shown as using queues 370 - 373, it is also feasible and contemplated that other components include similar queues. Queues 370 - 373 have attached control units (CUs) 374 - 377, which have hardware that performs various functions, such as control circuitry and storage elements. Examples of these functions are controlling access to queue entries, dispatching packets from queue entries, and any reordering of the storage of packets within queue entries. In various embodiments, queues 370 - 373 and the attached control units 374 - 377 provide ( Figure 1 of) the functionality of queue 100. In other embodiments, one or more of queues 370 - 373 and the attached control units 374 - 377 provide ( Figure 2 of) the functionality of fabric packet transmitter 200.

[0048] In various embodiments, communication structure 310 (or structure 310) transfers packets between CPU 360, GPU 362, multimedia engine 366, and peripherals 380 and 390. Structure 310 also transfers data between memory 350 and clients such as CPU 360, GPU 362, multimedia engine 366, and peripherals 380 and 390 and other peripherals (not shown). In various embodiments, interfaces 312 - 316 and 330 - 334 and memory controller 340 include hardware circuits for implementing algorithms to provide functionality. Interfaces 312 - 316 and 332 - 334 are used to transfer data, requests, and acknowledgment responses between routing network 320 and CPU 360, GPU 362, multimedia engine 366, and peripherals 380 and 390. One or more of interfaces 312 - 316 and 332 - 334 and control units 374 - 377 include circuits for generating packets, decoding packets, and supporting communication with routing network 320. In some embodiments, interfaces 312 - 316 and 330 - 334 use a communication protocol such as the PCIe (Peripheral Component Interconnect Express) interconnect communication protocol. In other embodiments, another communication protocol is used. In some embodiments, each of interfaces 312 - 316 and 332 - 334 communicates with a single client as shown. In other embodiments, one or more of interfaces 312 - 316 and 332 - 334 communicate with multiple clients and use an identifier that identifies the client to track data with the client.

[0049] Although a single memory controller 340 is shown for memory 350, in other embodiments, computing system 300 includes multiple memory controllers, where each memory controller supports one or more memory channels. Memory controller 340 includes circuitry for grouping requests to be sent to memory 350 and sending the requests to memory 350 based on the timing specifications of memory 350 in support of burst mode. In various embodiments, memories 350-390 include any of a variety of random access memories (RAMs). In some embodiments, memory 350 stores data and corresponding metadata in synchronous RAM (SRAM). In other embodiments, memory 350 stores data and corresponding metadata in one of a variety of dynamic RAMs (DRAMs). For example, memory 350 stores data in traditional DRAM or in multiple three-dimensional (3D) memory dies stacked on top of each other, depending on the embodiment. Although not shown, memory controller 340 or another memory controller provides access to non-volatile memory for storing data at a lower memory hierarchy level than memory 350. Examples of non-volatile memory are hard disk drives (HDDs), solid state drives (SSDs), and the like.

[0050] When processing an application, clients 360-366 and other peripheral devices (not shown) store frequently accessed data in one or more caches of the cache memory subsystem. Processors of clients 360-366 and other peripheral devices use linear (or "virtual") addresses to identify the requested data. Examples of the requested data are user data, final result data, intermediate result data, and instructions. Each software process executed by a processor has a virtual address space. The virtual address space is divided into pages of a specific size. For example, page sizes of 4 kilobytes (4KB) or 64 kilobytes (64KB) are feasible, but other sizes are also contemplated. Virtual pages are mapped to frames of physical memory. The mapping of virtual addresses to physical addresses tracks the location where the virtual page is stored in physical memory, such as page table 352 in memory 350. Although a single page table 352 is shown, in other embodiments, a different number of page tables are stored in memory 350.

[0051] To reduce access to the memory 350, a cache is used to store copies of one or more subsets of the page table 352. For example, the memory controller 340 uses a translation lookaside buffer (TLB) 342 to store the copies. One or more endpoints have access to these copies of subsets of the page table 352, depending on the one or more applications that are running. As shown, at least the hub 364 and the peripheral devices 380 and 390 have this access, and they store copies of subsets of the page table 352 in an address translation cache (ATC) 366, ATC 382, and ATC 392. The processor or other circuitry uses the virtual address of a given memory access request to access the corresponding one of the ATC 366, ATC 382, and ATC 392 to determine whether the corresponding address translation cache stores the associated physical address of the memory location holding the target data.

[0052] When the virtual-to-physical mapping is not found, the processor or other circuitry generates an address translation request to send to the owner of the address translation. In some examples, the memory controller 340 is the owner. In other examples, a peripheral device such as the peripheral device 380 is designated as the owner by the operating system or an application. For example, when an application starts and later ends, one or more of the operating system and the application perform a dynamic reconfiguration of the virtual channel allocation and set the access rights to a specific page table for a particular client. If the peripheral device 390 determines that a miss occurs during access to the ATC 392, the peripheral device 390 generates an address translation request to send to the peripheral device 380 for access to the ATC 382. This address translation access request and its corresponding response are transmitted within a packet through one or more of the queues 370 - 373. Similarly, when a running application is completed and no longer requires address translation, the peripheral device 390 generates an address translation invalidation request to send to each of the hub 364 and the peripheral device 390. This address translation invalidation request and its corresponding response are transmitted within a packet through one or more of the queues 370 - 373. Based on the specific implementation of the queues 370 - 373 and the accompanying control units 374 - 377, the latency for servicing address translation requests is reduced.

[0053] The methods 400 and 500 described below are for the circuitry of a queue. The queue stores requests targeting a shared resource in its entries. Multiple requesters generate requests for accessing the shared resource. In some embodiments, the requesters are clients of a computing system, and the queue is in the path between the requesters and the shared resource. For example, the queue is within the communication fabric of the computing system. In various embodiments, access by the multiple requesters to the shared resource is based on the specifications of an interconnection communication protocol. In some embodiments, the shared resource is a copy of a portion of a shared page table stored in a memory controller or an address translation cache at the other endpoint, and the specification is the address translation services (ATS) specification of the PCIe (Peripheral Component Interconnect Express) interconnection communication protocol. The ATS specification supports remote caching (storage) of address translations at the endpoint. In other embodiments, another communication protocol is supported. When a request traverses from a requester to the shared resource, the request is stored in the queue. In addition to determining when to issue requests from the queue, the circuitry of the queue also controls data storage in the entries of the queue. The circuitry of the queue assigns at least one entry to each virtual channel in a group of virtual channels, where the group includes virtual channels for each address translation request type from a single source among multiple sources. Any of the previously described apparatuses, packet transmitters, queues, and systems can be used to implement the steps of methods 400 - 500. Further description of these steps is provided in the following discussion.

[0054] Now referring to Figure 4 , an embodiment of method 400 for efficiently allocating entries of a queue that stores at least address translation requests is shown. For purposes of discussion, the steps in this embodiment (and Figure 5 ) are shown in sequential order. However, in other embodiments, some steps occur in a different order than shown, some steps are performed simultaneously, some steps are combined with other steps, and some steps do not exist.

[0055] When a particular application starts, one or more of the operating system and the application perform a dynamic reconfiguration of virtual channel allocation and set permissions for a particular source to access a particular page table. One or more sources with permissions are endpoints such as peripheral devices. The source(s) granted permission can generate access requests for a copy of the address translation in the particular page table. The circuitry of the queue assigns at least one entry of the queue to each virtual channel in a set of virtual channels, the set of virtual channels including virtual channels for each address translation request type from a single source among multiple sources (block 402). Here, the virtual channels are assigned to address translation requests from a particular source. In one example, virtual channel (VC) 1 is assigned to address translation requests from a first type of peripheral device, VC 2 is assigned to address translation requests from a second type of peripheral device, VC 3 is assigned to address translation requests from a memory controller, and so on. Although virtual channel identifiers 1 - 3 are used here, in other embodiments, other identifiers are used. Thus, the control circuitry of the queue assigns at least one entry to each virtual channel in a set of virtual channels, where the set includes virtual channels for each address translation request type from a single source among multiple sources.

[0056] When an entry has not been assigned, the circuitry holds one or more unassigned entries of the buffer as available for any of the multiple sources (block 404). The circuitry receives an address translation request from a given virtual channel (block 406). An example of an address translation request is a read request to retrieve a copy of one or more virtual - to - physical address translations from a destination. Another example of an address translation request is an invalidate request to invalidate one or more stored virtual - to - physical address translations stored at a destination. If the circuitry determines that an assigned entry is available for the given virtual channel (the "yes" branch of conditional block 408), the circuitry selects an available assigned entry of the queue (block 410). Subsequently, the circuitry allocates the selected entry with the received request (block 414). Since the control circuitry of the queue assigns at least one entry to each virtual channel in a set of virtual channels, the set of virtual channels including virtual channels for each address translation request type from a single source among multiple sources, an assigned entry is unavailable only when it has been allocated. The arbitration circuitry of the queue examines the assigned entries of the queue. Thus, no virtual channel having access to the shared address translation has a request that would be blocked from arbitration due to there being no available assigned entries.

[0057] If the circuit determines that the allocated entry is not available for a given virtual channel (the "no" branch of conditional block 408), the circuit selects an available unallocated entry of the queue (block 412). Subsequently, the circuit allocates the selected entry using the received request (block 414). In various embodiments, the queue has available unallocated entries because when unallocated entries are not available, the queue sends an indication to other sources or queues. In some embodiments, the queue, as well as other sources and queues, maintain a plurality of credits that indicate how many requests the circuit can receive and how many requests can be sent to each of the other sources and queues in a particular clock cycle.

[0058] Now referring to Figure 5 , an embodiment of a method 500 for efficiently issuing requests for entries stored in a queue that stores at least address translation requests is shown. The circuit of the queue maintains allocated and unallocated entries of the queue that store requests from requesters targeting an address translation cache (block 502). The circuit checks the allocated entries of the queue (block 504). If the circuit determines that none of the allocated entries exceed an epoch threshold (the "no" branch of conditional block 506), the circuit selects an allocated entry for issuance based on arbitration attributes (block 510). Examples of such attributes are the priority level of the received request, quality of service (QoS) parameters, source identifier (ID), application ID or type (such as a real-time application), virtual channel ID, bandwidth requirement or latency tolerance requirement, indication of an epoch, and so on.

[0059] In various embodiments, the control circuit of the queue allocates at least one entry to each virtual channel in a group of virtual channels, where the group includes virtual channels for each address translation request type from a single source among multiple sources. In some embodiments, during each arbitration phase, when the arbitration circuit of the queue selects two or more requests from the allocated entries of the queue for issuance, the arbitration circuit selects at least one request from the allocated entries of the above virtual channel group. The arbitration phase requires one or more clock cycles, depending on the implementation. Thus, the arbitration circuit selects at least one address translation request for issuance from the above virtual channel group during each arbitration phase, but the selected address translation request has an arbitration value lower than that of one or more requests of other virtual channels.

[0060] If the circuit determines that the allocated entries exceed the epoch threshold (the "yes" branch of conditional block 506), then the circuit selects the allocated entries that exceed the epoch threshold for issuance (block 508). If two or more requests stored in the allocated entries have an epoch that exceeds the epoch threshold and the arbitration circuit cannot issue all the requests, then the circuit selects one or more requests based on the attributes as described above. The circuit issues the requests of the selected allocated entries (block 512). For example, the circuit issues the selected requests to an endpoint that includes a shared address translation cache, and one or more intermediate queues may be on the path toward the endpoint.

[0061] If the circuit determines that there are no unallocated entries allocated to the virtual channel for the issued entry (the "no" branch of conditional block 514), then the circuit keeps the selected allocated entry empty (block 516). For example, even if the selected allocated entry is the only available entry in the queue, the circuit does not allocate the selected allocated entry to other virtual channels. If the circuit determines that there are unallocated entries allocated to the virtual channel for the issued entry (the "yes" branch of conditional block 514), then the circuit allocates the selected allocated entry with a request from the unallocated entry of the virtual channel (block 518). Subsequently, the control flow of method 500 returns to block 502, where the circuit maintains the allocated and unallocated entries of the queue.

[0062] Note that one or more of the above-described embodiments include software. In such embodiments, the program instructions that implement the method and / or mechanism are conveyed or stored on a computer-readable medium. Many types of media are available for storing program instructions and include hard disks, floppy disks, CD-ROMs, DVDs, flash memories, programmable ROMs (PROMs), random access memories (RAMs), and various other forms of volatile or non-volatile storage devices. Generally, a computer-accessible storage medium includes any storage medium that can be accessed by a computer during use to provide instructions and / or data to the computer. For example, a computer-accessible storage medium includes storage media such as magnetic or optical media, such as disks (fixed or removable), tapes, CD-ROMs or DVD-ROMs, CD-Rs, CD-RWs, DVD-Rs, DVD-RWs, or Blu-ray, etc. The storage medium also includes volatile or non-volatile storage media, such as RAM (e.g., synchronous dynamic RAM (SDRAM), double data rate (DDR, DDR2, DDR3, etc.) SDRAM, low power DDR (LPDDR2, etc.) SDRAM, Rambus DRAM (RDRAM), static RAM (SRAM), etc.), ROMs accessible via a peripheral device interface (such as a universal serial bus (USB) interface, etc.), flash memories, non-volatile memories (e.g., flash). The storage medium includes microelectromechanical systems (MEMS), as well as storage media accessible via a communication medium such as a network and / or a wireless link.

[0063] Additionally, in various embodiments, the program instructions include a behavioral-level description or a register transfer level (RTL) description of a hardware function in a high-level programming language (such as C) or a design language (HDL) (such as Verilog, VHDL, or a database format (such as the GDS II stream format (GDSII))). In some cases, the description is read by a synthesis tool that synthesizes the description to produce a netlist that includes a list of gates from a synthesis library. The netlist includes a set of gates that also represents the functionality of the hardware that includes the system. The netlist is then placed and routed to produce a data set that describes the geometry to be applied to a mask. The mask is then used in various semiconductor manufacturing steps to produce a semiconductor circuit or circuit corresponding to the system. Alternatively, the instructions on the computer-accessible storage medium are a netlist (with or without a synthesis library) or a data set as desired. Additionally, the instructions are for the purpose of being simulated by a hardware-based type emulator from such vendors as and Mentor for simulation purposes.

[0064] Although the above embodiments have been described in considerable detail, many variations and modifications will become apparent to those skilled in the art once the above disclosure is fully understood. It is intended that the following claims be interpreted to cover all such variations and modifications.

Claims

1. A device, comprising: a plurality of entries, each entry configured to store a request corresponding to one of a first set of virtual channels and a second set of virtual channels different from the first set of virtual channels, wherein each virtual channel of the second set of virtual channels is allocated for communicating address translation requests from a different one of a plurality of requesters; and A circuit, the circuit being configured as: assigning a given entry of the plurality of entries to each virtual channel of the second set of virtual channels; as well as One or more requests are selected from the plurality of entries for publication using at least the assigned selection criteria for each of the assigned given entries. 2 . The apparatus of claim 1 , wherein the circuit is further configured to select at least one request from the allocated given entry each time a request is issued from the plurality of entries.

3. The apparatus of claim 1, wherein the address translation request type is an access request targeting a copy of at least a portion of a shared page table storing address translations.

4. The apparatus of claim 3, wherein the address translation request type is an invalidation request targeting a copy of at least a portion of a shared page table storing address translations.

5. The apparatus of claim 1, wherein the virtual channel represents multiple time division multiplexed channels on a single physical channel.

6. The apparatus of claim 3, wherein based at least in part on determining that an address space including the address translation is redefined, the circuitry is further configured to: redefine the second set of virtual channels; and At least one entry of the plurality of entries is reallocated to each virtual channel in the redefined second group.

7. The apparatus of claim 1 , wherein the circuitry is further configured to maintain an entry in the allocated given entries as de-allocated in response to determining that: The entry in the allocated given entry is de-allocated; and There are no pending requests for virtual channels in the second group assigned to the entry in the assigned given entry.

8. The apparatus of claim 1, wherein the circuit is further configured to: receiving an allocation request indicating that a first request of a first type different from the address translation type is ready for allocation in the plurality of entries; and Sending a response indicating awaiting allocation of the first request is based at least in part on determining: No entry in the plurality of entries allocated to requests of the first type is available for allocation; and No unallocated entry in the plurality of entries is available for allocation.

9. A method comprising: storing, in each of a plurality of entries of a queue, a request from one of a first set of virtual channels and a second set of virtual channels different from the first set of virtual channels, wherein each virtual channel of the second set of virtual channels is allocated for transmitting address translation requests from a different requestor of a plurality of requestors; assigning, by circuitry of the queue, a given entry of the plurality of entries to each virtual channel of the second set of virtual channels; as well as One or more requests are selected from the plurality of entries for issuance by the circuitry of the queue using at least the assigned selection criteria for each of the assigned given entries.

10. The method of claim 9, further comprising selecting at least one request from the allocated given entry each time a request is issued from the plurality of entries.

11. The method of claim 9, wherein the address translation request type is an access request targeting a copy of at least a portion of a shared page table storing address translations.

12. The method of claim 11, wherein a requestor assigned a virtual channel of the second set of virtual channels is one of a plurality of clients and peripheral devices having access to the shared page table.

13. The method of claim 11 , wherein based at least in part on determining that an address space including the address translation is redefined, the method further comprises: redefine the second group of virtual channels; as well as At least one entry of the plurality of entries is reallocated to each virtual channel in the redefined second group.

14. The method of claim 9, further comprising maintaining an entry in the allocated given entries as de-allocated in response to determining: The entry in the allocated given entry is de-allocated; and There are no pending requests for virtual channels in the second group assigned to the entry in the assigned given entry.

15. A computing system, comprising: a plurality of requesters configured to generate requests; A first queue, wherein the first queue comprises: a first plurality of entries, each entry configured to store a request from one of a first set of virtual channels and a second set of virtual channels different from the first set of virtual channels, wherein each virtual channel in the second set of virtual channels is allocated for communicating an address translation request from a different one of a plurality of requesters; and First Circuit; The first circuit is configured as follows: assigning a given entry of the plurality of entries to each virtual channel of the second set of virtual channels; as well as One or more requests are selected from the plurality of entries for publication using at least the assigned selection criteria for each of the assigned given entries. 16 . The computing system of claim 15 , wherein the first circuit is further configured to select at least one request from the allocated given entry each time a request is issued from the plurality of entries.

17. The computing system of claim 15, wherein the address translation request type is an access request targeting a copy of at least a portion of a shared page table storing address translations.

18. The computing system of claim 17, wherein based at least in part on determining that an address space including the address translation is redefined, the circuit is further configured to: redefine the second set of virtual channels; and At least one entry of the plurality of entries is reallocated to each virtual channel in the redefined second group.

19. The computing system of claim 15, wherein the first circuit is further configured to maintain an entry in the allocated given entries as deallocated in response to determining that: The entry in the allocated given entry is de-allocated; and There are no pending requests for virtual channels in the second group assigned to the entry in the assigned given entry.

20. The computing system of claim 15, wherein the computing system further comprises a second queue, the second queue comprising: a second plurality of entries, each entry configured to store a response to one of a third group of virtual channels and a fourth group of virtual channels, wherein each virtual channel in the fourth group is assigned to a response of an address translation type for a single requestor of the plurality of requestors; and Second Circuit; The second circuit is configured as follows: assigning a set of entries from the second plurality of entries to each virtual channel in the fourth set; as well as During each arbitration phase, a selection criterion from the allocated set of allocated entries is utilized.

Citation Information

Patent Citations

  • Address translation and address translation memory for storage class memory

    US10810133B1