Unordered Memory Request Trace Structures and Techniques

By adopting a memory request tracking structure of multiple trace queues in the streaming cache, the head blocking and stagnation caused by a single FIFO is solved, achieving higher parallelism and performance.

CN113986779BActive Publication Date: 2025-05-30NVIDIA CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202110244828.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-07-27
Filing Date
2021-03-05
Publication Date
2025-05-30
Estimated Expiration
2041-05-30

AI Technical Summary

Technical Problem

In the existing streaming cache design, the strict orderly nature of a single FIFO leads to the head blocking of the line and the L1Tag stagnation, limiting the parallelism and efficiency of the system.

Method used

A memory request tracking structure with multiple tracking queues allows the head of each tracking queue to be removed every cycle, dynamically allocating pending requests to avoid head blocking.

Benefits of technology

By allowing some types of requests to be processed out of order, parallelism within the thread bundle and overall system performance are improved, and SM stagnation time and memory request delays are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113986779B_ABST
    Figure CN113986779B_ABST
Patent Text Reader

Abstract

This application relates to unordered memory request tracking structures and techniques. In a streaming cache, multiple dynamically sized tracking queues are employed. Request tracking information is distributed across the multiple tracking queues to selectively enable unordered memory request returns. A dynamically controlled policy assigns outstanding requests to the tracking queues, for example providing ordered memory returns in some cases and / or for some traffic, and unordered memory returns in other cases and / or for other traffic.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology herein relates to streaming cache memories and, more particularly, to a streaming cache memory that includes a memory request tracking structure having a plurality of tracking queues that allow for ordered or unordered tracking of memory requests. Background Art

[0002] The workload of a graphics processing unit (GPU) typically has a very large working set - typically on the order of hundreds of megabytes or more to generate a single image frame. Providing sufficient on-chip cache memory capacity to store such a large working set is impractical. Additionally, high-performance massively parallel GPUs cannot stall while waiting for missing data to become available - they need to be able to push forward and process available data (i.e., cache hits), even if other data (cache misses) is not yet available and is still being retrieved by the memory system.

[0003] A cache architecture that continues with a mix of hits and misses is called a streaming cache architecture. See, e.g., Patterson et al., Computer Organization and Design: The Hardware / Software Interface, Appendix C (Nickolls et al., Graphics and Computing GPUs) (Elsevier 2012). In the past, such streaming caches have been used in a variety of GPU architectures, including NVIDIA's Volta and Turing architectures. See, e.g., Qiu et al.'s USP 10459861. Titled "Unified Cache For Diverse Memory Traffic," which is incorporated herein by reference.

[0004] In such an architecture, the streaming multiprocessor (SM) level 1 (L1) data cache includes a streaming cache that serves as a bandwidth filter from the SM to the memory system. The unified cache subsystem disclosed by Qiu et al. includes a data memory configured to serve as both a shared memory and a local cache memory. To handle memory transactions that are not targeted at the shared memory, the unified cache subsystem includes a tag processing pipeline configured to identify cache hits and cache misses. When the tag processing pipeline identifies a cache miss for a given memory transaction, the transaction is pushed into a first-in, first-out (FIFO) tracking queue until the requested data is returned from the L2 cache or external memory.

[0005] In this design, the first miss that occurs can be resolved before the second miss occurs, the second miss that occurs can be resolved before the third miss occurs, and so on - even though in some cases the memory system (e.g., L2 cache) may resolve a later miss (e.g., via main memory) before resolving an earlier miss. This means that the latency to resolve a particular cache miss becomes the worst-case latency to resolve any cache miss.

[0006] Although traditional streaming cache designs typically only allow a small number of outstanding memory requests, the L1 cache in more advanced high-performance GPUs (e.g., NVIDIA Volta and NVIDIA Turing) is designed to be latency-tolerant and can have many memory requests concurrently in flight (e.g., up to 1024 independent outstanding requests can be issued simultaneously). Allowing a large number of concurrent outstanding memory requests helps avoid SM stalls, as SMs execute many threads massively in parallel and concurrently (e.g., in one embodiment, an SM executes between 32 and 64 independent warps, each warp including 32 concurrent threads), and all of these threads share (and often compete for) the common L1 cache. Since the latency associated with the L1 cache fetching data from the L2 cache or main memory is relatively high, a tracking FIFO is used to hide the latency, and the SM and its processes are designed to anticipate the latency caused by such cache misses and can perform useful work while waiting for the missing data to be retrieved from the memory system. Meanwhile, a memory request tracking structure is used to track all requests in flight and schedule a return to the SM when all data required for an operation has returned to the cache and the request process can continue execution.

[0007] GPUs of the prior art as described above tend to use a single FIFO as a memory request tracking structure, where only the oldest outstanding memory request in each cycle is eligible for processing. The advantages of a single FIFO are simplicity, space and power savings, and reduced statistics for matching requests to returns. Additionally, when filtering textures, designers typically choose to resolve multiple samples into a texture in the same order each time (representing a single thread for a single texture query) to avoid differences due to rounding errors when filtering (merging) multiple samples into a single resulting color. For example, wavefronts in a single texture instruction are typically processed in order to filter consistently. A single FIFO is usually sufficient for such workloads. However, in some cases, a single FIFO can create head-of-line blocking to delay ready requests that are not at the head of the FIFO, thus preventing the system from taking advantage of the parallelism in the design. Once the single FIFO is full, the L1 cache stops sending requests to the memory system, ultimately causing the SM to stall.

[0008] More specifically, the conceptual delay FIFO between the tag level (T) and the data level (D) of L1 is referred to as the "T2D FIFO" (tag-to-data FIFO). In this document, "T2D" will refer to such a tag-to-data FIFO. This conceptual T2D FIFO has a limited length, which was 512 entries for some prior GPU designs, but can be any desired length. In one embodiment, the FIFO is very wide, meaning that increasing it may require a large amount of chip area. The T2D FIFO is designed to cover the average L1 miss latency to prevent the SM from stalling.

[0009] Typically, all entries are pushed to the tail of the T2D FIFO from L1 Tag. In each cycle, the hardware supporting the streaming cache checks the head of the FIFO to see if the data of the head entry is available in L1 Data. If the data is ready, the FIFO pops the head entry from the FIFO and sends it to L1 Data to perform a data read.

[0010] The strictly ordered nature of this FIFO creates two problems:

[0011] Head-of-line blocking: In each cycle, the only available entry popped from the FIFO is the head entry. If there are entries in the middle of the FIFO with data ready, they will wait until all previous entries are popped - then they are eligible for removal only after reaching the FIFO head. This causes head-of-line blocking of the line and increases the latency tolerated by the SM. This means that the observed average latency is close to the worst-case L1 miss latency, and the design does not see the lower latency of operations hitting in the L2 cache. If new requests are to be streamed continuously, this latency can theoretically be hidden, but this does not always happen in practice.

[0012] L1Tag stall when the FIFO is full: As long as there are more entries (e.g., 512 entries) to be processed in T2D than the length of the FIFO, the L1Tag stalls. This prevents all traffic in the L1Tag from advancing, including hits that can normally bypass T2D. This also prevents the memory system from servicing any new memory requests.

[0013] Thus, although a single FIFO with a memory request tracking structure for a streaming cache has the advantage of simplicity and is generally sufficient for many applications (especially traffic that needs to maintain serialization) in a GPU design that tolerates latency, additional complexity may be needed to proceed faster in cases where the constraints of a single FIFO affect efficiency and performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Please refer to the following detailed description of exemplary non-limiting embodiments in conjunction with the accompanying drawings, wherein:

[0015] Figure 1 A comparison between an existing design of a memory request tracking structure for a streaming cache and an example embodiment is schematically shown.

[0016] Figure 2 is a schematic diagram showing the logical architecture of an embodiment of a memory request tracking structure for a streaming cache.

[0017] Figure 3 shows Figure 2 An example physical view of the memory request tracking structure for a streaming cache of the embodiment.

[0018] Figure 4 shows an example non-limiting trace queue mapping pattern.

[0019] Figure 5 shows an example trace structure entry filling in an embodiment.

[0020] Figure 6-22The description of U.S. Patent USP10459861 titled "Unified Cache For Diverse Memory Traffic" by Qiu et al., which is incorporated by reference. Specific embodiments

[0021] The example non - restrictive techniques herein create multiple trace queues and allow the head of each trace queue to be removed each cycle if ready. The GPU checks the heads of all trace queues and frees an entry from the ready trace queues instead of simply checking the T2D head each cycle. Thus, one embodiment replaces or supplements the previous single FIFO ordered trace structure with an unordered trace structure that improves performance at minimal hardware cost and without the need for software changes.

[0022] In one embodiment, this unordered tracing allows certain types of requests (e.g., general - purpose compute load / store and ray - tracing acceleration data structure requests) to be processed out - of - order while having sufficient flexibility to handle other requests (e.g., texture data), thus providing the ability to support both ordered and unordered memory request traffic flows. At the same time, the example embodiments still allow out - of - order memory request traffic flows across warps for any workload such as texture workloads. Out - of - order memory access and the associated execution to the extent tolerated by the system can result in significant performance and efficiency improvements. For example, out - of - order return processing can improve the processing efficiency of certain workload types within a warp that do not require maintaining ordered processing or ordered returns and can support out - of - order returns within the same warp.

[0023] Figure 1Shows a comparison between an existing design that provides a single FIFO (labeled "Today" on the left) and a new design that provides N trace queues (where N > 1) (labeled "Proposal" on the right). As shown on the left, instead of using a single FIFO to track all outstanding memory requests (or all outstanding requests for a particular type of workload), one embodiment provides multiple (N) trace queues, where N is a dynamic size. An example non-limiting embodiment has N (e.g., 48) trace queues, which correspond to the number N (e.g., 48) of warps that an SM can execute simultaneously. However, the number N of trace queues can be any integer greater than 1 and does not have to match the number of warps. Additionally, the number N can change from one design to another, depending on various factors such as the number of parallel warps that can be processed at one time, the mix of memory request traffic processed by the system, the type and number of SM memory interfaces, and the amount of chip area that will be dedicated to this part of the design to trade off between area and performance and other factors. And the number k of the N trace queues allocated to a particular workload can be dynamically allocated on a programming basis as needed to provide ordered execution or any desired degree of unordered processing. For example, some workloads may require strictly ordered execution, while other workloads can execute in any order, and still other workloads may still require ordered execution relative to a particular workgroup (e.g., warp), but can execute unordered relative to other workgroups (warps). One embodiment is flexible enough to dynamically accommodate and enable any or all of these execution models.

[0024] In one embodiment, each of the N trace queues includes a FIFO that stores pointers to a larger structure. This arrangement allows for dynamic partitioning among the N trace queues. One embodiment uses a dynamically controlled policy that assigns outstanding requests to specific trace queues. In one embodiment, the simplest policy assigns all work from the same warp to the same trace queue. This provides ordered memory returns within a warp, while unordered memory returns across warps. Since the returns within a warp are still ordered, there is no need for any software changes to obtain performance advantages. Additionally, since in one embodiment, the requests within a warp are drained in order, no additional accumulator precision or storage is required to ensure arithmetic consistency for filtered texture operations and other requests. Furthermore, the assignment of requests to the trace queues can be dynamic and based on many factors. Some such assignments may result in requests being almost evenly distributed among the N trace queues, while other assignments may result in uneven distribution of work among the trace queues, thus providing flexibility.

[0025] Specifically, in the embodiments herein, any work assigned to a particular trace queue will be processed in order by that trace queue. This functionality can be utilized to provide an ordered service for workloads, such as certain texture mapping processes that expect an ordered return, and thus benefit from such an ordered service. On the other hand, some other types of workloads (e.g., ray tracing bounding volume hierarchy compression small tree memory requests) may not require an ordered service and may benefit from an unordered service. In such cases, unordered accesses can be distributed across N trace queues to reduce the chance that any single long-latency access can block a large number of other accesses and thus stall the ray tracer. See, e.g., USP10580,196.

[0026] In one embodiment, during each cycle, the front entry of each of the N (e.g., 48) trace queues is checked to see if the cached fill data is ready. In one embodiment, this check is performed in parallel for each of the N trace queues. Once it is determined that any one of the heads of the various trace queues is ready, those entries can be removed from the trace queue and sent to the SM, thereby unblocking those corresponding trace queues.

[0027] In one embodiment, a trace structure is utilized to accomplish the ability to check all N (e.g., 48) queues each cycle, which stores the sectors that each queue is waiting on. In one embodiment, this trace structure is updated each cycle when "GNIC fill" (see below) returns data to the L1 cache. In some cycles, multiple trace queues will have ready entries. In one embodiment, a round robin arbiter is used to select the ready trace queue. Once a ready entry is selected, the entry is removed from the trace structure, the request is processed in the cache, and then the data is sent back to the SM.

[0028] Example non-limiting novel aspects of this design include:

[0029] Dynamically sized trace queues: Instead of each trace queue having a fixed capacity, an example design uses three different storage tables that are traversed using a series of linked list pointers, allowing for dynamic capacity across the trace queues. If only one queue is active, it can allocate all of the storage in the trace structure. If all N (e.g., 48) queues are active, each of them can allocate some portion of the trace structure. This scheme allows the maximum number of memory requests to always be in transit, regardless of how many trace queues are active.

[0030] Dynamically Configured Queue Mapping Policies for Different Traffic Volume Categories: One embodiment has a dynamic runtime decision policy that controls the queue mapping policy. This allows for different mapping decisions to be made for different types of memory access requests (e.g., local / global (“L / G”) memory transaction traffic versus texture / surface (“Tex / Surf”) traffic versus tree traversal unit (TTU) traffic or others). In this case, local / global, texture / surface, and TTU traffic are different memory traffic classes with different tracking and ordering requirements. For example, local / global traffic is related to loads in local or global memory; texture / surface traffic is related to accessing data stored in memory (e.g., texture or shared) (commonly used by shaders) for rendering textures and / or surfaces; and TTU traffic involves memory accesses initiated by a hardware-based “tree traversal unit” (TTU) to traverse an acceleration data structure (ADS), such as a bounded volume hierarchy (BVH) for ray tracing. For these different policies, different workloads will see different performance, and the dynamic control allows for runtime optimization. In one embodiment, TTU requests from a single warp are mapped to all tracking queues, thus maximizing the number of unordered returns of TTU traffic that do not have to be serviced in order. In one embodiment, local / global requests from a single warp are mapped to the same tracking queue. Other embodiments may map local / global requests from a single warp across multiple tracking queues.

[0031] Support for Sorting Across Global Events: In one embodiment, the L1 cache also processes events such as texture header / sampling status update packets and data slot reference counter clear token events. Specifically, as described in USP 10459861 in connection with Figures 11A-11D and reference counter clear token events, in one embodiment, if the reference counter 604 saturates relative to the maximum value data memory 430, a copy of the data slot 600 is created to handle future memory transactions and / or the reference counter 604 is locked to the maximum value and does not increase or decrease. When the tag in the tag memory 414 associated with the data slot 600 is subsequently evicted, a token with a pointer to the data slot 600 is pushed onto the eviction FIFO 424 and scheduled for eviction. After dequeueing, the eviction FIFO 424 detects that the data slot 600 is saturated and then queues the token to the t2d FIFO 420 using the pointer to the data slot. When the token reaches the head of the t2d FIFO 420, the reference counter 604 in the data slot 600 is reset to 0.

[0032] These events create a global ordering requirement across all in-flight requests (not just those within a single warp). One embodiment adds ordering requirements to the tracking queues to functionally handle these global ordering requirements with minimal impact on performance.

[0033] Support for ordered and unordered dispatch in the tracking structures: In one embodiment, ordered allocate and deallocate maintain the storage structures for tracking information while requests from these structures are processed unordered. The impact of unordering the allocate and deallocate of these structures can be simulated and quantified, which is characteristic of other embodiments.

[0034] Support for work items of different granularities based on different traffic classes: Different traffic classes may require different granularities of items to be atomically released from the tracking structures. For texture operations that may need to be filtered in the TEX DF pipeline, one embodiment releases all requests from the tracking structures for an entire structure together. Releasing entries from T2D to L1Data at the instruction granularity allows downstream stages (e.g., dstage / fstage) to continue working on instruction granularity units. For other operations such as local / global or TTU instructions, one embodiment may release only a single wavefront (a dispatchable unit created by grouping together multiple execution thread data such as 64 threads) at a time. In one embodiment, a special mechanism is utilized to handle instructions that generate a large number of wavefronts.

[0035] In one embodiment, each of the N tracking queues contains pointers to entries in T2D, where the data for T2D packets is stored in the same way as in previous ordered designs. In one embodiment, the tracking queues are linked lists of pointers, such that each tracking queue can have a dynamic capacity between 1 and N (e.g., 512) entries.

[0036] One embodiment exploits unordering across warps while maintaining all operations within a warp in program order. This simplification avoids software changes such as instruction execution order allocation and still obtains a significant number of possible performance benefits. Other architectures and embodiments may consider unordered execution of warps.

[0037] One embodiment performs ordered T2D dispatch / deallocate, which means that entries can be removed from anywhere within T2D, but their storage cannot be reclaimed until they reach the head of T2D. This simplification is done to reduce design effort at the cost of reduced performance. In other embodiments, other trade-offs are possible.

[0038] Feature microarchitecture design details

[0039] In the context of an example non - limiting embodiment, Figure 2 illustrates the life of a T2D data packet from TagCheck 2002 until it is delivered to L1Data (2004a, 2004b, 2004c). Figure 2 The structure can be used for Figure 9 the "T2D FIFO 420" block of Figure 9 The unified cache 316 processes different types of memory transactions along different data paths. The unified cache 316 processes shared memory transactions along data path 440 with low latency. The unified cache 316 processes non - shared memory transactions via the tag pipeline 410. When a cache hit occurs, the tag pipeline 410 directs non - shared memory transactions along the low - latency data path 442, and when a cache miss occurs, routes the memory transactions through the T2D FIFO 420 and data path 444. The unified cache 316 also processes texture - oriented memory transactions received from the texture pipeline 428. In the case of a cache hit or a cache miss, the tag pipeline 410 directs texture - oriented memory transactions through the T2D FIFO 420 and data path 446.

[0040] In one embodiment, local / global & TTU cache - miss traffic and all texture ("TEX") (cache - hit and cache - miss) traffic pass through T2D 420, where the local / global & TTU traffic and the texture traffic diverge at the drain logic blocks 1018(a), 1018(b) and the associated separate streaming multi - processor ("SM") interfaces. Meanwhile, local / global & TTU cache hits bypass the T2D FIFO and instead pass through a fast path to the SM L1data hit interface.

[0041] Improved Figure 4 The main stages of the T2D FIFO 420 of

[0042] · Tag Check 2002 is the tag referred to in "T2D". This tag check 2002 is similar to the tag check present in a conventional data cache - in response to a memory access, the tag check determines whether the requested data exists in the cache based on the stored tag. If the data exists, the tag check 2002 determines a hit. If the data does not exist, the tag check 2002 declares a miss and initiates a memory access to the L2 cache and the rest of the memory subsystem.

[0043] · For more details about the "fast path" of T2D around local / global and TTU hits / stores, see Qiu et al.'s USP10459861, which can reduce the latency of hits that do not require serialization.

[0044] · Queue mapping 2006: Tag check 2002 sends the miss traffic to queue mapping 2006. Queue mapping 2006 determines which trace queue 2008 the incoming miss requests should be mapped to. A dynamically controlled policy is used to determine this mapping (different policy options can be used depending on the traffic type). The simplest policy is "all misses from warp 0 go into trace queue 0, all misses from warp 1 go into trace queue 1, and so on."

[0045] In one embodiment, all packets sent to a trace queue maintain their order from the time they are inserted into the trace queue until they are pushed into the commit FIFO 442 (see Figure 9 ). Out-of-order occurs between different trace queues. In one embodiment, for all local / global and single-warp texture packets will always be mapped to the same trace queue. In one embodiment, TTU packets are mapped across multiple trace queues in a round-robin fashion.

[0046] · Trace queues ("TQ") 2008(0), 2008(1), … 2008(N = 47): Trace queues 2008 contain a dynamically sized number of packets that have completed tag checking and are waiting to be selected by the checker selector 2009. In one embodiment, each trace queue 2008 constitutes a FIFO that maintains the first-in, first-out order of the miss traffic assigned to it by queue mapping 2006.

[0047] · Checker selector 2009: In one embodiment, checker selector 2009 selects, for each cycle, a trace queue 2008 that has a valid entry and an empty check slot 2010. Then, one T2D entry corresponding to a single wavefront is moved from the trace queue 2008 to the check phase 2010. It fills the check phase 2010 with the dslot and sector that the entry is still waiting for. Other embodiments may move multiple entries in one cycle.

[0048] Check(2010(0), 2010(1), … 2010(N = 47): Is operatively connected to the heads of the trace queues 2008(0) - 2008(N). The check block 2010 stores the status information of the waiting response from the memory system regarding the trace queues 2008. In one embodiment, the "GNIC fill information" information comes from the GNIC interface of the memory system and is returned to the cache, and the GNIC interface provides the response of the memory system to the access request. Figure 2 Shows a Figure 9 more detailed view of 420, and the GNIC fill interface itself (returned from memory) can be at Figure 9is shown as returning the T2D FIFO 420 instead of the M-Stage 426. These responses from the GNIC interface are referred to as "GNIC" fill information", and in one embodiment, one sector is provided per clock. Each time a GNIC fill occurs, the check 2010 at the head of each trace queue 2008 will be updated to reflect that the filled dslot / sector is now valid in the cache and the entry should no longer wait for it. Note that the same GNIC fill response from the storage system may potentially satisfy the heads of multiple trace queues 2008, because in one embodiment, the GNIC fill can indicate that multiple data values including data blocks are now available in the cache, and there may also be multiple trace queues that may be waiting for the same data values. Once an entry in the check phase 2010 is no longer waiting for any sector, it indicates that it is ready to be removed from the check.

[0049] · Pop selector: Each cycle, the pop selector selects a check entry 2010(k) from the pool of ready entries to be removed from the check phase 2010 and inserted into the checked queue 2012. This "pop" results in obtaining a value from the check phase 2010. In one embodiment, N entries are moved from the check phase 2010 to the checked queue 2012 each cycle, where N is any positive integer such that the pop selector selects which one or more entries are to be moved to the checked queue 2012 in the current cycle. Other embodiments may allow multiple or all available entries to be moved from the check phase 2010 to the checked queue 2012 each cycle. The pop selector may use polling or any other more complex selector algorithm to make the selection. Note: In one embodiment, an entry from trace queue 2008(1) is inserted into checked queue 2012(1), an entry from trace queue 2008(2) is inserted into checked queue 2012(2), and so on.

[0050] · Checked queue 2012: The checked queue 2012 is a linked list of T2D packets, where all the data referenced by the packets exists in the cache. In one embodiment, an entry remains in the checked queue until all the entries of that instruction or T2D commit group are inserted into the checked queue 2012 and the interlock between the checked queue and the status packet queue is satisfied.

[0051] · Submission Selector 2014: The submission selector 2014 selects a queue. In each cycle, the submission selector 2014 selects an examined queue having ready instructions or a T2D submission group and moves these entries from the examined queue 2012 to the submission FIFO. In one embodiment, there is a submission selector 2014(a) for local / global and TTU traffic and a submission selector 2014(b) for texture traffic. In one embodiment, the submission selector 2014 ensures that the interlocks between the fast path and the slow path are adhered to. The submission selector 2014 can use polling or any other more complex selector algorithm. Once a queue is selected, the deallocation of the resources used for the de-dispatch of the head entry on that queue is independent of the selector itself.

[0052] · Submission FIFO 2016: The submission FIFO is the final order of the packets and then they are drained to the L1Data and SM memory interfaces through the drain logic 2018. In one embodiment, the values of K examined queues 2012 are moved to the drain logic 2018 in each cycle where K is any positive integer, but other arrangements are also possible.

[0053] · The drain logic 2018 cycles between the submitted values and sends the associated L1 data back to the appropriate SM interface (in this embodiment, there is one drain logic 2018(a) shared between local / global and TTU traffic to send such traffic back to the local / global & TTU interface of the SM, and a second consumption logic 2018(b) for sending texture traffic back to the texture traffic interface of the SM).

[0054] Figure 2 is a logical view. The trace queue 2008 and the examined queue 2012 do not have to be physical FIFOs with a fixed size. Instead, they can include a dynamically sized linked list of pointers to T2D.

[0055] Figure 3 shows an exemplary non - restrictive physical view that includes new pointers to be added to the design and which parts of the control logic operate on which sets of trace pointers. Figure 3 The structures shown in represent different physical storage structures for tracing. Thus, there is a stage for each trace queue list; a state for each wavefront table; and global state information.

[0056] In one embodiment, the state of each trace queue list can store the current trace queue insert pointer; the next - to - put pointer; the check ready flag; the examined queue insert pointer; the next - to - submit FIFO pointer; and the number of instruction packets at the end of the examined queue through the trace queue number.

[0057] In one embodiment, the state of each wavefront table can store the wavefront number; valid bit; identification of the next wavefront in the tracking queue; and a T2D pointer pointing to the previous T2D FIFO. In one embodiment, the table is typically written in order and cannot be recycled until the last wavefront in the table has finished processing a miss. In other embodiments, such a table can be randomly accessed or otherwise accessed out of order, with an associated increase in potential complexity.

[0058] In one embodiment, the global state table is used to track the status information of global ordering events (global across all tracking queues) flowing through the cache. The global state table can include two sub-tables. The first table can include a status packet tracking queue pointer associated with the tracking queue insertion pointer and the next one to be placed in the commit FIFO; and a T2D pointer table providing the T2D head and T2D tail.

[0059] Mapped to the tracking queue

[0060] Mode #1 ( Figure 4 top diagram in) maps all packets to tracking queue 0 and does not provide out-of-order miss tracking at all, thus simulating the existing design. This method can be used for various reasons (e.g., in case of potential errors).

[0061] Mode #2 ( Figure 4 next diagram down in) provides ordered service for local / global and texture traffic to traffic queue 0 and provides out-of-order miss tracking for TTU packets. Again, this mode can be used for various reasons (e.g., in case of potential errors).

[0062] Mode #3 ( Figure 4 third diagram in) can be used as a fallback in case of unexpected interference between local / global and texture traffic and TTU traffic. Mode 3 provides mapping to separate tracking queues based on the warp number, and TTU packets are mapped in turn in the remaining queues.

[0063] Mode #4 ( Figure 4 bottom diagram) is expected to be the most flexible and provide the highest performance. It provides its own local / global and texture packet tracking queues for each warp and maps the TTU traffic cycle across all tracking queues (since TTU traffic has no ordered requirements and thus no ordered service is needed).

[0064] The above four modes are non-limiting examples; many other mappings are possible.

[0065] Inspector Selector 2009

[0066] The checker selector 2009 and the checker pipeline are designed to process the wavefronts (excluding the status packet queue 2020 and the eviction long queue) queued in the trace queue (TQ) 2008 and pass these wavefronts to the check phase 2010. In one embodiment, the checker pipeline (after selection) is itself two pipeline stages deep. Due to the depth of this pipeline, in one embodiment, the checker selector 2009 has no way of knowing when to pick a given wavefront and whether it is safe to pick again from the same TQ 2008 in the next cycle, because the first wavefront may be stalled in the check phase 2010 due to a miss. In one embodiment, this limitation leads to a design with two different and competing tasks: 1) maximizing performance 2) minimizing wasted effort. In this way, the checker selector 2009 design in one embodiment has been implemented as a set of algorithms designed to attempt and balance these tasks. In one embodiment, the checker selector 2009 pipeline can make selector selections based on a combination of multiple different independent criteria.

[0067] After selection, the selected wavefront enters the checker pipeline. In the first cycle of this pipeline (CP2 in RTL), the WAVE and CHKLINK RAMs are read using the wave_id of the wavefront stored in the state of each wavefront table. Figure 3 In one embodiment, the WAVE RAM reads the information (its dslots and half_sector_mask) necessary to determine whether the wavefront is a hit or a miss, and the CHKLINK RAM provides the information necessary to update the TQ check head pointer. In the second cycle of the checker pipeline (CP3 in RTL), the tag library is indexed by the dslot and produces hit / miss information that is combined with the half_sector_masks read from the WAVE ram. This hit / miss information is then passed to the check phase 2010. In addition, the TQ check pointer is updated from CP3 for use by the age arbiter.

[0068] Check Phase 2010

[0069] To provide the ability to check the head entries of all trace queues, we introduce Figure 5The new tracking structure shown. In one embodiment, each tag bank has a tracking structure, and there are multiple (e.g., four) tag banks. The tracking structure stores the dslot ID and the sectors still needed by the headline entry in each tracking queue 2008. In one embodiment, to fill the check phase 2010, it is helpful to know all the dslots needed by the T2D entry and which sectors are needed in other dslots. The overall strategy is: 1) Each request is checked once when passing through the checker selector. At that time, it records which of the required dslot IDs & sectors already exist and which are still missing. From there, it resides in the Figure 5 tracking structure until all the remaining dslot IDs and sectors are ready. For a given request, it can be marked as ready in two ways: either it is ready when passing through "check" (when selected by the checker selector), or it is in the Figure 5 tracking structure until a dslot / sector is seen to be filled by the GNIC. The tracking structure updates the tracking only based on GNIC fills. Thus, in one embodiment, the hardware ensures that the valid bit is updated upon insertion to reflect the current cache state. This may involve adding a bypass to cover 1 or 2 cycles between reading the sector valid state from the cache to filling the tracking structure and starting to receive GNIC fill updates. This ensures that GNIC fills are not lost during the time between querying the cache state and filling the entry.

[0070] On each GNIC fill, the hardware in one embodiment matches in the tracking structure corresponding to the tag bank of the dslot ID and then updates the valid bit for the matching dslot ID. When all 4 sector valid bits for each of the 4 different entries (for different tag banks) are valid, the hardware declares that the tracking queue 2008 is ready to move the entry out of memory and fill the memory with the next entry in that tracking queue. This is Figure 5 described in detail in, which figure shows the "ready" and "not ready" states of different TQs 2008.

[0071] Commit selector 2014 / Comply with instruction boundaries

[0072] For local / global and texture / surface operations, a single instruction can generate multiple T2D entries. In one embodiment, all these entries from a single instruction will stay in the pipeline together. Specifically, in one embodiment, the TEXdstage / fstage wants to act on all the entries of an entire quad before switching to another quad.

[0073] In one embodiment, entries are retained in the inspection queue 2012 until the last entry of the instruction arrives. Once all the entries of the instruction are in the inspection queue 2012, the hardware moves them all to the commit FIFO 2016. There is a separate commit FIFO 2016(a) for local / global & TTU requests, and there is another separate commit FIFO 2016(b) for texture requests. In one embodiment, the commit FIFO 2016 has a small fixed capacity and is used to buffer requests between the L1Tag and L1Data.

[0074] L1Data pops an entry from the commit FIFO and performs a data read. Then, the entry in T2D is marked as invalid.

[0075] Once the invalid entries become the oldest (by reaching the head of the T2D FIFO), their storage is reclaimed and they can be reallocated for new entries. This is called ordered dispatch / dispatch cancellation.

[0076] Status packet

[0077] In one embodiment, the status packet is sent down the texture pipeline and contains information about the texture header / sampler. In one embodiment, some of this status is used in the dstage / fstage. In a previous sequential T2D design, the status packet was sent through the T2D as a packet to ensure ordering relative to the texture instructions that require the status. In one embodiment that provides out-of-order processing, the number of out-of-order (OOO) is limited such that the status packets remain in program order relative to the texture requests that reference them.

[0078] To ensure proper ordering, one embodiment adds an additional tracking queue 2020 ("status packet queue") dedicated to holding status packets. Then, when the hardware determines whether an entry is suitable to move from the inspected queue 2012 to the commit queue 2016, it compares the deadline of the status packet with the deadline of the entry in the inspected queue. In one embodiment, only textures / surfaces must adhere to the ordering relative to the status packets; local / global and TTU may ignore this status packet deadline test. Example non-limiting features in one embodiment are:

[0079] · Local / global and TTU packets do not compare the deadline with the status packet queue

[0080] · Entries in the status packet queue are only eligible to move into the commit queue 2016 if they are earlier than all other entries

[0081] · If the texture / surface operation is earlier than the head of the state packet queue (or the state packet queue is empty), it can be moved to the commit FIFO.

[0082] L1Tag 2002 may also receive state packets between wavefronts belonging to the same instruction. To handle this situation, in one embodiment, the hardware relies on the T2D commit group concept with the following extensions:

[0083] · When a state packet arrives at the L1Tag, if the L1Tag is currently in the middle of processing an instruction, it sets the end of the T2D commit group flag on the last wavefront filled in the "per-wavefront state" wavefront storage table.

[0084] By setting the end of the T2D commit group flag, the system ensures that all wavefronts before the state packet are removed from the T2D before processing the state packet.

[0085] Processing instructions that generate a large number of wavefronts

[0086] Some texture instructions can generate thousands of T2D entries for a single texture instruction. Therefore, if the system waits until the last entry arrives, a deadlock situation may occur. To handle these events, one embodiment introduces an additional bit to indicate that a packet is the last packet in a T2D commit group. When L1Tag 2002 sends a request to the trace queue 2008, it calculates how many packets have been sent for the current instruction. In one embodiment, a programmable dynamic runtime decision value controls the maximum number of wavefronts in a T2D commit group. In one embodiment, the default value of this dynamic runtime decision value can be 32. Each time more than a specific number (e.g., 32) of T2D packets are generated, the last (e.g., the 32nd) packet is marked as "end of the T2D commit group". Once the queue with the T2D commit group end flag set starts to deplete, other queues cannot deplete until the instruction end flag is seen for the selected queue. In one embodiment, the end-T2D commit group flag is used when determining whether an entry can be moved from the inspected queue 2012 to the commit FIFO 2016.

[0087] In one embodiment, an entry is eligible to be moved from the inspected queue 2012 to the commit FIFO 2016 if the following conditions are met:

[0088] · All entries of an instruction are in the inspected queue 2012, which means that the packet with the instruction end flag set has arrived at the inspected queue.

[0089] · The state packet deadline check is satisfied.

[0090] ··All entries in the T2D submission group are in the inspection queue 2012, which means that the packets at the end of the T2D submission group flag have reached the inspection queue. In one embodiment, it is effective to move these packets from the inspection queue 2012 to the submission queue only when the first packet of the T2D submission group is the oldest packet among all 48 trace queues. In one embodiment, this is done by comparing the T2D arrangement of the packet with the T2D tail pointer. (This method is effective because the dispatch / dispatch of T2D is executed in order).

[0091] Interlock between fast path and slow path

[0092] One embodiment guarantees that items inserted into T2D (slow path) after inserting items into the fast path (see Figure 2 ) will always reach L1Data after items from the fast path, i.e., the slow path is always slower than the fast path. This guarantee works in one embodiment to ensure that the correct value is returned. For example, in one embodiment, if the system performs surface store (SUST) and then surface load (SULD) to the same address, SULD needs to return the data written by SUST. Since SUST uses the fast path and SULD uses the slow path, if the slow path is faster than the fast path, SULD may read stale data instead of the data from SUST.

[0093] Previously, the guarantee was made through the interlock between the slow path and the fast path. When pushing an entry on the fast path, it takes a snapshot of the T2D tail pointer and the CBUF epoch. Then the following rules are followed:

[0094] · Always be able to process entries in the fast path

[0095] · Only when its T2D pointer is different from the captured T2D pointer stored in the head entry in the fast path can an entry be popped from the slow path.

[0096] If the above two rules are followed, the fast path will always be faster than the slow path of the ordered T2D.

[0097] Assuming ordered dispatch / dispatch, the above scheme can be extended with an unordered T2D. When pushing an entry on the fast path, it records the T2D tail pointer (which is the tail pointer for ordered dispatch / dispatch). Then, the following rules can be followed in one embodiment:

[0098] · Always be able to process entries in the fast path

[0099] · An entry in the slow path can only be processed if its T2D pointer is earlier than the captured T2D pointer stored in the starting entry in the fast path.

[0100] In some embodiments, the above rules are sufficient to ensure that the fast path is always faster than the slow path with unordered T2Ds.

[0101] One embodiment provides the following two possible options to implement the above comparison test:

[0102] · Option 1: If an instruction or T2D submission group passes the comparison between the T2D pointer and the captured pointer in the fast path, the commit selector 2014 is only suitable for selecting the check queue 2012. This will use one comparator per checked queue, but with the highest performance.

[0103] · Option 2: The commit selector 2014 selects the checked queue 2012 that meets the following three conditions, which specify when an entry can be moved from the checked queue to the commit queue 2016, and then the T2D pointer is compared with the T2D pointer captured from the fast path. This will only require one comparator. If the comparison fails, no entries will be moved to the commit queue for that cycle, and in the next cycle, the commit selector will select another checked queue. In one embodiment, it is not sufficient for the commit selector 2014 to stall until the comparison with the T2D pointer captured from the fast path becomes valid.

[0104] Reference counter clear token

[0105] In one embodiment, each dslot contains a reference counter that counts the number of in-flight references to that dslot in the T2D. This reference counter is incremented when a wavefront is pushed into the T2D. In one embodiment, references in the fast path do not manipulate these counters. When a dslot is read in L1Data, the reference counter is decremented. In one embodiment, only dslots with a counter equal to 0 are eligible to be reallocated to another tag. The width of the reference counter determines the number of in-flight requests that can exist for a single dslot. When there are more in-flight references than this number, the reference counter will saturate and hold the maximum value. In one embodiment, a saturated reference count cannot be decreased. When the tag of a dslot with a saturated reference counter becomes invalid in L1Tag, a special reference count (refcount) flush token ("RCFT") is pushed into the T2D. When this token reaches the head of the T2D, it can be ensured that there are no longer any in-flight references to that data line, and the dslot can be reallocated to other tags.

[0106] Since RCFT relies on the fact that T2D is ordered, RCFT may cause problems with T2D. The scheme for status packets can be extended to handle RCFT. In one embodiment, a similar mechanism for RFCT is used compared to the status packet, which has slightly different ordering requirements specific to the ordering requirements of RFCT. Using this scheme, the performance impact of ordering requests on complying with the RFCT ordering requirements can be negligible.

[0107] In one embodiment, the techniques herein improve GPU-level performance in a set of workloads deemed important. The achievable performance improvement depends on the nature of the workload - some individual "buckets" (processing groups, such as graphics processing) see higher performance improvements. The execution system also becomes more efficient because threads or warps that have satisfied memory requests can access the streaming cache and continue execution without having to wait for another thread or warp at the head of the queue.

[0108] All patents and publications cited above are hereby incorporated by reference.

[0109] Although the invention has been described in connection with the presently considered most practical and preferred embodiments, it is to be understood that the invention is not limited to the disclosed embodiments. On the contrary, it is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims.

Claims

1. A memory request tracking circuit for use with a streaming cache memory, the streaming cache memory being configured to receive memory requests for data in a memory system and to return memory system data in response to the received requests, the memory request tracking circuit comprises: a tag check configured to detect a miss in the streaming cache memory; a plurality of tracking queues, each tracking queue being configured to hold miss traffic in a first-in-first-out order; and a queue mapper coupled to the tag check and the plurality of tracking queues, the queue mapper being configured to provide a plurality of memory request tracking information entries for miss traffic to the plurality of tracking queues to enable ordered and unordered memory request returns, the queue mapper being configured to allocate a first subset of the plurality of memory request tracking information entries to the same tracking queue to enable ordered memory request return for the first subset, and to allocate a second subset of the plurality of memory request tracking information entries across the plurality of tracking queues to enable unordered memory request return for the second subset, wherein the queue mapper is configured to allocate the unordered memory request returns across the plurality of tracking queues to reduce the chance that any single long-latency access will block several other accesses, thereby enabling a consumer ray tracer to advance when a single long-latency access occurs.

2. The memory request tracking circuit according to claim 1, wherein, the queue mapper is programmable to: maintain ordered memory request return processing for a first type of memory request, and enable unordered memory request return processing for a second type of memory request, the second type of memory request being different from the first type of memory request.

3. The memory request tracking circuit according to claim 2, wherein, the first type of memory request and the second type of memory request are selected from the group consisting of loads from local or global memory, texture memory / storage, and accelerated data structure storage.

4. The memory request tracking circuit according to claim 1, wherein, the plurality of tracking queues includes first to Nth tracking queues, and the queue mapper assigns the first tracking queue to a particular warp, and evenly distributes the tracking information for unordered memory requests among the second to Nth tracking queues.

5. The memory request tracking circuit according to claim 1, wherein, each of the plurality of tracking queues includes first-in-first-out storage.

6. The memory request tracking circuit according to claim 1, further comprising a pipelined checker selector that, in response to a cache miss fill indication, selects a tracking queue output for application to a check, the pipelined checker selector being dynamically configured to perform the selection.

7. The memory request tracking circuit according to claim 6, wherein, The checking includes a tracking structure that indicates the memory system return data required for the head of each tracking queue and is configured to track when the memory system has returned all sectors required for the head of each tracking queue.

8. The memory request tracking circuit according to claim 7, wherein, the checking is configured to provide a plurality of sector valid bits for each of a plurality of tag libraries.

9. The memory request tracking circuit according to claim 1, further comprising: A first commit selector configured to handle a first traffic type.

10. The memory request tracking circuit according to claim 9, further comprising a second commit selector configured to handle a second traffic type different from the first traffic type.

11. The memory request tracking circuit according to claim 1, configured to provide a first mode and a second mode, the first mode providing an ordered dispatch and undispatch of tracking resources, and the second mode providing an unordered dispatch and undispatch of tracking resources.

12. The memory request tracking circuit according to claim 1, further comprising a status packet queue configured to hold status packets and further configured to limit unordered status packet processing so that the status packets maintain program order with respect to memory access requests referencing the status packets.

13. The memory request tracking circuit according to claim 1, further comprising an indicator encoder that encodes an indicator of the last packet in a commit group.

14. The memory request tracking circuit according to claim 13, further comprising a commit selector configured to select a tracked memory request entry to move to a commit queue in response to the indicator indicating that all memory request entries in the commit group have been serviced and all memory request entries are in the cache.

15. The memory request tracking circuit according to claim 1, further comprising: A first traffic path and a second traffic path, where the first traffic path bypasses the plurality of tracking queues, and the second traffic path passes through at least one of the plurality of tracking queues; and An interlock circuit that ensures that the first traffic path is always faster than the second traffic path.

16. The memory request tracking circuit according to claim 15, wherein, the interlock circuit includes a comparator configured to compare the traffic epochs in the fast path with the traffic epochs in the slow path.

17. The memory request tracking circuit according to claim 1, further comprising a reference counter that counts the number of active references to memory slots, and the output and / or operation of the reference counter is configured to accommodate unordered processing.

18. A streaming cache configured to receive memory requests for data in a memory system and return memory system data in response to the received requests, the streaming cache comprising: A tag pipeline including a merger, a tag memory, and a tag processor, the tag pipeline detecting cache misses; A commit FIFO; An eviction FIFO; An out-of-order memory request tracking circuit includes a plurality of tracking queues, each tracking queue configured to hold unhit traffic in a first-in-first-out order, the out-of-order memory request tracking circuit configured to queue and assign memory request tracking entries for out-of-order memory requests to different tracking queues, and queue and assign memory request tracking entries for in-order memory requests to the same tracking queue; and a pipelined selector circuit coupled to the plurality of tracking queues, the pipelined selector circuit configured to select among the plurality of tracking queues having ready requests and enqueue the selected ready ones into another first-in-first-out queue for delivery to a requester.

19. The streaming cache according to claim 18, further comprising a queue mapper coupled to the tag pipeline and the plurality of tracking queues, the queue mapper configured to distribute a first subset of a plurality of memory request tracking information entries across the plurality of tracking queues to selectively enable out-of-order memory request returns for the first subset, and enqueue a second subset of the plurality of memory request tracking information entries onto the same tracking queue to enable in-order memory request returns for the second subset, wherein each of the plurality of tracking queues is capable of holding any one of the plurality of memory request tracking information entries.

20. A memory request tracking method for use with a streaming cache memory configured to receive memory requests for data in a memory system and return memory system data in response to the received requests, the memory request tracking method comprises: detecting a cache miss; providing a plurality of tracking queues, each tracking queue configured to hold unhit traffic in a first-in-first-out order; and assigning request tracking information entries for unhit traffic to the plurality of tracking queues to support in-order and out-of-order memory request returns, including queuing and assigning out-of-order memory requests across different tracking queues and queuing and assigning in-order memory requests to the same tracking queue, wherein the assignment includes: distributing the out-of-order memory request returns across the plurality of tracking queues to reduce the chance that any single long-latency access will block several other accesses, thereby enabling a consumer ray tracer to advance while a single long-latency access occurs.

21. The memory request tracking method according to claim 20, further comprises: programming a queue mapper to maintain in-order memory request return processing for a first type of memory request and enable out-of-order memory request return processing for a second type of memory request different from the first type of memory request.

22. The memory request tracking method according to claim 21, wherein the first type of memory request and the second type of memory request are each selected from the group consisting of local or global memory load requests, texture memory requests, surface memory requests, and accelerated data structure memory requests.

23. The memory request tracking method according to claim 20, wherein, the plurality of tracking queues includes a first to an Nth tracking queue, and the allocation includes: dispatching the first tracking queue to a specific warp; and evenly distributing the tracking information for out-of-order memory requests among the second to Nth tracking queues.

24. The memory request tracking method according to claim 20, wherein, each of the plurality of tracking queues includes a first-in-first-out storage.

25. The memory request tracking method according to claim 20, further including: in response to a cache miss fill indication, selecting a tracking queue output for application to the check.

26. The memory request tracking method according to claim 25, further including: using a tracking structure to indicate the memory system return data required for the head of each tracking queue, and tracking when the memory system has returned all sectors required for the head of each tracking queue.

27. The memory request tracking method according to claim 26, further including: providing a plurality of sector valid bits for each of a plurality of tag libraries.

28. The memory request tracking method according to claim 20, further including: configuring a first commit selector to handle a first traffic type.

29. The memory request tracking method according to claim 28, further including: configuring a second commit selector to handle a second traffic type different from the first traffic type.

30. The memory request tracking method according to claim 20, further including: providing both ordered and out-of-order dispatch and undispatch of tracking resources.

31. The memory request tracking method according to claim 20, further including: configuring a status packet queue to hold status packets; and restricting out-of-order status packet processing so that the status packets maintain program order with respect to the memory access requests referencing the status packets.

32. The memory request tracking method according to claim 20, further including: encoding an indicator for the last packet in a commit group.

33. The memory request tracking method according to claim 32, further including: in response to the indicator indicating that all memory request entries in the commit group have been serviced and all memory request entries are in the cache, selecting the tracked memory request entries to move to the commit queue.

34. The memory request tracking method according to claim 20, further including: using a first traffic path to bypass the plurality of tracking queues; using a second traffic path through at least one of the plurality of tracking queues; and ensuring that the first traffic path is always faster than the second traffic path.

35. The memory request tracking method according to claim 34, further including: comparing the traffic periods in the fast path and the slow path.

36. The memory request tracking method according to claim 20, further including: Count the number of active references to memory slots, and configure the output and / or operation of the reference counter to accommodate out-of-order processing.

Citation Information

Patent Citations

  • Command ordering based on dependencies

    US20040210679A1

  • Data Rebuild on Feedback from a Queue in a Non-Volatile Solid-State Storage

    US20160041868A1

  • Unified cache for diverse memory traffic

    US20180322078A1