Hybrid allocation of data lines in a stream cache memory
Patent Information
- Application Number
- CN202211152088.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-05-04
- Filing Date
- 2022-09-21
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2042-09-21
AI Technical Summary
结果,高速缓存存储器的利用率为50%,留下一半的高速缓存存储器未使用和不可用
[0009] At least one technical advantage of the disclosed technology over the prior art is that, using the disclosed technology, the cache memory can share a cache line with two or more allocations, where each allocation includes fewer sectors than the entire cache line. As a result, the cache memory can have fewer unused sectors compared to prior art that does not employ cache line sharing. This improves cache memory utilization, leading to improved cache memory performance and faster execution of software applications. These advantages represent one or more technical improvements over prior art methods.
Smart Images

Figure CN117056247B_ABST
Abstract
Description
Technical Field
[0001] The various embodiments generally relate to computer memory architecture, and more specifically, to the hybrid allocation of data rows in a streaming cache memory. Background Technology
[0002] Among other things, a computing system typically includes one or more processing units, such as a central processing unit (CPU) and / or a graphics processing unit (GPU), and one or more memory systems. The processing units execute user-mode software applications that submit and initiate computational tasks, which are executed on one or more computing engines included within the processing unit. The processing units include multi-tiered memory systems to improve performance when loading data from and storing data in memory.
[0003] A multi-tiered memory system includes a relatively large, lower-performance system memory for storing the large number of program instructions included in user-mode software applications, as well as data accessed by the user-mode software applications over time during execution. Additionally, a multi-tiered memory system includes a relatively small, higher-performance cache memory for storing those program instructions and data that the user-mode software applications are currently accessing or about to access. Typically, the cache memory can be organized as a set of cache lines, where each cache line contains tens or hundreds of bytes of data. When data is loaded into the cache, the cache controller allocates one or more cache lines, then loads the data from system memory and stores it in the cache lines. The cache controller loads instructions and data from system memory into the cache memory as needed or just before use. As a result, the processing unit can load instructions and data more frequently from the high-performance cache memory used for instructions and data compared to the lower-performance system memory. The processing unit can also store data into a higher-performance cache. The cache controller ultimately stores these cache lines into the lower-performance system memory. The processing unit thus achieves improved memory performance compared to a non-tiered memory system with only system memory.
[0004] Typically, the available memory transfer bandwidth between system memory and cache memory is limited. Cache performance can be improved by reducing the data transfer traffic between system memory and cache memory. One technique for reducing this data transfer traffic is to divide cache lines into sectors and then load and store only the necessary sectors, rather than loading the entire cache line. For example, if a cache line has four sectors, the cache controller can load one, two, or three sectors as needed, instead of the entire cache line. Therefore, the cache controller does not consume memory transfer bandwidth that would otherwise be used to load sectors not needed by the software application.
[0005] One problem with this technique for reducing memory transfer bandwidth consumption is that cache memory utilization decreases when fewer than a full cache line is loaded. For example, if, on average, each cache line loads data into only two of the four available sectors, then two sectors in each cache line are unused. The unused sectors on the cache line cannot be reallocated or used for other purposes without first evicting those two used sectors. As a result, cache memory utilization is 50%, leaving half of the cache memory unused and unavailable.
[0006] As mentioned above, there is a need in the art for more efficient techniques for managing cache memories in computing systems. Summary of the Invention
[0007] Various embodiments of this disclosure illustrate a computer-implemented method for managing cache memory in a computing system. The method includes detecting a first cache line allocation request to allocate a first logical sector. The method further includes determining that the first cache line allocation request can be combined with a second cache line allocation request to allocate a second logical sector. The method also includes storing first data associated with the first logical sector in a first physical sector of a first cache line of the cache memory. Second data associated with the second logical sector is stored in a second physical sector of the cache line.
[0008] Other embodiments include, but are not limited to, systems that implement one or more aspects of the disclosed technology, one or more computer-readable media including instructions for performing one or more aspects of the disclosed technology, and methods for performing one or more aspects of the disclosed technology.
[0009] At least one technical advantage of the disclosed technology over the prior art is that, using the disclosed technology, the cache memory can share a cache line with two or more allocations, where each allocation includes fewer sectors than the entire cache line. As a result, the cache memory can have fewer unused sectors compared to prior art that does not employ cache line sharing. This improves cache memory utilization, leading to improved cache memory performance and faster execution of software applications. These advantages represent one or more technical improvements over prior art methods. Attached Figure Description
[0010] To gain a more detailed understanding of the features of the various embodiments described above, the briefly summarized inventive concept can be described in more specific terms with reference to various embodiments (some of which are illustrated in the accompanying drawings). However, it should be noted that the accompanying drawings illustrate only typical embodiments of the inventive concept and are therefore not intended to limit the scope in any way; other equally effective embodiments exist.
[0011] Figure 1 It is a block diagram of a computer system configured to implement one or more aspects of the various embodiments;
[0012] Figure 2 Included according to various embodiments Figure 1 A block diagram of the parallel processing unit (PPU) in the accelerator processing subsystem;
[0013] Figure 3 Included according to various embodiments Figure 2 A block diagram of the General Processing Cluster (GPC) in a Parallel Processing Unit (PPU);
[0014] Figure 4 Included according to various embodiments Figure 1 CPU and / or Figure 2 A block diagram of the cache memory system in the PPU;
[0015] Figure 5 It is based on various embodiments with a one-to-one cache line mapping Figure 4 A block diagram of the cache memory and cache tag memory;
[0016] Figure 6 It is based on various embodiments and has a flexible cache line mapping. Figure 4 A block diagram of the cache memory and cache tag memory;
[0017] Figure 7 It is according to various embodiments having unused sectors Figure 4 A block diagram of the cache memory and cache tag memory;
[0018] Figure 8 It is based on various embodiments with cache line sharing Figure 4 A block diagram of the cache memory and cache tag memory;
[0019] Figure 9 It is based on other embodiments with cache line sharing. Figure 4 A block diagram of the cache memory and cache tag memory; and
[0020] Figure 10It is according to various embodiments for managing such as Figure 1 CPU or Figure 2 A flowchart of the method steps for cache memory of processing units such as PPUs. Detailed Implementation
[0021] In the following description, numerous specific details are set forth in order to provide a more thorough understanding of the various embodiments. However, it will be apparent to those skilled in the art that the inventive concept can be practiced without one or more of these specific details.
[0022] System Overview
[0023] Figure 1 A block diagram is provided to illustrate a computer system 100 configured to implement one or more aspects of various embodiments. As shown, the computer system 100 includes, but is not limited to, a central processing unit (CPU) 102 and a system memory 104 coupled to an accelerator processing subsystem 112 via a memory bridge 105 and a communication path 113. The memory bridge 105 is further coupled to an I / O (input / output) bridge 107 via a communication path 106, and the I / O bridge 107 is in turn coupled to a switch 116.
[0024] In operation, I / O bridge 107 is configured to receive user input from input device 108 (such as a keyboard or mouse) and forward the input to CPU 102 for processing via communication path 106 and memory bridge 105. In some examples, input device 108 is used to authenticate one or more users to allow authorized users to access computer system 100 and deny unauthorized users access. Switch 116 is configured to provide connectivity between I / O bridge 107 and other components of computer system 100, such as network adapter 118 and various add-on cards 120 and 121. In some examples, network adapter 118 is used as a primary or dedicated input device to receive input data for processing using the disclosed techniques.
[0025] As also shown in the figure, I / O bridge 107 is coupled to system disk 114, which can be configured to store content, applications, and data for use by CPU 102 and accelerator processing subsystem 112. Generally, system disk 114 provides non-transitory storage for applications and data and may include fixed or removable hard disk drives, flash memory devices, and CD-ROMs (optical disc read-only memory), DVD-ROMs (digital versatile discs), Blu-ray, HD-DVDs (high-definition DVDs), or other magnetic, optical, or solid-state storage devices. Finally, although not explicitly shown, other components (such as universal serial bus or other port connections, optical disc drives, digital versatile disc drives, film recording devices, etc.) may also be connected to I / O bridge 107.
[0026] In various embodiments, memory bridge 105 may be a northbridge chip, and I / O bridge 107 may be a southbridge chip. Furthermore, communication paths 106 and 113, as well as other communication paths, can be implemented within computer system 100 using any technically suitable protocol (including but not limited to PCIe, HyperTransport, or any other bus or point-to-point communication protocol known in the art).
[0027] In some embodiments, the accelerator processing subsystem 112 includes a graphics subsystem that supplies pixels to a display device 110, which can be any conventional cathode ray tube, liquid crystal display, light-emitting diode display, etc. In this embodiment, the parallel processing subsystem 112 incorporates circuitry optimized for graphics and video processing, including, for example, video output circuitry. (The following is a continuation of the previous paragraph.) Figure 2 As described in more detail below, such circuitry can be combined across one or more accelerators included in the accelerator processing subsystem 112. An accelerator includes any or more processing units capable of executing instructions, such as a central processing unit (CPU). Figure 2-4 Parallel processing units (PPU), graphics processing units (GPU), intelligent processing units (IPU), neural processing units (NAU), tensor processing units (TPU), neural network processors (NNP), data processing units (DPU), vision processing units (VPU), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), etc.
[0028] In some embodiments, the accelerator processing subsystem 112 includes two processors, referred to herein as a main processor (typically a CPU) and an auxiliary processor. Typically, the main processor is a CPU and the auxiliary processor is a GPU. Additionally or alternatively, each of the main processor and auxiliary processor may be any one or more types of accelerators disclosed herein, in any technically feasible combination. The auxiliary processor receives secure commands from the main processor via an insecure communication path. The auxiliary processor accesses memory and / or other storage systems, such as system memory 104, a compute fast link (CXL) memory expander, memory-managed disk storage, on-chip memory, etc. The auxiliary processor accesses this memory and / or other storage system via an insecure connection. The main processor and auxiliary processor may communicate with each other via a GPU-to-GPU communication channel, such as NvidiaLink (NVLink). Furthermore, the main processor and auxiliary processor may communicate with each other via a network adapter 118. Typically, the distinction between insecure and secure communication paths depends on the application. Specific applications typically consider communication on the die or within the package to be secure. Communication of unencrypted data via standard communication channels (e.g., PCIe) is considered insecure.
[0029] In some embodiments, the accelerator processing subsystem 112 incorporates circuitry optimized for general and / or computational processing. Similarly, such circuitry may be incorporated across one or more accelerators included in the accelerator processing subsystem 112, which are configured to perform such general and / or computational operations. In other embodiments, one or more accelerators included in the accelerator processing subsystem 112 may be configured to perform graphics processing, general processing, and computational processing operations. The system memory 104 includes at least one device driver 103 configured to manage the processing operations of one or more accelerators in the accelerator processing subsystem 112.
[0030] In various embodiments, the accelerator processing subsystem 112 may be connected to Figure 1 One or more other elements can be integrated to form a single system. For example, the accelerator processing subsystem 112 can be integrated with the CPU 102 and other interconnect circuitry on a single chip to form a system-on-a-chip (SoC).
[0031] It should be understood that the system illustrated herein is illustrative, and variations and modifications are possible. The connection topology (including the number and arrangement of bridges, the number of CPUs 102, and the number of accelerator processing subsystems 112) can be modified as needed. For example, in some embodiments, system memory 104 may be directly connected to CPU 102 instead of being connected to CPU 102 via memory bridge 105, and other devices will communicate with system memory 104 via memory bridge 105 and CPU 102. In other alternative topologies, accelerator processing subsystems 112 may be connected to I / O bridge 107 or directly to CPU 102 instead of being connected to memory bridge 105. In other embodiments, I / O bridge 107 and memory bridge 105 may be integrated into a single chip rather than existing as one or more discrete devices. Finally, in some embodiments, Figure 1 One or more of the components shown may be absent. For example, it is possible to eliminate the direct connection of switch 116, network adapter 118, and add-on cards 120, 121 to I / O bridge 107.
[0032] Figure 2 According to various embodiments Figure 1 A block diagram of the parallel processing unit (PPU) 202 included in the accelerator processing subsystem 112. Although Figure 2 One PPU 202 is described above, but the accelerator processing subsystem 112 may include any number of PPUs 202. Furthermore, Figure 2 PPU 202 is included Figure 1 An example of an accelerator in the accelerator processing system 112. Alternative accelerators include, but are not limited to, CPUs, GPUs, IPUs, NPUs, TPUs, NNPs, DPUs, VPUs, ASICs, FPGAs, and / or similar devices. Figure 2-4 The technology disclosed in the document regarding PPU 202 is equally applicable, in any combination, to any type of accelerator included within accelerator processing subsystem 112. As shown, PPU 202 is coupled to local parallel processing (PP) memory 204. PPU 202 and PP memory 204 may be implemented using one or more integrated circuit devices, such as programmable processors, application-specific integrated circuits (ASICs), or memory devices, or any other technically feasible method.
[0033] In some embodiments, PPU 202 includes a graphics processing unit (GPU) configured to implement a graphics rendering pipeline to perform various operations related to generating pixel data based on graphics data provided by CPU 102 and / or system memory 104. When processing graphics data, PPU 204 can be used as graphics memory to store one or more regular frame buffers (and one or more other rendering targets if needed). Among other things, PPU 204 can be used to store and update pixel data and transmit the final pixel data or display frame to display device 110 for display. In some embodiments, PPU 202 can also be configured for general processing and computational operations.
[0034] In operation, CPU 102 is the main processor of computer system 100, controlling and coordinating the operation of other system components. Specifically, CPU 102 issues commands to control the operation of PPU 202. In some embodiments, CPU 102 writes the command stream for PPU 202 into a data structure (...). Figure 1 or Figure 2 (Not explicitly shown) This data structure may reside in system memory 104, PP memory 204, or another storage location accessible to both CPU 102 and PPU 202. Alternatively, a processor and / or accelerator other than CPU 102 may write one or more command streams for PPU 202 into the data structure. A pointer to the data structure is written to a push buffer to initiate processing of the command streams within the data structure. PPU 202 reads the command streams from the push buffers and then executes the commands asynchronously relative to the operation of CPU 102. In embodiments that generate multiple push buffers, the application may specify an execution priority for each push buffer via device driver 103 to control the scheduling of different push buffers.
[0035] As also shown in the figure, PPU 202 includes an I / O (input / output) unit 205 that communicates with the rest of computer system 100 via communication path 113 and memory bridge 105. I / O unit 205 generates data packets (or other signals) for transmission on communication path 113 and also receives all incoming data packets (or other signals) from communication path 113, directing the incoming data packets to the appropriate components of PPU 202. For example, commands related to processing tasks may be directed to host interface 206, while commands related to memory operations (e.g., reading from or writing to PP memory 204) may be directed to crossbar switch unit 210. Host interface 206 reads each push buffer and sends the command stream stored in the push buffer to front end 212.
[0036] As described above Figure 1The connection between PPU 202 and the rest of computer system 100 can vary. In some embodiments, accelerator processing subsystem 112 (which includes at least one PPU 202) is implemented as an add-in card that can be inserted into an expansion slot of computer system 100. In other embodiments, PPU 202 may be integrated on a single chip using a bus bridge, such as memory bridge 105 or I / O bridge 107. Similarly, in other embodiments, some or all of the components of PPU 202 may be included together with CPU 102 in a single integrated circuit or system-on-a-chip (SoC).
[0037] In operation, front-end 212 sends processing tasks received from host interface 206 to a work allocation unit (not shown) within task / work unit 207. The work allocation unit receives pointers to processing tasks, which are encoded as Task Metadata (TMDs) and stored in memory. Pointers to TMDs are included in a command stream, stored as a push buffer, and received by front-end 212 from host interface 206. Processing tasks, which can be encoded as TMDs, include an index associated with the data to be processed, as well as state parameters and commands defining how the data should be processed. For example, state parameters and commands can define a program to be executed on the data. Task / work unit 207 receives tasks from front-end 212 and ensures that GPC 208 is configured to an active state before initiating the processing task specified by each TMD. A priority can be assigned to each TMD, which is used to schedule the execution of processing tasks. Processing tasks can also be received from processing cluster array 230. Optionally, the TMD may include parameters that control whether the TMD is added to the head or tail of the list of processing tasks (or a list of pointers to processing tasks), thereby providing another level of control over execution priority.
[0038] The PPU 202 leverages a highly parallel processing architecture based on a processing cluster array 230, which comprises a set of C general-purpose processing clusters (GPCs) 208, where C ≥ 1. Each GPC 208 is capable of executing a large number (e.g., hundreds or thousands) of threads simultaneously, where each thread is an instance of a program. In various applications, different GPCs 208 can be allocated to handle different types of programs or perform different types of computations. The allocation of GPCs 208 can vary depending on the workload generated by each type of program or computation.
[0039] The memory interface 214 includes a set of D partition units 215, where D ≥ 1. Each partition unit 215 is coupled to one or more dynamic random access memories (DRAMs) 220 residing within the PP memory 204. In one embodiment, the number of partition units 215 is equal to the number of DRAMs 220, and each partition unit 215 is coupled to a different DRAM 220. In other embodiments, the number of partition units 215 may differ from the number of DRAMs 220. Those skilled in the art will recognize that the DRAMs 220 can be replaced with any other technically suitable storage device. In operation, various rendering targets (such as texture maps and framebuffers) can be stored across the DRAMs 220, allowing the partition units 215 to write portions of each rendering target in parallel, thereby efficiently utilizing the available bandwidth of the PP memory 204.
[0040] A given GPC 208 can process data to be written to any DRAM 220 in PP memory 204. A crossbar switch unit 210 is configured to route the output of each GPC 208 to the input of any partition unit 215 or any other GPC 208 for further processing. GPCs 208 communicate with memory interface 214 via crossbar switch unit 210 to read from or write to the respective DRAMs 220. In one embodiment, crossbar switch unit 210 is connected to I / O unit 205 in addition to being connected to PP memory 204 via memory interface 214, thereby enabling processing cores in different GPCs 208 to communicate with system memory 104 or other memory not native to PPU 202. Figure 2 In some embodiments, the crossbar switch unit 210 is directly connected to the I / O unit 205. In various embodiments, the crossbar switch unit 210 may use a virtual channel to separate the service flow between the GPC 208 and the partition unit 215.
[0041] Similarly, GPC 208 can be programmed to perform various application-related processing tasks, including but not limited to linear and nonlinear data transformations, filtering of video and / or audio data, modeling operations (e.g., applying physical laws to determine the position, velocity, and other properties of objects), image rendering operations (e.g., tessellation shaders, vertex shaders, geometry shaders, and / or pixel / fragment shaders), general computational operations, etc. In operation, PPU 202 is configured to transfer data from system memory 104 and / or PP memory 204 to one or more on-chip memory units, process the data, and write the resulting data back to system memory 104 and / or PP memory 204. The resulting data can then be accessed by other system components (including CPU 102, another PPU 202 in accelerator processing subsystem 112, or another accelerator processing subsystem 112 in computer system 100).
[0042] As described above, the accelerator processing subsystem 112 may include any number of PPUs 202. For example, multiple PPUs 202 may be provided on a single add-on card, or multiple add-on cards may be connected to the communication path 113, or one or more PPUs 202 may be integrated into a bridge chip. The PPUs 202 in a multi-PPU system may be the same or different from each other. For example, different PPUs 202 may have different numbers of processing cores and / or different numbers of PP memories 204. In an implementation with multiple PPUs 202, these PPUs can operate in parallel to process data at a higher throughput than that possible with a single PPU 202. Systems including one or more PPUs 202 can be implemented in various configurations and form factors, including but not limited to desktop computers, laptops, handheld personal computers or other handheld devices, servers, workstations, game consoles, embedded systems, etc.
[0043] Figure 3 According to various embodiments Figure 2A block diagram of a General Processing Cluster (GPC) 208 included in a Parallel Processing Unit (PPU) 202 is provided. In operation, the GPC 208 can be configured to execute a large number of threads in parallel to perform graphics, general processing, and / or computational operations. As used herein, a “thread” refers to an instance of a specific program executed on a specific input dataset. In some embodiments, Single Instruction Multiple Data (SIMD) instruction issuing techniques are used to support the parallel execution of a large number of threads without providing multiple independent instruction units. In other embodiments, Single Instruction Multiple Threading (SIMT) techniques are used to support the parallel execution of a large number of typically synchronous threads using a general instruction unit configured to issue instructions to a set of processing engines in the GPC 208. Unlike SIMD execution architectures, where all processing engines typically execute the same instructions, SIMT execution allows different threads to more easily follow different execution paths through a given program. Those skilled in the art will recognize that a SIMD processing architecture represents a subset of the functionality of a SIMT processing architecture.
[0044] The operation of GPC 208 is controlled via pipeline manager 305, which distributes processing tasks received from work assignment units (not shown) within task / work unit 207 to one or more streaming multiprocessors (SMs) 310. Pipeline manager 305 can also be configured to control work assignment crossbar switch 330 by specifying the destination of the processed data output from SM 310.
[0045] In one embodiment, GPC 208 includes a set of M SMs 310, where M ≥ 1. Furthermore, each SM 310 includes a set of functional execution units (not shown), such as execution units and load-memory units. Processing operations specific to any functional execution unit can be pipelined, allowing new instructions to be issued for execution before previous instructions have completed. Any combination of functional execution units in a given SM 310 can be provided. In various embodiments, functional execution units can be configured to support a wide variety of operations, including integer and floating-point arithmetic (e.g., addition and multiplication), comparison operations, Boolean operations (e.g., AND, OR, XOR), bit shifting, and computation of various algebraic functions (e.g., plane interpolation and trigonometric functions, exponential and logarithmic functions, etc.). Advantageously, the same functional execution unit can be configured to perform different operations.
[0046] In operation, each SM 310 is configured to handle one or more thread groups. As used herein, a “thread group” or “warp” refers to a group of threads that execute the same program simultaneously on different input data, where a thread in the group is assigned to a different execution unit in the SM 310. A thread group may include fewer threads than the number of execution units in the SM 310, in which case some executions may be idle during the cycle while the thread group is being processed. A thread group may also include more threads than the number of execution units in the SM 310, in which case processing may occur in consecutive clock cycles. Since each SM 310 can support up to G thread groups simultaneously, up to G*M thread groups can be executed in the GPC 208 at any given time.
[0047] Furthermore, multiple related thread groups can be active simultaneously (at different stages of execution) within the SM 310. This collection of thread groups is referred to herein as a “cooperative thread array” (“CTA”) or “thread array”. The size of a particular CTA is equal to m*k, where k is the number of threads executing concurrently within the thread group, which is typically an integer multiple of the number of execution units in the SM 310, and m is the number of concurrently active thread groups within the SM 310. In various embodiments, software applications written in the Computing Unified Device Architecture (CUDA) programming language describe the behavior and operations of threads executing on the GPC 208, including any of the aforementioned behaviors and operations. A given processing task can be specified in a CUDA program, allowing the SM 310 to be configured to execute and / or manage general-purpose computing operations.
[0048] although Figure 3 Not shown, but each SM 310 includes a Level 1 (L1) cache, or uses space in a corresponding L1 cache outside the SM 310 to support load and store operations performed by the execution unit, etc. Each SM 310 can also access a Level 2 (L2) cache (not shown) shared among all GPCs 208 in the PPU 202. The L2 cache can be used to transfer data between threads. Finally, the SM 310 can also access off-chip “global” memory, which may include PP memory 204 and / or system memory 104. It should be understood that any memory outside the PPU 202 can be used as global memory. Furthermore, as Figure 3As shown, GPC 208 may include a Level 1.5 (L1.5) cache 335, which is configured to receive and store data requested from memory by SM 310 via memory interface 214. This data may include, but is not limited to, instructions, uniform data, and constant data. In embodiments where GPC 208 has multiple SMs 310, the SMs 310 may advantageously share common instructions and data cached in the L1.5 cache 335.
[0049] Each GPC 208 may have an associated memory management unit (MMU) 320 configured to map virtual addresses to physical addresses. In various embodiments, the MMU 320 may reside within the GPC 208 or the memory interface 214. The MMU 320 includes a set of page table entries (PTEs) for mapping virtual addresses to physical addresses of tiles or memory pages, optionally mapping to cache line indices. The MMU 320 may include an address translation back buffer (TLB) or a cache residing within the SM 310, one or more L1 caches, or the GPC 208.
[0050] In graphics and computing applications, the GPC 208 can be configured to couple each SM 310 to a texture unit 315 to perform texture mapping operations, such as determining texture sampling locations, reading texture data, and filtering texture data.
[0051] In operation, each SM 310 sends the processed task to the work allocation crossbar switch 330 so that the processed task can be provided to another GPC 208 for further processing, or the processed task can be stored in the L2 cache (not shown), the parallel processing memory 204, or the system memory 104 via the crossbar switch unit 210. Furthermore, the pre-raster operation (preROP) unit 325 is configured to receive data from the SM 310, direct the data to one or more raster operation (ROP) units in the partitioning unit 215, perform color mixing optimization, organize pixel color data, and perform address translation.
[0052] It should be understood that the core architecture described herein is illustrative and can be varied and modified. Among other things, the GPC 208 may include any number of processing units, such as SM 310, texture units 315, or preROP units 325. Furthermore, as combined with the above... Figure 2The PPU 202 may include any number of GPCs 208, which are configured to be functionally similar to each other, such that execution behavior is independent of which GPC 208 receives a specific processing task. Furthermore, each GPC 208 operates independently of the other GPCs 208 in the PPU 202 to execute tasks of one or more applications. In view of the foregoing, those skilled in the art should understand that... Figures 1-3 The architecture described herein does not limit the scope of the various embodiments of this disclosure.
[0053] Please note that, as used herein, references to shared memory may include any one or more technically feasible memories, including but not limited to local memory shared by one or more SM 310s or memory accessible via memory interface 214, such as cache memory, parallel processing memory 204, or system memory 104. Also note that, as used herein, references to cache memory may include any one or more technically feasible memories, including but not limited to L1 cache, L1.5 cache, and L2 cache.
[0054] Cache line sharing among multiple allocations in sector cache memory
[0055] Various embodiments include techniques for managing cache memory in computing systems that include sector cache memory. Prior to the disclosed methods, instead of the entire cache line, the cache controller would allocate the entire cache line and then load and store only the necessary sectors. As a result, using these prior art techniques, the cache controller does not consume memory transfer bandwidth used to load sectors not needed by the software application. However, these techniques can lead to reduced cache memory utilization when fewer than a full cache line is loaded. Consequently, the cache memory may have low utilization, potentially resulting in a large portion of the cache memory being unused and unusable. Generally, prior art techniques for managing memory transfer bandwidth and cache memory utilization involve determining the optimal cache line size.
[0056] In contrast, the disclosed technique is a different method for improving cache memory utilization without excessively increasing memory transfer bandwidth consumption. The disclosed technique includes a sector cache memory that provides a mechanism for software applications to share portions of a cache line between two or more individual allocations. A first allocation may allocate one or more sectors of an empty and available cache line. If the first allocation results in one or more unused sectors, a second allocation may allocate one or more unused sectors of the cache line. If the second allocation also results in one or more unused sectors, an additional allocation may allocate one or more unused sectors of the cache line. Furthermore, if two allocations for the same cache line have overlapping logical sectors, the sectors of one of the allocations may be moved, for example, by a rotation function, to eliminate the overlap before loading the sectors into the physical cache memory. In this way, multiple allocations can share the same cache line in the cache memory.
[0057] Figure 4 Included according to various embodiments Figure 1 CPU 102 and / or Figure 2 A block diagram of the cache memory system 400 in the PPU 202 is shown. As illustrated, the cache memory system 400 includes, but is not limited to, a backup store 410, a cache memory 420, a cache tag store 430, an allocation tracker 440, and a next-level cache memory 450. The cache memory 420 may include any one or more of the technically feasible memories described herein, including but not limited to L1 cache, L1.5 cache, or L2 cache. The cache memory 420 maintains cache lines 422 loaded from the backup store 410.
[0058] Backup storage 410 may include any or more technically feasible memories described herein, including but not limited to system memory 104 or PP memory 204. Alternatively or additionally, backup storage 410 may include cache memory remote from CPU 102 and / or PPU 202 relative to cache memory 420. In some examples, cache memory 420 may be an L2 cache and backup storage 410 may be system memory 104 or PP memory 204. Cache memory 420 may be an L1.5 cache and backup storage 410 may be an L2 cache. Cache memory 420 may be an L1 cache and backup storage 410 may be an L1.5 cache, and so on.
[0059] Unless cache memory 420 is the cache memory closest to CPU 102 and / or PPU 202, the next-level cache memory 450 includes cache memories that are closer to CPU 102 and / or PPU 202 than cache memory 420. In some examples, cache memory 420 may be an L2 cache and the next-level cache memory 450 may be an L1.5 cache. Cache memory 420 may be an L1.5 cache, while the next-level cache memory 450 may be an L1 cache, and so on.
[0060] In operation, cache controller 470 manages cache memory 420. Cache controller 470 loads cache lines 422 or portions thereof from cache memory 420, loading data from backup storage 410 when needed or just before it is used by a processing unit. In general, cache memory 420 includes N+1 cache lines 422, numbered cache lines 422(0), 422(1), 422(2), ..., 422(N). Each cache line 422 includes four sectors stored in sector 0 memory 424(0), sector 1 memory 424(1), sector 2 memory 424(2), and sector 3 memory 424(3), respectively. When loading cache line 422, cache controller 470 may load a single sector memory 424 of the cache line or may load two, three, or all four sectors of the cache line 422.
[0061] In some examples, each cache line consists of 128 bytes and the cache line consists of four sectors, resulting in 128 bytes per cache line divided by four sectors per cache line equaling 32 bytes per sector. The communication channel 460 between the backing store 410 and the cache memory 420 can have the same data width as a sector, i.e., 32 bytes.
[0062] To load one or more sectors of memory 424 into cache line 422, cache controller 470 starts with a virtual address within the software application's virtual address space. Cache controller 470 divides the virtual address into several parts, including a cache line tag, sector number, and sector offset. The sector number and sector offset together form the cache line offset. For a cache line with 128 bytes (or 2... 7 A cache memory 420 has a cache line of 422 bytes, and the cache line offset is the seven least significant bits (LSBs) of the virtual address. Correspondingly, the cache line tag is the portion of the virtual address excluding the seven LSBs. For a cache memory 420 with four (or 2) bytes, the cache line offset is the seven least significant bits (LSBs) of the virtual address. 2The cache memory 420 consists of cache lines 422 (or 2) sectors, where the two most significant bits (MSB) of the cache line offset are the sector number. Furthermore, because each sector comprises 32 bytes (or 2)... 5 (5 LSBs), so the sector offset is five LSBs of the cache line offset.
[0063] The memory management unit translates the cache line tag portion of the virtual address into a physical address that addresses the beginning of the corresponding 128 bytes in the backing store 410. If the entire cache line 422 is being loaded, the cache controller 470 generates four load transactions on the communication channel 460 to retrieve four sectors of each 32-byte size and stores these four sectors in the four-sector memory 424 of the cache line 422. If a single sector of the cache line 422 is being loaded, the cache controller 470 combines the physical address at the beginning of the 128 bytes in the backing store 410 with a 2-bit sector number to identify the beginning address of the sector in the backing store 410. The cache controller 470 generates a single load transaction via the communication channel 460 to retrieve the 32-byte sector and stores it in the corresponding sector memory 424 of the cache line 422. Similarly, the cache controller 470 can load two or three sectors of the cache line by generating two or three 32-byte load transactions respectively. The cache controller 470 stores cache line tags along with status indicators for each sector in the cache tag memory 430. Each sector is associated with at least two status indicators: a valid indicator and a dirty indicator.
[0064] In summary, the valid indicator and the dirty indicator indicate one of three potential state conditions for the corresponding sector. First, if the valid indicator indicates the sector is invalid, the cache controller 470 cannot rely on any data stored in the corresponding sector memory 424, regardless of the state of the dirty indicator. Second, if the valid indicator indicates the sector is valid and the dirty indicator indicates the sector is clean (i.e., not dirty), the corresponding sector memory 424 of cache line 422 contains valid data. Furthermore, the data is clean, indicating that the processing unit has not yet written new data to the sector memory that has not yet been written to the backing store 410. Therefore, the data in sector memory 424 has not changed since the last time data was retrieved from backing store 410. Third, if the valid indicator indicates the sector is valid and the dirty indicator indicates the sector is dirty, the corresponding sector memory 424 of cache line 422 contains valid data, but the data in this sector memory includes new data that has not yet been written to backing store 410. Therefore, the data in sector memory 424 does not match the corresponding data in backing store 410. At some point in the future, cache controller 470 writes valid, dirty sectors from cache memory 420 to backing memory 410, ensuring that the data in backing memory 410 matches the data in sector memory 424. Until a write-back occurs, the processing unit accesses the sector memory 424 of cache line 422 to ensure that the processing unit accesses an updated version of the data.
[0065] When a processing unit subsequently accesses data contained in one or more valid sectors, memory management accesses the cache line tag in cache tag memory 430 to access the corresponding sector memory 424 of cache line 422 via communication channel 452. Similarly, the next-level cache memory 450 can load an entire cache line or its sectors by generating a load transaction on communication channel 454 to load data from cache memory 420 into the next-level cache memory 450.
[0066] In some examples, cache controller 470 allocates cache line 422 without storing the tag in cache tag memory 430; this is referred to herein as transient allocation and / or transient cache line allocation. Transient allocation is useful in various use cases where a software application only accesses the allocated sector once, such as streaming applications. For transient allocations, memory management does not store the tag in cache tag memory 430. However, memory management does store a status indicator for transient cache line 422 in cache tag memory 430.
[0067] To load one or more sectors of memory 424 into cache line 422, cache controller 470 starts with a virtual address within the software application's virtual address space. Cache controller 470 divides the virtual address into several parts, including a cache line label, a sector number, and a sector offset. The sector number and sector offset together form the cache line offset. For transient cache line 422, the software application directly accesses the sectors of transient cache line 422.
[0068] Allocation tracker 440 monitors the most recent partial cache line allocations that meet the conditions for cache line sharing, as described herein. In doing so, allocation tracker 440 monitors cache memory 420 via communication channel 456 and cache tag memory 430 via communication channel 458. One type of partial cache line allocation suitable for cache line sharing is transient allocation. Additionally or alternatively, partial cache line allocations may qualify for cache line sharing if a mechanism exists to handle potential conflicts in any potential subsequent sector allocations within the same line. For example, cache controller 470 may choose to allocate a new cache line 422 when inter-line cache conflicts evolve over time within a single shared cache line 422.
[0069] More generally, cache line sharing is enabled when the cache controller 470 makes a one-time selection at allocation time regarding which sectors to use for a given allocation. Traditional sector caches allow this selection to be revisited over time to expand partial cache line allocations sector by sector until the partial allocation is a complete cache line allocation. However, to achieve cache line sharing, the cache controller 470 performs a one-time allocation that is unaffected by expansion. Additionally or alternatively, the cache controller 470 uses a subset of allocations that typically involve a single allocation, such as transient cache line allocations.
[0070] Allocation tracker 440 stores the last N qualified partial cache line allocations eligible for cache line sharing. When N=1, allocation tracker 440 stores the last qualified partial cache line allocation. When N=2, allocation tracker 440 stores the last two qualified partial cache line allocations, and so on. When cache controller 470 receives a request for a new qualified partial cache line allocation, cache controller 470 determines whether the new qualified partial cache line allocation can be combined with any qualified partial cache line allocation stored in allocation tracker 440. If the new qualified partial cache line allocation can be combined with at least one of the qualified partial cache line allocations stored in allocation tracker 440, cache controller 470 stores a cache tag in cache tag memory 430 such that the new qualified partial cache line allocation points to the same cache line 422 in cache memory 420 that is combined with the new qualified partial cache line allocation.
[0071] If a new qualified partial cache line allocation cannot be combined with at least one of the qualified partial cache line allocations stored in allocation tracker 440, cache controller 470 stores a cache tag in cache tag memory 430 such that the new qualified partial cache line allocation points to an unused cache line 422 in cache memory 420. Cache controller 470 stores a copy of the new qualified partial cache line allocation in allocation tracker 440. If allocation tracker 440 is full, cache controller 470 evicts the oldest qualified partial cache line allocation from allocation tracker 440 before storing a new qualified partial cache line allocation.
[0072] Figure 5 It is based on various embodiments with a one-to-one cache line mapping Figure 4A block diagram of cache memory 420 and cache tag memory 430 is shown. As shown, cache memory 420 includes, but is not limited to, a set of cache lines 422, each cache line 422 being divided into multiple sectors. The first cache line 422(0) is divided into four sectors, labeled as sector 0 522(0), sector 1 524(0), sector 2 526(0), and sector 3 528(0). The second cache line 422(1) is also divided into four sectors, labeled as sector 0 522(1), sector 1 524(1), sector 2 526(1), and sector 3 528(1). The third cache line 422(2) is divided into four sectors, labeled as sector 0 522(2), sector 1 524(2), sector 2 526(2), and sector 3 528(2), and so on. In general, the cache memory 420 includes N+1 cache lines, and the last cache line 422(N) is divided into four sectors, labeled as sector 0 522(N), sector 1 524(N), sector 2 526(N) and sector 3 528(N).
[0073] The cache tag memory 430 includes, but is not limited to, N+1 cache line tags 510, numbered cache line tags 510(0), 510(1), 510(2), ..., 510(N). Each cache line tag 510 has a one-to-one correspondence with a cache line 422 in the cache memory 420. In such an example, cache line tag 510(0) corresponds to cache line 422(0), cache line tag 510(1) corresponds to cache line 422(1), cache line tag 510(2) corresponds to cache line 422(2), and so on.
[0074] As shown in the figure, three cache lines 422(0), 422(1), and 422(3) are filled with cached data. Therefore, the cache line label 510(0) with label address A0 points to the corresponding cache line 422(0). The cache line label 510(1) with label address A1 points to the corresponding cache line 422(1). The cache line label 510(3) with label address A3 points to the corresponding cache line 422(3).
[0075] All sectors in cache lines 422(0), 422(1), and 422(3) are filled with cache data. In this respect, sectors 0 522(0), 1 524(0), 2 526(0), and 3 528(0) are filled with cache data D0a, D0b, D0c, and D0d, respectively. Sectors 0 522(1), 1 524(1), 2 526(1), and 3 528(1) are filled with cache data D1a, D1b, D1c, and D1d, respectively. Sectors 0 522(3), 1 524(3), 2 526(3), and 3 528(3) are filled with cache data D3a, D3b, D3c, and D3d, respectively. Sectors in cache lines 422(2) and 422(N) are not filled with cache data. Therefore, cache line labels 510(2) and 510(N) do not point to the corresponding cache lines 422(2) and 422(N).
[0076] In some examples, the valid sector status indicators of cache lines can be combined into a binary valid sector mask, which indicates which sectors of the cache line are valid. The valid sector mask can be ordered from sector 0 on the left to sector 3 on the right. Furthermore, the valid status indicator can be set to 1 when a sector is valid and to 0 when a sector is invalid. In such examples, the valid sector mask for cache lines 422(0), 422(1), and 422(3) is set to 0b1111, indicating that all four sectors are valid. The valid sector mask for cache lines 422(2) and 422(N) is set to 0b0000, indicating that all four sectors are invalid.
[0077] Figure 6 It is based on various embodiments and has a flexible cache line mapping. Figure 4 A block diagram of cache memory 420 and cache tag memory 430 is shown. As illustrated, there is no one-to-one correspondence between cache line tags 510 and cache lines 422 in cache memory 420. Therefore, each cache line tag 510 can correspond to any cache line 422 in cache memory 420. Each cache line tag 510 is associated with a cache line address (…). Figure 6 (Not shown in the image) is associated with cache line address identifier 422, which is associated with the corresponding cache line label 510.
[0078] As shown in the figure, three cache lines 422(0), 422(2), and 422(N) are partially filled with cached data. Cache line label 510(0) with label address A0 includes a cache line address pointing to cache line 422(N). Cache line label 510(1) with label address A1 includes a cache line address pointing to cache line 422(2). Cache line label 510(3) with label address A3 includes a cache line address pointing to cache line 422(0).
[0079] Figure 7 It is according to various embodiments having unused sectors Figure 4 A block diagram of cache memory 420 and cache tag memory 430. As shown, cache line tags 510 do not have a one-to-one correspondence with cache lines 422 in cache memory 420. Furthermore, three cache line tags 510(0), 510(1), and 510(3) represent transient cache line allocations. These transient cache line tags 510(0), 510(1), and 510(3) include cache line addresses pointing to the corresponding cache lines 422 in cache memory 420, but do not include tag addresses. More specifically, cache line tag 510(0) does not have a tag address, but includes a cache line address pointing to cache line 422(N). Cache line tag 510(1) does not have a tag address, but includes a cache line address pointing to cache line 422(2). Cache line tag 510(3) does not have a tag address, but includes a cache line address pointing to cache line 422(0).
[0080] Cache line 422(0) has one valid sector, sector 3 528(0), which contains the data D3d. The remaining sectors, sector 0 522(0), sector 1 524(0), and sector 2 526(0), are invalid and therefore contain valid data. As a result, the valid sector mask of cache line 422(0) is 0b0001. Cache line 422(2) has two valid sectors, sector 0 522(2) and sector 3 528(2), which contain the data D1a and D1d, respectively. The remaining sectors, sector 1 524(2) and sector 2 526(2), are invalid and therefore contain valid data. As a result, the valid sector mask of cache line 422(2) is 0b1001. Cache line 422(N) also has two valid sectors, sector 0522(N) and sector 1524(N), which contain data D0a and D0b respectively. The remaining sectors, sector 2526(N) and sector 3528(N), are invalid and therefore contain valid data. As a result, the valid sector mask of cache line 422(2) is 0b1100. Since cache line labels 510(0), 510(1), and 510(3) represent transient cache line allocations, these transient cache line allocations can be combined to share one or more cache lines 422.
[0081] Figure 8 It is based on various embodiments with cache line sharing Figure 4A block diagram of cache memory 420 and cache tag memory 430. When cache controller 470 receives a request for a new qualified partial cache line allocation, cache controller 470 determines whether the new qualified partial cache line allocation can be combined with any qualified partial cache line allocation stored in allocation tracker 440. If the total number of sectors in the new qualified partial cache line allocation and the previously qualified partial cache line allocation does not exceed the number of sectors in each cache line 422, the two allocations are candidates for cache line sharing. If the two allocations do not intersect, meaning that the two allocations do not store data in the same sector, then the two allocations can share the same cache line 422. In this regard, cache line tag 510(0) includes a cache line address that points to the corresponding cache line 422(N) in cache memory 420. Cache line label 510(0) includes a valid mask 0b1100, indicating that the cache line is allocated to store data D0a in sector 0522(N) and data D0b in sector 1524(N). Cache line label 510(3) includes a cache line address, which also points to the corresponding cache line 422(N) in cache memory 420. Cache line label 510(3) includes a valid mask 0b0001, indicating that the cache line is allocated to store data D3d in sector 3528(N).
[0082] Figure 9 It is based on other embodiments with cache line sharing. Figure 4 A block diagram of cache memory 420 and cache tag memory 430. When cache controller 470 receives a request for a new qualified partial cache line allocation, cache controller 470 determines whether the new qualified partial cache line allocation can be combined with any qualified partial cache line allocation stored in allocation tracker 440. If the total number of sectors in the new qualified partial cache line allocation and the previously qualified partial cache line allocation does not exceed the number of sectors per cache line 422, the two allocations are candidates for cache line sharing. If the two allocations are not disjoint, the two allocations can still share a cache line if the sectors of the new qualified partial cache line allocation can be moved to be disjoint with the previously qualified partial cache line allocation. To distinguish between sectors in a cache allocation request and sectors in the physical cache, sectors in a cache allocation request are referred to herein as logical sectors, while sectors in the physical cache are referred to herein as physical sectors.
[0083] One mechanism for moving sectors allocated in a new qualified partial cache line is to rotate the sectors allocated in the new qualified partial cache line before storing the sectors in the cache line. Given a four-sector cache, this rotation value can be a two-digit rotation value indicating the number of sector positions used to rotate the sectors allocated in the new qualified partial cache line. In some examples, a rotation value of 0b00 could indicate that the new qualified partial cache line allocation was not rotated before storing the sectors in cache line 422. A rotation value of 0b01 could indicate that the new qualified partial cache line allocation was rotated once, three positions to the right or equivalently to the left, before storing the sectors in cache line 422. A rotation value of 0b10 could indicate that the new qualified partial cache line allocation was rotated two positions to the right, or equivalently, two positions to the left, before storing the sectors in cache line 422. A rotation value of 0b11 could indicate that the new qualified partial cache line allocation was rotated three positions to the right, or equivalently, one position to the left, before storing the sectors in cache line 422.
[0084] In this respect, cache line label 510(0) includes a cache line address pointing to the corresponding cache line 422(N) in cache memory 420. Cache line label 510(0) includes a valid mask 0b1100 indicating that the cache line allocation stores data D0a in sector 0 522(N) and data D0b in sector 1 524(N). The sectors allocated by the cache line are not rotated. Therefore, the rotation value of cache line label 510(0) is 0b00. Cache line label 510(1) includes a cache line address that also points to the corresponding cache line 422(N) in cache memory 420. Cache line label 510(1) includes a valid mask 0b1001 indicating that the cache line allocation stores data D1a and D1d, corresponding to sectors 0 and 3 before rotation, respectively. The sectors allocated by the cache line are rotated three positions to the right, or equivalently one position to the left. Therefore, the rotation value of cache line tag 510(1) is 0b11. As a result, the cache line allocation stores data D1d in sector 2 526(N) and data D1a in sector 3 528(N). Cache line tag 510(3) includes the cache line address pointing to the corresponding cache line 422(0) in cache memory 420. Cache line tag 510(3) includes a valid mask 0b0001, indicating that the cache line allocation stores data D3d in sector 3 528(0). The sectors of the cache line allocation are not rotated. Therefore, the rotation value of cache line tag 510(3) is 0b00.
[0085] This rotation technique allows cache lines to be shared under various conditions. In some examples, up to four partial cache line allocations that meet the single-sector condition can be combined. Two allocations can be combined if they allocate contiguous sectors (including allocations of sectors 0 and 3 as contiguous allocations). Two allocations can also be combined if they allocate two sectors every other sector. Three-sector allocations can be combined with any single-sector allocation. Other combinations are also possible using this rotation technique.
[0086] This rotation technique covers many useful combinations for cache line sharing, but not all possible combinations. Within the scope of this disclosure, other techniques can be employed to cover other cases not covered by rotation. In extreme cases, a full sector-by-sector allocation could be used to cover all possible combinations, although this increases complexity. Even so, the disclosed rotation technique covers most possible combinations with a relatively simple hardware implementation.
[0087] As described herein, cache line sharing is possible when two allocations meet the cache line sharing criteria. In some examples, transient cache line allocations meet the cache line sharing criteria. Transient cache line allocations can be useful for data streaming applications and other similar applications where data is stored in cache memory 420, read from cache memory 420 once, and then discarded. Transient cache line allocations are also useful for applications with regular, predictable memory access patterns. Such applications use the data storage capacity of cache memory as a temporary register memory rather than general cache storage. Utilizing transient cache line allocations, cache controller 470 assembles uncached storage into transient cache line 422 before transferring uncached storage to backing storage 410. As a result, cache controller 470 can access transient storage and tagged storage using similar data paths. When cache controller 470 uses transient cache line 422, the software application cannot directly observe this temporary use of cache line 422. In some examples, a portion of cache memory 420 is partitioned off and separated from a portion of cache memory used for tagged cache line 422 and transient cache line. This partitioning is referred to herein as scratchpad memory or shared memory. Scratchpad memory is persistent memory that software applications can request during the kernel's lifetime. Software applications can use and share scratchpad memory as needed during the kernel's lifetime.
[0088] In some examples, the cache tag memory 430 may include tagged cache line tags 510 and transient cache line tags 510. In such examples, the cache controller 470 may restrict cache line sharing to allocations represented by transient cache line tags 510. Where possible, sharing cache lines for transient cache line allocations can improve cache memory utilization and performance. Furthermore, the likelihood of evicting tagged cache allocations can be reduced because fewer cache lines 422 are consumed by transient cache allocations.
[0089] It should be understood that the system shown herein is illustrative and variations and modifications are possible. The techniques described herein are in the context of a cache memory 420 having a cache line of 128 bytes and four sectors of memory 424, each of 32 bytes. However, the disclosed techniques can be used with cache lines 422 of any size, any number of sectors per cache line, and / or sector memories 424 of any size. The techniques described herein can be applied in any combination to any one or more cache levels in a cache memory system. In this respect, the cache line size, number of sectors, and sector memory size of each cache level can be the same or different from each other and can be combined arbitrarily. The techniques described herein are in the context of a communication channel 460 between a backing store 410 and a cache memory 420 having the same data width as the sector memories 424, i.e., 32 bytes. However, the communication channel 460 between the backup storage 410 and the cache memory 420 can have any data width, including fewer bytes than the sector memory 424 or more bytes than the sector memory 424. The techniques described herein can be applied in any combination to any CPU 102, PPU 202, and / or any other processing unit.
[0090] In some examples, the memory system may include multiple cache memories implementing the techniques described herein. For example, an L1 cache implementing the disclosed techniques may be coupled to an L2 cache that also implements the disclosed techniques. In this case, the L1 cache may have cache lines with specific cache line tags, while the L2 cache may have cache lines with the same cache line tags. Cache lines in the L1 cache and cache lines in the L2 cache may or may not be in the same relative position in their respective cache memories. Furthermore, sectors of cache lines in the L1 cache and corresponding sectors of cache lines in the L2 cache may have the same status indicator or different status indicators, and such combinations are possible.
[0091] Furthermore, cache lines in the L1 cache and cache lines in the L2 cache can have the same number of bytes and / or the same sector size, or they can have different numbers of bytes and / or different sector sizes, and can be combined arbitrarily. In this respect, a cache line in one cache memory can be mapped to multiple cache lines in another cache memory, and vice versa. Moreover, multiple cache lines in one cache memory can store one or more identical sectors from a given cache line in another cache memory in any combination.
[0092] In some examples, cache line sharing can be restricted to operating with a limited set of instructions and without using other instructions. In some cases, restricting cache line sharing to certain instructions simplifies the implementation of the disclosed techniques. In some examples, cache line sharing can be implemented to operate between threads within the same warp instruction. In such examples, cache controller 470 can allow combined allocations only when the allocation is requested by the same warp instruction. Additionally or alternatively, cache line sharing can be implemented to operate between threads spanning multiple warps. In such examples, cache controller 470 can allow combined allocations regardless of which one or more warps request the allocation.
[0093] In some examples, cache memory 420 can be divided into independent tag groups, each managed by a different cache tag memory 430 instance. Furthermore, each tag group has a separate and independent allocation tracker 440. In such examples, four independent cache line sharing operations can be performed in the same clock cycle. Typically, any given cache line 422 can be allocated by one tag group at a time. In some examples, cache lines 422 are physically bound to a specific tag group, where cache lines 422 are divided into cache line pools, with each cache line pool pinned to a specific tag group. Additionally or alternatively, tag groups can manage the same set of shared cache lines 422. Additionally or alternatively, a single allocation tracker 440 can manage multiple tag groups.
[0094] In some examples, the cache controller 470 shares cache lines in a manner that helps reduce data set conflicts, also referred to herein as memory data set conflicts. A data set conflict occurs when the cache controller 470 attempts to access a memory device multiple times within the same clock cycle. When a data set conflict occurs, the multiple device accesses are serialized to occur over multiple clock cycles. For example, two data set conflicts occur within two clock cycles, three data set conflicts occur within three clock cycles, and so on.
[0095] Typically, cache memory 420 is implemented such that each sector for a set of cache lines is stored on a different set of memory devices. Thus, sector 0 is stored on one set of memory devices, sector 1 on a second set, sector 2 on a third set, and so on. As a result, cache controller 470 can access two or more sectors in a single clock cycle without causing a data set conflict. On the other hand, multiple accesses to the same sector can lead to a data set conflict. For example, if an instruction accesses sector 0 on a cache line 422 and concurrently accesses sector 0 on one or more additional cache lines 422, these accesses will result in a data set conflict.
[0096] As described herein, access to cache line 422 associated with transient allocations is typically limited to storing a sector into cache memory once, loading a sector from cache memory once, and then deallocating that sector. Furthermore, the access patterns of sectors in transient allocations are generally regular and predictable. These access patterns can be understood at design time when one or more software engineers build the software application. Additionally or alternatively, cache controller 470 can observe access patterns at runtime. In either case, when sharing cache lines, cache controller 470 may attempt to allocate sectors to cache lines in a manner that distributes concurrent access across multiple sectors and reduces the likelihood of multiple concurrent accesses to the same sector.
[0097] In some examples, cache controller 470 allocates sectors in a manner that serves two somewhat conflicting optimization objectives. The first optimization objective is to pack multiple partial cache line allocations, for example, through sector rotation, to reduce the number of cache lines consumed, or equivalently, to increase the number of sectors used per cache line, thereby improving cache memory utilization. The second optimization objective is to reduce the likelihood of set conflicts when accessing data in cache memory 420. Generally, the number of sectors packed and the likelihood of set conflicts are positively correlated. As the number of sectors used per cache line increases, the likelihood of set conflicts also increases. Therefore, one of the optimization objectives can be prioritized relative to the other, while a balance can be struck between the two objectives. In some examples, cache controller 470 may prioritize increasing cache line utilization, with the secondary objective of reducing set conflicts for various anticipated and / or common access patterns.
[0098] In the example method, cache memory 420 can be divided into tag groups (denoted as B), where each tag group can perform cache line allocation within the same clock cycle. Furthermore, thread bundles are typically executed as a wavefront sequence (denoted as N), where each wavefront executes thread bundle instructions for a subset of threads within the thread bundle. For example, a 32-thread thread bundle can execute thread bundle instructions over four waves, where each wavefront executes thread bundle instructions for a group of eight threads. Therefore, thread bundle instructions are executed over four clock cycles. Typically, the amount of data accessed by each wavefront is equal to the amount of data stored in a full cache line. If the access pattern does not cause data set conflicts, each wavefront can execute within a single clock cycle. Between consecutive waves, the general goal is to share cache lines between two waves on the same tag group.
[0099] Because each wavefront executes within a single clock cycle, data set conflicts occur within a single wavefront. Typically, data set conflicts occur when a wavefront generates multiple accesses from different tag groups accessing the same sector. Therefore, the cache controller 470 allocates sectors in a manner that attempts to avoid sector conflicts within a wavefront and from multiple tag groups, while packing sectors across multiple wavefronts within each tag group.
[0100] In one example technique, cache controller 470 determines whether a new cache line allocation request is eligible for cache line sharing. Cache controller 470 may determine suitability for cache line sharing based on whether the allocation is used for a portion of the cache line, whether the allocation comes from the same thread bundle (but a different wavefront) as another allocation, which thread bundle instruction is being executed, and / or similar factors. If the allocation is eligible for cache line sharing, cache controller 470 calculates the sector number from the equation (B+N)%4, where "%" is a modulo operation, B is the tag group number, and N is the wavefront number.
[0101] By including the tag group number B in the equation, the sector number increments for each consecutive tag group. As a result, each tag group prioritizes access to different sectors, thus reducing the likelihood of data group collisions. By including the wavefront number N in the equation, the sector number increments with each consecutive wavefront of the thread bundle instruction, resulting in the packing of valid sectors within each tag group.
[0102] The cache controller 470 allocates sectors for cache line allocation based on the number of sectors in the allocation and the number of sectors calculated from the equation (B+N)%4. For the case where one sector is allocated per cache line, the cache controller 470 assigns (B+N)%4 sectors to the allocation. For the case where two sectors are allocated per cache line, the cache controller 470 first determines whether the allocation is for two consecutive sectors or for every other sector.
[0103] For two consecutive sector allocations, the valid allocation mask is 0b0011, 0b0110, 0b1100, or 0b1001. If the allocation tracker 440 does not contain any previous allocations, no other allocation can be combined with the new allocation. The cache controller 470 assigns sector (B+N)%4 and the immediately preceding sector to the allocation. As a result, if the next allocation is a sector allocation, the cache controller 470 can assign the next sector to that allocation without causing a conflict between the database and the two sector allocations.
[0104] If allocation tracker 440 does indeed contain a previous allocation, cache controller 470 assigns sector (B+N)%4 and the immediately following sector to the current allocation. If the previous allocation was a single-sector allocation, the previous allocation may have assigned sectors before the current allocation. The current allocation assigns the current and subsequent sectors to avoid conflicts with the data set allocated by the previous one-sector allocation. If the previous allocation was a two-sector allocation, then the previous allocation likely assigned two sectors before the current allocation, as described above. The current allocation assigns the current and subsequent sectors to avoid conflicts with the data set allocated by the previous two-sector allocation. If the current allocation cannot be combined with a previous allocation stored in allocation tracker 440, cache controller 470 may evict one or more allocations stored in allocation tracker 440 and store the current allocation in allocation tracker 440. Cache controller 470 can then proceed to the case where allocation tracker 440 does not store any previous allocations.
[0105] For a two-sector allocation every other sector, the valid allocation mask is 0b0101 or 0b1010. The cache controller 470 assigns sector (B+N)%4 and a second subsequent sector to the current allocation. A subsequent one-sector allocation can be assigned to unused sectors between the two sectors allocated to the current allocation. A subsequent two-sector allocation every other sector can be assigned to unused sectors between the two sectors and the sector following the sector allocated to the current allocation.
[0106] For the case where each cache line allocates three sectors, the cache controller 470 assigns sectors as follows: If the allocation tracker 440 does not contain any previous allocations, the cache controller 470 rotates the currently allocated sectors so that unused sectors are located in sector ((B+N+1)%4), which is the preferred sector for the subsequent sector allocation. If the allocation tracker 440 contains a previous sector allocation, the cache controller 470 rotates the currently allocated sectors so that unused sectors are in the same position as sectors in the previous allocation. If the allocation tracker 440 contains previous multi-sector allocations, the current allocation cannot be combined with previous allocations stored in the allocation tracker 440. The cache controller 470 may evict one or more allocations stored in the allocation tracker 440 and store the current allocation in the allocation tracker 440. The cache controller 470 can then proceed to the case where the allocation tracker 440 does not store any previous allocations.
[0107] In some examples, other similar techniques different from the described techniques may be implemented within the scope of this disclosure. In some examples, the number of sectors calculated by formula (B+N)%4 may be replaced by formulas (BN)%4, (NB)%4, etc. Other variations from the described techniques are contemplated within the scope of this disclosure. Furthermore, the described techniques are in the context of cache memory 420, which includes four sectors per cache line 422. Various additional and / or alternative techniques may be employed for cache memory 420 with a different number of sectors per cache line 422. Even so, the general framework of (B+N)% (number of sectors) of reference sectors can be used as a starting point for such cache memory 420.
[0108] After a portion of cache line 422 is combined and stored in cache memory 420, data items stored in the portion of cache line 422 can subsequently be accessed, with or without rotation. Typically, a requester accessing a data item stored in the portion of cache line 422 provides a cache line index to the specific cache line 422 in cache memory 420. Additionally, the requester provides a 2-bit rotation value for conversion between logical and physical sectors. For tagged cache lines 422 that share cache lines 422, the cache line label 510 includes a 2-bit rotation value. Cache controller 470 uses the 2-bit rotation value to rotate access requests pointing to cache line 422 in cache memory 420. Similarly, when cache controller 470 performs a read access to data stored in shared cache line 422, cache controller 470 rotates requests pointing to transient cache line 422. This 2-bit rotation code is an enclosing element for accesses to cache memory 420 because the requester does not know whether the accessed data has been rotated. In some examples, a spin value is assigned when processing a request. The spin value is received via a downstream storage system along with the returned data. Alternatively or additionally, spin values can be generated, stored, and propagated using any technically feasible technique.
[0109] Figure 10 It is according to various embodiments for managing such as Figure 1 CPU 102 or Figure 2 The flowchart illustrates the method steps for using a cache memory for a processing unit such as the PPU 202. Additionally or alternatively, the method steps may be executed by one or more alternative accelerators, including but not limited to CPUs, GPUs, IPUs, NPUs, TPUs, NNPs, DPUs, VPUs, ASICs, FPGAs, etc., and can be arbitrarily combined. Although combined... Figure 1-9 The system describes the method steps, but those skilled in the art will understand that any system configured to perform the method steps in any order is within the scope of this disclosure.
[0110] As shown in the figure, method 1000 begins at step 1002, in which cache controller 470 detects a cache line allocation request to allocate a sector in a cache line of the cache memory.
[0111] In step 1004, the cache controller 470 determines whether a cache line allocation is suitable for cache line sharing. One type of partial cache line allocation eligible for sharing a cache line is a transient allocation. Additionally or alternatively, a cache line allocation may be eligible for cache line sharing if a mechanism exists to handle potential conflicts in any potential subsequent sector allocations within the same line. For example, the cache controller 470 may choose to allocate a new cache line 422 as inter-line conflicts within sectors of a single shared cache line 422 evolve over time. Furthermore, a cache line allocation is suitable for cache line sharing if it allocates fewer sectors than the entire cache line allocation.
[0112] More generally, cache line sharing is enabled when cache controller 470 makes a one-time selection at allocation time regarding which sectors to use for a given allocation. Traditional sector caches allow this selection to be revisited over time to expand partial cache line allocations sector by sector until the partial allocation is a complete cache line allocation. However, to achieve cache line sharing, cache controller 470 performs a one-time allocation that is unaffected by expansion. Additionally or alternatively, cache controller 470 uses a subset of allocations that typically involve a single allocation, such as transient cache line allocations.
[0113] If a cache line allocation allocates an entire cache line, then by definition, that allocation cannot be shared with another allocation. In some examples, the cache controller 470 may determine whether a cache line allocation is suitable for cache line sharing based on certain other factors. These factors may include whether the allocation comes from the same thread bundle (but a different wavefront) as another allocation, the specific thread bundle instruction being executed, etc.
[0114] If, in step 1004, the cache line allocation is not suitable for cache line sharing, method 1000 proceeds to step 1008, as described herein. However, if the cache line allocation meets the conditions for cache line sharing, method 1000 proceeds to step 1006, where cache controller 470 determines whether the cache line allocation can be combined with any previous allocation stored in allocation tracker 440. As described herein, tracker 440 monitors the most recent partial cache line allocations that meet the conditions for cache line sharing. Allocation tracker 440 stores the last N partial cache line allocations eligible to share cache lines. When N=1, allocation tracker 440 stores the last eligible partial cache line allocation. When N=2, allocation tracker 440 stores the last two eligible partial cache line allocations, and so on. When cache controller 470 receives a request for a new eligible partial cache line allocation, cache controller 470 determines whether the new eligible partial cache line allocation can be combined with any eligible partial cache line allocation stored in allocation tracker 440.
[0115] If, in step 1006, the cache line allocation cannot be combined with any previous allocation stored in the allocation tracker 440, method 1000 proceeds to step 1008, where cache controller 470 allocates one or more sectors of an empty cache line. In the previous method, the allocation of cache line 422 included requesting the allocation of the cache line and, in response, receiving an index X (between 0 and N) referencing cache line 422(X). In contrast, using the disclosed technique, the allocation of cache line 422 includes requesting the allocation of the cache line and simultaneously generating a sector mask for that allocation request. Furthermore, the disclosed technique includes receiving an index X (between 0 and N) referencing cache line 422(X) in response. Additionally, the disclosed technique includes receiving a two-bit rotation value for use whenever the allocation is referenced. In step 1010, cache controller 470 locates the logical sectors for the cache line allocation based on certain factors. These factors may include optimization objectives. The first optimization objective is to pack multiple partial cache line allocations, such as through sector rotation, to reduce the number of cache lines consumed, or equivalently, to increase the number of sectors used per cache line, thereby improving cache memory utilization. The second optimization objective is to reduce the likelihood of data set conflicts when accessing data in cache memory 420. Generally, the number of sectors packed and the likelihood of data set conflicts are positively correlated. As the number of sectors used per cache line increases, the likelihood of data set conflicts also increases. Therefore, one of the optimization objectives can be prioritized relative to the other, and a balance can be struck between the two objectives. In some examples, cache controller 470 may prioritize increasing cache line utilization, with the secondary objective being reducing memory set conflicts for various anticipated and / or common access patterns. This paper describes techniques for balancing these two optimization objectives.
[0116] In step 1012, the cache controller 470 loads data into the physical sector allocated to the cache line according to the sector location determined in step 1010. Upon successful allocation, a two-bit rotation value is returned to the agent requesting cache line 422 so that the request agent knows how to subsequently access the allocated cache line 422. For tagged accesses, the rotation value can be returned by storing the two-bit rotation value within the tag itself. For transient requests, the rotation value can be returned by returning the two-bit rotation value along with the allocation index X, where the allocation index X identifies the corresponding cache line 422(X). Method 1000 proceeds to step 1002 as described above.
[0117] Returning to step 1006, if the cache line allocation can be combined with any previous allocation stored in allocation tracker 440, method 1000 proceeds to step 1014, where cache controller 470 allocates one or more sectors of the existing cache as described in conjunction with step 1008. In step 1016, cache controller 470 locates the logical sector for the cache line allocation based on certain factors and previous allocations. These factors are described in conjunction with step 1010. Furthermore, memory management locates the currently allocated sector such that the currently allocated sector is aligned with unused sectors of previously allocated sectors stored in allocation tracker 440. As a result, the current allocation and previous allocations are disjoint and can therefore share the same cache line.
[0118] In step 1018, the cache controller 470 loads data into the physical sector allocated to the cache line based on the sector location determined in step 1016. Upon successful allocation, a two-bit rotation value is returned to the agent requesting cache line 422 so that the request agent knows how to subsequently access the allocated cache line 422. For tagged accesses, the rotation value can be returned by storing the two-bit rotation value within the tag itself. For transient requests, the rotation value can be returned by returning the two-bit rotation value along with the allocation index X, where the allocation index X identifies the corresponding cache line 422(X). Method 1000 proceeds to step 1002 as described above.
[0119] In summary, sectorized cache memory in a computing system allows each cache line to be shared among multiple cache line allocations. Sectorized cache memory provides a mechanism for software applications to share portions of a cache line between two or more separate allocations. A first allocation may allocate one or more sectors of an empty and available cache line. If the first allocation results in one or more unused sectors, a second allocation may allocate one or more unused sectors of the cache line. If the second allocation also results in one or more unused sectors, an additional allocation may allocate one or more unused sectors of the cache line. Furthermore, if two allocations for the same cache line have overlapping logical sectors, the sectors of one of the allocations may be moved, for example, by a rotation function, to eliminate the overlap before being loaded into the physical cache memory. In this way, multiple allocations can share the same cache line in the cache memory.
[0120] At least one technical advantage of the disclosed technology over the prior art is that, using the disclosed technology, the cache memory can share a cache line with two or more allocations, where each allocation includes fewer sectors than the entire cache line. As a result, the cache memory can have fewer unused sectors compared to prior art that does not employ cache line sharing. This improves cache memory utilization, leading to improved cache memory performance and faster execution of software applications. These advantages represent one or more technical improvements over prior art methods.
[0121] Any and all combinations of any claim element recited in any way in any claim and / or any element described in this application fall within the scope of this disclosure and protection.
[0122] Various embodiments have been described for illustrative purposes, but are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments.
[0123] Aspects of this embodiment may be embodied as a system, method, or computer program product. Therefore, aspects of this disclosure may take the form of a completely hardware embodiment, a completely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, all collectively referred to herein as a "module" or "system." Furthermore, aspects of this disclosure may take the form of a computer program product embodied in one or more computer-readable media containing computer-readable program code.
[0124] Any combination of one or more computer-readable media may be used. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. For example, a computer-readable storage medium includes, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any suitable combination thereof. More specific examples (not an exhaustive list) of computer-readable storage media may include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable optical disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In the context of this document, a computer-readable storage medium can be any tangible medium that may include or store a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0125] Aspects of this disclosure have been described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to create a machine such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such processors can be, but are not limited to, general-purpose processors, special-purpose processors, application-specific processors, or field-programmable gate arrays.
[0126] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code comprising one or more executable instructions for implementing one or more specified logical functions. It should also be noted that in some alternative embodiments, the functions indicated in the blocks may not occur in the order shown in the figures. For example, two blocks shown consecutively may actually be executed substantially simultaneously, or these blocks may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified function or action or a combination of dedicated hardware and computer instructions.
[0127] While the foregoing description is directed to embodiments of this disclosure, other and further embodiments of this disclosure may be designed without departing from its essential scope, as defined by the appended claims.
Claims
1. A computer-implemented method for managing cache memory in a computing system, the method comprising: Detect the first cache line allocation request to allocate the first logical sector; It is determined that the first cache line allocation request can be combined with the second cache line allocation request to allocate the second logical sector; The first data associated with the first logical sector is stored in the first physical sector of the first cache line of the cache memory; as well as The first cache line tag is stored in the cache tag memory, and the first cache line tag is associated with the first logical sector that references the first cache line. The second data associated with the second logical sector is stored in the second physical sector of the first cache line, and The second cache line tag is stored in the cache tag memory and is associated with the second logical sector that references the first cache line.
2. The computer-implemented method of claim 1, wherein determining that the first cache line allocation request can be combined with the second cache line allocation request includes: It is determined that the first logical sector and the second logical sector do not overlap.
3. The computer-implemented method of claim 1, wherein determining that the first cache line allocation request can be combined with the second cache line allocation request includes: Determine that the first logical sector and the second logical sector overlap in the first cache line; as well as It is determined that the first logical sector can be moved to not overlap with the second logical sector.
4. The computer-implemented method of claim 1 further includes determining that the first cache line allocation request is a transient cache line allocation request.
5. The computer-implemented method of claim 1, wherein the first logical sector is allocated via a first tag group associated with the cache memory, and the second logical sector is allocated via a second tag group associated with the cache memory.
6. The computer-implemented method as described in claim 1, wherein: The first cache line allocation request and the second cache line allocation request are associated with thread bundle instructions executed as a set of wavefronts; The first logical sector is allocated via a first wavefront included in the set of wavefronts; and The second logical sector is allocated via a second wavefront included in the set of wavefronts.
7. The computer-implemented method of claim 1 further includes the feature that concurrent access to the first physical sector and the second physical sector does not cause memory data group conflicts.
8. The computer-implemented method of claim 1, wherein the first cache line comprises 128 bytes and the first physical sector comprises 32 bytes.
9. The computer-implemented method of claim 1, wherein the first cache line comprises four physical sectors, the four physical sectors including the first physical sector and the second physical sector.
10. The computer-implemented method of claim 1, wherein the cache memory includes a Level 1 L1 cache, a Level 1.5 L1.5 cache, or a Level 2 L2 cache.
11. The computer-implemented method of claim 1, wherein the first cache line allocation request is issued by a software application.
12. A system comprising: Cache memory; as well as A cache controller, coupled to the cache memory and configured to: Detect the first cache line allocation request to allocate the first logical sector; It is determined that the first cache line allocation request can be combined with the second cache line allocation request to allocate the second logical sector; The first data associated with the first logical sector is stored in the first physical sector of the first cache line of the cache memory; as well as The first cache line tag is stored in the cache tag memory, and the first cache line tag is associated with the first logical sector that references the first cache line. The second data associated with the second logical sector is stored in the second physical sector of the first cache line, and The second cache line tag is stored in the cache tag memory and is associated with the second logical sector that references the first cache line.
13. The system of claim 12, wherein, in order to determine that the first cache line allocation request can be combined with the second cache line allocation request, the cache controller is further configured to determine that the first logical sector and the second logical sector do not overlap.
14. The system of claim 12, wherein, in order to determine that the first cache line allocation request can be combined with the second cache line allocation request, the cache controller is further configured to: It is determined that the first logical sector and the second logical sector overlap in the first cache line; and It is determined that the first logical sector can be moved to not overlap with the second logical sector.
15. The system of claim 12, wherein the cache controller is further configured to determine that the first cache line allocation request is a transient cache line allocation request.
16. The system of claim 12, wherein the first logical sector is allocated via a first tag group associated with the cache memory, and the second logical sector is allocated via a second tag group associated with the cache memory.
17. The system of claim 12, wherein: The first cache line allocation request and the second cache line allocation request are associated with thread bundle instructions executed as a set of wavefronts; The first logical sector is allocated via a first wavefront included in the set of wavefronts; and The second logical sector is allocated via a second wavefront included in the set of wavefronts.
18. The system of claim 12, wherein the cache controller is further configured to concurrently access the first physical sector and the second physical sector does not cause memory data group conflicts.
19. The system of claim 12, wherein the first cache line comprises 128 bytes and the first physical sector comprises 32 bytes.
20. The system of claim 12, wherein the first cache line comprises four physical sectors, the four physical sectors including the first physical sector and the second physical sector.
Citation Information
Patent Citations
Reducing memory traffic in dram ECC mode
US20150012705A1
Coalescing texture access and load / store operations
US20150046662A1
Cache and associated method with frame buffer managed dirty data pull and high-priority clean mechanism
US8464001B1