Prefetch Invalidation of Memory Requests for Data Lacking Locality
By using a non-cache storage buffer to handle data lacking temporal and spatial locality and controlling data prefetching, the processor core addresses cache pollution issues, enhancing performance and reducing latency in memory-bound applications.
Patent Information
- Application Number
- JP2023518252
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-12-23
- Filing Date
- 2021-09-23
- Publication Date
- 2025-06-30
- Estimated Expiration
- 2041-09-23
AI Technical Summary
As semiconductor manufacturing advances and on-die geometric dimensions decrease, interconnect delay and high electrical impedance between chips increase latency, and memory-bound applications face performance limitations due to cache pollution from data lacking temporal and spatial locality.
The processor core employs a non-cache storage buffer to store data lacking temporal and spatial locality, preventing cache pollution by directing memory requests for such data directly to the non-cache buffer instead of the cache, and controlling data prefetching based on indicators within memory access instructions.
This approach reduces the penalty of cache pollution, improving performance by minimizing cache misses and power consumption, while allowing for efficient processing of memory requests lacking locality.
Smart Images

Figure 0007700219000001 
Figure 0007700219000002 
Figure 0007700219000003
Abstract
Description
Background Art
[0001] (Description of Related Art) As semiconductor manufacturing processes advance and the on-die geometric dimensions decrease, semiconductor chips containing one or more processing units provide more functions and performance. For example, a semiconductor chip may include one or more processing units. A processing unit may represent any of various data processing integrated circuits. Examples of processing units are general-purpose central processing units (CPUs), multimedia engines for audio / video (A / V) data processing, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), and parallel data processing engines such as graphics processing units (GPUs), digital signal processors (DSPs), etc. Processing units process instructions such as general-purpose instruction set architectures (ISAs), digital, analog, mixed signal, and radio-frequency (RF) functions.
[0002] However, in modern technologies in processing and integrated circuit design, design problems that can limit potential benefits still occur. One problem is that the interconnect delay per unit length continues to increase, and the high electrical impedance between individual chips also increases latency, and most software applications that access a lot of data are typically memory-bound in that the computation time is generally determined by the memory bandwidth. The performance of one or more computing systems depends on rapid access to stored data. Memory access operations include read operations, write operations, memory-to-memory copy operations, etc.
[0003] To address many of the above problems, one or more processing units use one or more levels of the cache hierarchy as part of the memory hierarchy to reduce the latency of memory access requests that access a copy of data in system memory for read or write operations. A typical memory hierarchy ranges from small, relatively fast volatile static memories, such as registers on a semiconductor die of a processor core and caches that are either located on the die or connected to the die, to larger volatile dynamic off-chip memories and relatively slow non-volatile memories. Generally, a cache stores one or more blocks, each of which is a copy of data stored at a corresponding address in system memory.
[0004] One or more caches are filled both by demand memory requests generated by the instructions being processed and by prefetch memory requests generated by a prefetch engine. The prefetch engine attempts to hide off-chip memory latency by detecting patterns of memory accesses to off-chip memory that could potentially stall the processor and by initiating memory accesses to off-chip memory prior to each requested instruction. The simplest pattern is a sequence of memory accesses that reference a consecutive set of cache lines (blocks) in a monotonically increasing or decreasing manner. In response to detecting a sequence of memory accesses, the prefetch unit begins to prefetch a specific number of cache lines prior to the currently requested cache line. However, since the cache has a finite size, the total number of cache blocks is inherently limited. Additionally, there is a limit to the number of blocks that map to a given set within a set-associative cache. In some cases, there are conditions that benefit from a limit on the number of cache blocks associated with a particular instruction type that is finer than the limits provided by cache capacity or cache associativity.
[0005] In view of the above, an efficient method and mechanism for efficiently processing memory requests are desired.
Brief Description of the Drawings
[0006]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Best Mode for Carrying Out the Invention
[0007] Although the present invention is susceptible to various modifications and alternative forms, specific embodiments are shown by way of example in the drawings and are described in detail herein. However, it should be understood that the drawings and their detailed description are not intended to limit the present invention to the particular forms disclosed, but on the contrary, the present invention is to cover all modifications, equivalents, and alternative forms falling within the scope of the present invention as defined by the appended claims.
[0008] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present invention. However, those skilled in the art should recognize that the present invention may be practiced without these specific details. In some instances, well-known circuits, structures, and techniques have not been shown in detail in order to avoid obscuring the present invention. Further, it will be understood that the elements shown in the figures are not necessarily drawn to scale for the sake of simplicity and clarity of the description. For example, the dimensions of some elements are exaggerated relative to other elements.
[0009] A system and method for efficiently processing memory requests are contemplated. The semiconductor chip may include one or more processing units. The processing unit may represent any of a variety of data processing integrated circuits. Examples of processing units are general-purpose central processing units (CPUs), multimedia engines for audio / video (A / V) data processing, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), and parallel data processing engines such as graphics processing units (GPUs), digital signal processors (DSPs), etc. The processing unit of the semiconductor chip includes at least one processor core and access to at least one cache (on-die or off-die), such as a level-one (L1) cache.
[0010] Also, the processor core uses a non-cache storage buffer that can store data that is prevented from being stored in the cache. In various embodiments, the processor core prevents storing specific data in the cache based on an indicator within the opcode of a memory access instruction that specifies that the requested data lacks one or more of temporal locality and spatial locality. The processor core prevents the requested data from being stored in the cache to reduce the penalty of cache pollution. Examples of these penalties are performance degradation, increased power consumption, and increased memory bus utilization. If requested data lacking one or more of temporal locality and spatial locality is stored in the cache, the number of cache misses increases. For example, if a particular set of a set-associative cache configuration is full, data having locality and already stored in the particular set is evicted to make room for the data. This is an example of cache pollution. Similarly, due to the recent placement of data within a cache set, other cacheable data within the cache set may be a candidate for earlier eviction than that data. This is another example of cache pollution.
[0011] By preventing data lacking one or more of temporal locality and spatial locality from being stored in a cache, the penalty of cache pollution is reduced. In various embodiments, one or more of a software designer, a compiler, or other software determines which memory access instructions target data lacking one or more of temporal locality and spatial locality. In one example, the compiler samples access counters to make this determination and then changes the opcodes of these memory access instructions to include an opcode that specifies that the requested data of the memory access instruction lacks one or more of temporal locality and spatial locality.
[0012] Rather than issuing to the cache memory access instructions that request data lacking one or more of temporal locality and spatial locality, the load / store unit (LSU) of a processor core issues these memory access instructions to a non-cache storage buffer that stores data that is prevented from being stored in the cache. As used herein, "memory request" is also referred to as "memory access request" and "access request". "Memory access request" includes "read access request", "read request", "load instruction", "write access request", "write request", "store instruction", and "snoop request".
[0013] As used herein, "data prefetch" or "prefetch" refers to the operation of prefetching data from a lower-level memory into one or more of a cache and a non-cache storage buffer. The prefetched data includes any of application instructions, application source data, intermediate data, and result data. In various embodiments, the circuitry of a processor core examines an indicator stored in a tag of a memory request that specifies whether data prefetch is to be prevented or permitted during the processing of this instance of the received memory request. In some embodiments, the tag of the received memory request includes this indicator. In various embodiments, if the tag indicates that data prefetch is to be prevented during the processing of this instance of the received memory request, other memory requests such as younger program-order memory requests having the same target address and other instances of this memory request that have not yet been processed may have data prefetch permitted for the target address. Each of the tags of those memory requests includes a flag indicating whether to permit or prevent data prefetch during the processing of those instances of the memory request. Thus, control of data prefetch is performed at the granularity of the instance of the memory request being processed, rather than at the granularity of the target address.
[0014] A prefetch engine can perform the operation of prefetching data, i.e., execute a data prefetch. As used herein, a "prefetch engine" is also referred to as a "prefetcher". The operation of prefetching includes performing one or more of prefetch training and generation of prefetch requests. For example, the prefetch engine can generate one or more memory access requests (prefetch requests) for data before the processor issues a demand memory access request for the prefetched data. In addition, the prefetch engine can perform prefetch training to determine whether to generate one or more prefetch requests based on the received target address.
[0015] During training, the prefetch engine identifies one or more of a sequence of memory accesses that reference a contiguous set of cache lines and a sequence of memory accesses that reference a set of cache lines separated by a stride between them. The prefetch engine can and is intended to be able to identify other types of memory access patterns. When the prefetch engine identifies a memory access pattern, the prefetch engine generates one or more prefetch requests based on the identified memory access pattern. In some embodiments, the prefetch engine stops generating prefetch requests if the prefetch engine determines that a demand memory access provided by a processor core contains a request address that does not match the identified memory access pattern.
[0016] The LSU also executes prefetch hint instructions. As used herein, a "prefetch hint instruction" is one of various types of memory access instructions inserted into an application's instructions to prefetch certain data into one or more of cache and non-cache data storage before a younger (in program order) demand memory access instruction requests the prefetched data. Examples of these prefetch hint instructions are the AMD64-bit instruction PREFETCHNTA directed to data lacking temporal locality, PREFETCH1 that prefetches data into level 1 (L1) cache, etc. However, due to the non-temporal indicators within certain prefetch hint instructions such as the AMD64-bit instruction PREFETCHNTA directed to data lacking temporal locality, the processor core processes this prefetch hint instruction differently from other types of prefetch hint instructions. In various embodiments, the circuitry of the processor core examines an indicator stored in a tag of the prefetch hint instruction that specifies whether data prefetching is prevented or permitted during the processing of this instance of the prefetch hint instruction. The indicator stored in the tag of the prefetch hint instruction is processed in a manner similar to that described above for memory requests targeting data lacking one or more of temporal locality and spatial locality.
[0017] Referring to FIG. 1, a generalized block diagram of one embodiment of a computing system 100 is shown. The computing system 100 includes processing units 110A-110B, an interconnect unit 118, a shared cache memory subsystem 120, and a memory controller 130 communicable with a memory 140. Clock sources such as a phase lock loop (PLL), an interrupt controller, a power manager, an input / output (I / O) interface, and devices are not shown in FIG. 1 to simplify the description. Note that the number of components of the computing system 100 and the number of sub-components shown in FIG. 1 within each of the processing units 110A, 110B may vary from embodiment to embodiment. There may be more or fewer components / sub-components than those shown for the computing system 100.
[0018] In one embodiment, the exemplary functions of the computing system 100 are incorporated in a single integrated circuit. For example, the computing system 100 is a system on chip (SoC) that includes multiple types of integrated circuits in a single semiconductor die. The multiple types of integrated circuits provide individual functions. In other embodiments, the multiple integrated components are individual dies within a package such as a system-in-package (SiP), a multi-chip module (MCM), or a chipset. In still other embodiments, the multiple components are individual dies or chips on a printed circuit board.
[0019] As shown, processing units 110A, 110B include one or more processor cores 112A, 112B and corresponding cache memory subsystems 114A, 114B. Processor cores 112A, 112B include circuitry for processing instructions. Processing units 110A, 110B represent any of a variety of data processing integrated circuits. Examples of processing units include general-purpose central processing units (CPUs), multimedia engines for audio / video (A / V) data processing, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), and parallel data processing engines such as graphics processing units (GPUs), digital signal processors (DSPs), etc. Each of processing units 110A, 110B processes instructions of a general-purpose instruction set architecture (ISA), digital, analog, mixed signal, and radio frequency (RF) functions, etc.
[0020] Processor cores 112A, 112B support simultaneous multithreading. Multiple threads executed by processor cores 112A, 112B share at least a shared cache memory subsystem 120, other processing units of different processing types (not shown), and I / O devices (not shown). Interconnection unit 118 includes one or more of a router, switch, bus, point-to-point connection, queue, and arbitration circuitry for routing packets between components. In some embodiments, interconnection unit 118 is a communication fabric. Packets include various types such as memory access requests, responses, commands, messages, and snoop requests. Interfaces 116A, 116B support communication protocols used to route data through interconnection unit 118 to shared cache memory subsystem 120, memory controller 130, I / O devices (not shown), power manager (not shown), and other components.
[0021] The address space of computing system 100 is divided among multiple memories. A memory map is used to determine which addresses map to which memories. In one embodiment, the coherence point of an address is memory controller 130 connected to memory 140 that stores data corresponding to the address. Memory controller 130 includes circuitry for interfacing with memory 140 and a request queue for queuing memory requests and memory responses. Although a single memory controller 130 and memory 140 are shown, in other embodiments, computing system 100 uses a different number of memory controllers and memories. In various embodiments, memory controller 130 receives memory requests from processing units 110A, 110B, schedules the memory requests, and issues the scheduled memory requests to memory 140.
[0022] Memory 140 is used as system memory within computing system 100. Memory 140 stores an operating system that includes a scheduler for allocating software threads to hardware within processing units 110A, 110B. Memory 140 also stores one or more of a hypervisor, basic input output software (BIOS) control functions, one or more applications that use an application programmer interface (API), a page table that stores mappings from virtual addresses to physical addresses and access permissions, source data and result data of an application, etc. Memory 140 uses any of various types of memory devices.
[0023] A copy of a portion of the data stored in the memory 140 is stored in one or more of the caches (114A, 114B, 120). The memory hierarchy transitions from relatively fast volatile memories such as registers on the processor die and caches that are either located on the processor die or connected to the processor die to non-volatile relatively slow memories. In some implementation examples, the faster volatile memory is considered to be at the top or highest level of the memory hierarchy, while the slower non-volatile memory is considered to be at the bottom or lowest level of the memory hierarchy. In these implementation examples, the first level of the memory hierarchy that is closer to the faster volatile memory of the hierarchy than the second level is considered to be at a "higher" level than the second level. In other implementation examples, the slower non-volatile memory is considered to be at the top or highest level of the memory hierarchy. Both ways of explaining the memory hierarchy are possible and contemplated, but in the following description, the faster volatile memory is considered to be at the top or highest level of the memory hierarchy. Thus, the higher levels of the memory hierarchy include faster volatile memories such as processor registers and level 1 (L1) local caches, while the lower levels of the memory hierarchy include non-volatile slower memories such as hard disk drives (HDDs) or solid-state drives (SSDs).
[0024] The cache memory subsystems 114A, 114B, 120 use high-speed cache memory to store blocks of data. In some embodiments, the cache memory subsystems 114A, 114B are integrated within their respective processor cores 112A, 112B. Alternatively, the cache memory subsystems 114A, 114B are connected to the processor cores 112A, 112B in a backside cache configuration or an in-line configuration as needed. In various embodiments, the cache memory subsystems 114A, 114B are implemented as a cache hierarchy. One or more levels of the cache hierarchy include a translation lookaside buffer (TLB) for storing mappings and access permissions from virtual addresses to physical addresses, a tag array, a data array, and a cache controller. One or more levels of the cache hierarchy also use a prefetch engine that can generate prefetch requests to fill the cache with data from a lower-level memory. The caches closer to the processor cores 112A, 112B (within the hierarchy) are integrated with the processor cores 112 as needed. In one embodiment, each of the cache memory subsystems 114A, 114B represents an L1 and L2 cache structure, and the shared cache subsystem 120 represents a shared L3 cache structure. Other cache configurations are possible and contemplated.
[0025] The circuits of processor cores 112A and 112B implement non-cache memory buffers 113A and 113B using one or more of the registers, one of the various flip-flop circuits, one of the various types of random access memory (RAM), and a content addressable memory (CAM). Processor cores 112A and 112B use non-cache memory buffers 113A and 113B that can store data that is prevented from being stored in the cache. Memory 140 stores one or more of high-performance computing (HPC) applications and data center applications that are irregular memory bandwidth bound, using data that caters to irregular memory access, such as data lacking one or more of temporal locality and spatial locality. HPC applications are used in, for example, computational fluid dynamics, high-density linear algebra libraries, climate and atmospheric modeling using a high-order method modeling environment (HOMME). Examples of data center applications are social network crawlers, cloud stream analysis applications, Internet map applications, and the like.
[0026] Rather than issuing memory requests for data that lacks one or more of temporal locality and spatial locality to caches such as the L1 caches of processor cores 112A and 112B, processor cores 112A and 112B issue these types of memory requests to non-cache storage buffers 113A and 113B. In various embodiments, non-cache storage buffers 113A and 113B use a prefetcher. The prefetcher of non-cache storage buffers 113A and 113B can generate a prefetch request for filling the non-cache storage buffers 113A and 113B with data from a lower-level memory. Similarly, the L1 caches of cache memory subsystems 114A and 114B use prefetcher 115A and 115B that can generate a prefetch request for filling the L1 caches with data from a lower-level memory. The prefetcher of computing system 100 can use various prefetching methods known to those skilled in the art.
[0027] The circuits of processor cores 112A and 112B inspect indicators such as flags stored in the tags of memory requests targeting data lacking one or more of temporal locality and spatial locality. The flag indicates whether to permit or prevent data prefetching based on the target address of the memory request during the processing of these instances of the memory request. The flag is a one or more bit field that specifies whether data prefetching is permitted or prevented. If the flag specifies preventing data prefetching during the processing of this instance of the received memory request, data prefetching is prevented during the processing of this instance of the received memory request. For example, the circuits of processor cores 112A and 112B prevent data prefetching by one or more of prefetcher 113A, 113B and prefetch engines 115A, 115B based on the target address of the received memory request during the processing of this instance of the received memory request. Similarly, the circuits of processor cores 112A and 112B inspect the flag stored in the tag of the prefetch hint instruction and execute data prefetching in the same manner based on the value stored in this flag.
[0028] Referring to FIG. 2, an embodiment of a general-purpose processor core 200 that performs out-of-order execution is shown. In one embodiment, processor core 200 simultaneously processes two or more threads within a processing unit such as either of processing units 112A and 112B (of FIG. 1). The functions of processor core 200 are implemented by hardware such as circuits. The instruction cache (i-cache) of block 202 stores instructions of a software application, and the corresponding instruction translation lookaside buffer (TLB) of block 202 stores the mapping from virtual address to physical address required to access the instructions. In some embodiments, the instruction TLB (i-TLB) stores access permissions corresponding to the address mapping.
[0029] The instruction fetch unit (IFU) 204 fetches a plurality of instructions from the instruction cache 202 every clock cycle if there is no miss in the instruction cache or the instruction TLB of the block 202. The IFU 204 includes a program counter that holds a pointer to the address of the next instruction to be fetched from the instruction cache 202, and this pointer is compared with the address mapping in the instruction TLB. Also, the IFU 204 includes a branch prediction unit (not shown) that predicts the result of a conditional instruction before the execution unit determines the actual result in a later pipeline stage.
[0030] The decoder unit 206 decodes the opcodes of the plurality of fetched instructions and allocates entries to an in-order retirement queue such as the reorder buffer 218, the reservation station 208, and the load / store unit (LSU) 220. In some embodiments, the decode unit 206 performs register renaming. In other embodiments, the reorder buffer 218 performs register renaming. In some embodiments, the decoder 206 generates multiple micro-operations from a single fetched instruction. In various embodiments, memory requests targeting data lacking at least one of temporal locality and spatial locality include indicators such as flags in a tag that specify whether to permit or prevent data prefetching based on the target address of the memory request during the processing of those instances of the memory request. In some embodiments, the decoder 206 inserts this flag into the tag of the decoded memory request. In various embodiments, the decoder 206 determines the value of the flag based on the opcode or other field of the received memory request. In one embodiment, when the flag is asserted, a prefetch based on the target address of the received memory request is executed in a later pipeline stage during the processing of this instance of the received memory request, but when the flag is negated, a prefetch based on the target address of the received memory request is prevented from being executed in a later pipeline stage during the processing of this instance of the received memory request. Note that in some designs, a binary logic high value is used as the asserted value and a binary logic low value is used as the negated value, while in other designs, a binary logic low value is used as the asserted value and a binary logic high value is used as the negated value. Other values indicating an asserted flag and a negated flag are also possible and contemplated. For example, in some embodiments, the flag includes two or more bits, and the flag can indicate that prefetching is permitted for one type of storage such as a non-cache memory buffer but is prevented for another type of storage such as a cache.Other combinations of the indicators are possible and contemplated.
[0031] Reservation station 208 functions as an instruction queue where instructions wait until the operands of the instructions become available. When the operands are available and the hardware resources are also available, the circuitry of reservation station 208 issues the instructions out of order to integer and floating point functional units 210 or to load / store unit 220. Functional unit 210 includes an arithmetic logic unit (ALU) for computer calculations such as addition, subtraction, multiplication, division, and square root. In addition, the circuitry within functional unit 210 determines the results of conditional instructions such as branch instructions.
[0032] Load / store unit (LSU) 220 receives memory requests such as load operations and store operations from one or more of decode unit 206 and reservation station 208. Load / store unit 220 includes a queue and circuitry for executing the memory requests. In one embodiment, load / store unit 220 includes verification circuitry to ensure that a load instruction receives the transferred data from the correct most recent store instruction. In various embodiments, load / store unit 220 uses non-cache memory buffer 222 having a function equivalent to non-cache memory buffers 113A, 113B (of FIG. 1). In various embodiments, the circuitry of load / store unit 220 implements non-cache memory buffer 222 using registers, any of various flip-flop circuits, any of various types of random access memory (RAM), or content addressable memory (CAM).
[0033] The non-cache memory buffer 222 can store data that is prevented from being stored in a cache such as the level 1 (L1) data cache (d-cache) of the block 230. In one embodiment, the non-cache memory buffer 222 stores data that lacks one or more of temporal locality and spatial locality. The prefetcher 224 can perform the operation of data prefetch as described above. In response to determining that a flag in the tag of the memory request specifies that a prefetch based on the target address of the memory request is prevented during the processing of this instance of the memory request, one or more of the circuits of the LSU 220 and the prefetcher 224 prevent a data prefetch based on the target address of the memory request during the processing of this instance of the memory request. In contrast, in response to determining that a flag specifies that a prefetch based on the target address of the memory request is permitted during the processing of this instance of the memory request, the prefetcher 224 performs the operation of data prefetch based on the target address of the memory request during the processing of this instance of the memory request.
[0034] The load / store unit 220 issues memory requests to the level 1 (L1) data cache (d-cache) of block 230. The L1 data cache of block 230 uses a TLB, a tag array, a data array, a cache controller, and a prefetch engine 232. The cache controller uses various circuits and queues such as a read / write request queue, a miss queue, a read / write response queue, a read / write scheduler, and a filter buffer. Similar to the prefetcher 224, the prefetch engine 232 can perform the operation of data prefetch as described above. In some embodiments, the prefetch engine 232 can generate a prefetch request to fill one or more of the L1 data cache of block 230 and the non-cache memory buffer 222 with data from a lower-level memory. In one embodiment, the prefetch engine 232 generates a prefetch request after monitoring some demand memory requests within an address range. The circuit of block 230 receives a demand memory request from the LSU 220.
[0035] For memory requests targeted at the non-cache memory buffer 222, although data is not stored in the L1 data cache of block 230, in some embodiments, the load / store unit 220 still sends an indicator of the memory request to the L1 data cache of block 230. To prevent performance degradation of the prefetch engine 232 due to training based on the target addresses of these types of memory requests, a flag is inserted into the tag of the memory request. As described above, the flag specifies whether to permit or prevent data prefetch based on the target address of the memory request during the processing of those instances of the memory request.
[0036] In response to determining that a flag specifies that prefetching based on the target address of a memory request is to be prevented during processing of this instance of the memory request, one or more of the circuitry of block 230 and the prefetch engine 232 prevent one or more steps of the operation of data prefetching based on the target address of this instance of the memory request. In other words, one or more of the circuitry of block 230 and the prefetch engine 232 prevent the prefetch engine 232 from generating a prefetch request based on the target address of this instance of the memory request and from performing prefetch training. In the case of a prefetch hint instruction, the use of the flag is similar to the use of the flag for memory requests targeting data lacking one or more of temporal locality and spatial locality.
[0037] In some embodiments, the processor core 200 includes a level-two (L2) cache 240 for processing memory requests from the L1 data cache 230 and the L1 instruction cache 202. The TLB of block 240 processes address mapping requests from the instruction TLB of block 202 and the data TLB of block 230. If the requested memory line is not found in the L1 data cache of block 230, or if the requested memory line is not found in the instruction cache of block 202, the corresponding cache controller sends a miss request to the L2 cache of block 240. If the requested memory line is not found in the L2 cache 240, the L2 cache controller sends a miss request to access memory in a lower-level memory such as a level-three (L3) cache or system memory.
[0038] The functional unit 210 and the load / store unit 220 present results on a common data bus 212. The reorder buffer 218 receives results from the common data bus 212. In one embodiment, the reorder buffer 218 is a first-in first-out (FIFO) queue that ensures in-order retirement of instructions according to program order. Here, an instruction that receives the result of an instruction is marked for retirement. When the instruction is at the head of the queue, the circuitry of the reorder buffer 218 sends the result of the instruction to the register file 214. The register file 214 holds the architectural state of the general-purpose registers of the processor core 200. Next, the instructions in the reorder buffer 218 retire in order, and the logic updates its queue head pointer to point to the subsequent instruction in program order. The results on the common data bus 212 are also sent to the reservation stations 208 to transfer values to the operands of the instructions waiting for the results. Multiple threads share the multiple resources within the core 200. For example, these multiple threads each share the blocks 202-240 shown in FIG. 2.
[0039] Note that in some embodiments, the processor core 200 implements a non-cache memory buffer 222 using individual buffers. For example, in one embodiment, the processor core 200 can use a "streaming load buffer" that stores data targeted by load instructions that target data lacking one or more of spatial locality and temporal locality, and an individual write-combining buffer that stores data targeted by store instructions that target data lacking one or more of spatial locality and temporal locality. These types of load instructions are also referred to as "streaming load instructions". Similarly, these types of store instructions are also referred to as "streaming store instructions".
[0040] Referring to FIG. 3, an embodiment of a method 300 for processing memory requests lacking locality is shown. For purposes of explanation, the steps in this embodiment (as well as in FIGS. 4 and 5) are shown in order. However, in other embodiments, some steps are performed in an order different from the illustrated order, some steps are executed simultaneously, some steps are combined with other steps, and some steps do not exist.
[0041] In various embodiments, the processing unit includes at least a processor core, a cache, and a non-cache storage buffer capable of storing data that is prevented from being stored in the cache. When processing a memory request targeted at the non-cache storage buffer, the processor core sends the memory request to the circuitry of the non-cache storage buffer. The circuitry of the non-cache storage buffer receives the issued memory request targeted at the non-cache storage buffer (block 302). In some embodiments, the non-cache storage buffer stores data that is not stored in the cache due to data targeted by a memory request lacking one or more of temporal locality and spatial locality. The access circuitry accesses the non-cache storage buffer indicated by the memory request (block 304). For example, the memory request indicates whether the memory access is a read access or a write access. The circuitry of the non-cache storage buffer examines the tag of the received memory request to determine whether to prevent or permit a data prefetch using the target address of the memory request during the processing of this instance of the memory request (block 306). For example, in one embodiment, the tag includes a flag that specifies whether to prevent or permit a data prefetch during the processing of this instance of the received memory request.
[0042] If the tag indicates that data prefetching is permitted during the processing of this instance of the received memory request (the "permit" in conditional block 308), the prefetch engine of the non-cache memory buffer generates one or more prefetch requests to prefetch data into the non-cache memory buffer based on the target address of the received memory request (block 310). In some embodiments, the target address of the memory request is sent to the prefetch engine, and the prefetch engine monitors some demand memory accesses within the address range and then initiates consecutive prefetching. This type of consecutive prefetching is stopped by the prefetch engine if the match to consecutive memory access operations transferred to the prefetch engine fails. In other embodiments, the prefetch engine automatically prefetches some cache lines based on the target address.
[0043] In some embodiments, one or more of the circuits of the load / store unit and the non-cache memory buffer send an indicator of the memory access of the non-cache memory buffer to the cache (block 312). The cache receives a flag indicating whether data prefetching is prevented or permitted during the processing of this instance of the received memory request. The cache does not store the data requested by the received memory request, but the prefetch engine of the cache can perform data prefetching or prefetch training based on the target address of the received memory request (request address). In various embodiments, if the flag indicates that prefetching is permitted during the processing of this instance of the received memory request, the prefetch engine of the cache uses the target address to perform one or more of generating a prefetch request based on the target address and training. For example, the prefetch engine performs monitoring of some demand memory accesses within the address range.
[0044] If the tag indicates to prevent data prefetch during the processing of this instance of the received memory request (the "prevent" in conditional block 308), data prefetch to the non-cache memory buffer and the cache is prevented during the processing of this instance of the received memory request (block 314). Other memory requests, such as memory requests in younger program order with the same target address and other instances of this memory request that have not yet been processed, may have data prefetch permitted for the target address. Each of the tags of those memory requests includes a flag indicating whether to permit or prevent data prefetch based on the target address of the memory request during the processing of those instances of the memory request. However, during the processing of this instance of the received memory request, data prefetch is prevented by an indicator specified by the flag stored in the tag of the received memory request. In some embodiments, the circuit of the non-cache memory buffer does not send the target address of the received memory request to either the prefetch engine of the non-cache memory buffer or the prefetch engine of the cache. In one embodiment, the circuit of the non-cache memory buffer does not send any information of the received memory request to the cache. In other embodiments, the non-cache memory buffer sends information of the received memory request, including a flag specifying whether prefetch is permitted or prevented during the processing of this instance of the received memory request, to the cache. In such embodiments, one or more circuits among the cache controller and the prefetch engine of the cache determine whether to permit or prevent generating a prefetch request and perform prefetch training based on the target address based on the received flag.
[0045] Referring to FIG. 4, an embodiment of a method 400 for processing memory requests lacking locality is shown. The cache receives an index of a memory access of a non-cache storage buffer (block 402). In various embodiments, the memory access corresponds to a memory request targeting a non-cache storage buffer. The cache bypasses the read or cache write processing of the received memory access (block 404). The cache circuitry examines the tag of the received memory access to determine whether to prevent or permit data prefetching using the target address of the received memory access during the processing of this instance of the memory access (block 406).
[0046] If the tag indicates that data prefetching is to be permitted during the processing of this instance of the received memory access ("permit" in conditional block 408), the cache prefetch engine performs one or more of generating one or more prefetch requests based on the target address of the received memory access and performing prefetch training (block 410). However, if the tag indicates that data prefetching is to be prevented during the processing of this instance of the received memory access ("prevent" in conditional block 408), the cache prefetch engine prevents performing one or more of generating one or more prefetch requests based on the target address of the received memory access and performing prefetch training (block 412). Other memory accesses, such as memory accesses corresponding to younger program-order memory requests having the same target address and other instances of this memory access that have not yet been processed, may have data prefetching permitted for the target address. However, during the processing of this instance of the received memory access, data prefetching is prevented by an index specified by a flag stored in the tag of the received memory access.
[0047] Note that in some embodiments, the circuitry that supports the non-cache memory buffer is external to the cache. Thus, in one embodiment, when the received memory request stores an indicator that specifies that data prefetching is to be prevented during the processing of this instance of the received memory request, the non-cache memory buffer prevents the target address (request address) from being sent to the cache. In other embodiments, the cache prefetch engine indicates that the flag prevents data prefetching during the processing of this instance of the received memory access, but the cache controller still generates one or more prefetch requests based on the requested address of the received memory access when providing a value for the least recently used (LRU) value used in a cache line replacement policy that limits the amount of time that the prefetched data exists in the cache. Alternatively, the cache controller restricts the placement of the prefetched data to a particular way of a multi-way cache configuration.
[0048] Referring to FIG. 5, an embodiment of a method 500 for processing memory requests lacking locality is shown. The circuitry of a unit of a processor, such as a load / store unit (LSU) or other unit of the processor, receives a prefetch hint instruction requesting data lacking locality (block 502). An example of a prefetch hint instruction is the AMD64-bit instruction PREFETCHNTA directed to data lacking temporal locality. Other examples are possible and contemplated. Due to an indicator within the prefetch hint instruction specifying that the requested data lacks locality, the processor, unlike other types of load instructions, fetches data for this prefetch hint instruction. In some embodiments, the processor fetches data from system memory and stores the fetched data at a particular level of the cache memory subsystem. For example, the level 2 (L2) cache stores a copy of the fetched data, while the level 1 (L1) cache and level 3 (L3) cache are bypassed if used. Also, the processor prefetches a particular amount of data based on the target address of the prefetch hint instruction. For example, in some cases, the processor fetches one or more cache lines based on the target address of the prefetch hint instruction. The processor does not store a copy of the fetched data or the prefetched data at any level of the cache memory subsystem if these types of data are later evicted from any cache. Other types of processing data for this prefetch hint instruction are possible and contemplated. By using a flag within the tag of the prefetch hint instruction, data prefetch as described above can be prevented.
[0049] The processor executes the read access indicated by the received prefetch hint instruction (block 504). The processor fetches the data stored at the memory location pointed to by the target address of the prefetch hint instruction. The processor stores a copy of the fetched data in one or more of the cache and the non-cache storage buffer (block 506). Since the specific level of the cache is based on design requirements, in some examples, the L1 cache located closest to the processor core does not store a copy of the fetched data, while the L2 cache stores a copy of the fetched data. Other storage arrangements are possible and contemplated. In some embodiments, when the prefetched data is stored in the cache, the corresponding cache controller sets restrictions on the storage time and / or storage location of the data targeted by the prefetch hint instruction in the cache data array. For example, the cache controller provides a value of the longest time unused (LRU) value used in the cache line replacement policy that limits the amount of time the prefetched data exists in the cache. Also, the cache controller can restrict the placement of the prefetched data to a specific way of the multi-way set associative cache configuration.
[0050] The processor examines the tag of the prefetch hint instruction to determine whether to prevent or permit data prefetch using the target address of the prefetch hint instruction during the processing of this instance of the prefetch hint instruction (block 508). If the tag indicates to permit data prefetch during the processing of this instance of the prefetch hint instruction ("permit" in conditional block 510), one or more prefetch engines among the cache and non-cache storage buffers generate one or more prefetch requests based on the target address of the prefetch hint instruction (block 512). However, if the tag indicates to prevent data prefetch during the processing of this instance of the prefetch hint instruction ("prevent" in conditional block 510), the processor prevents one or more prefetch engines among the cache and non-cache storage buffers from generating a prefetch request based on the target address of the prefetch hint instruction (block 514). In some embodiments, the processor prevents prefetch by preventing any information of the prefetch hint instruction from being sent to one or more prefetch engines among the cache and non-cache storage buffers.
[0051] Note that one or more of the above-described embodiments include software. In such embodiments, program instructions implementing the method and / or mechanism are transmitted or stored on a computer-readable storage medium. A number of types of media configured to store program instructions are available, including hard disks, floppy (registered trademark) disks, CD-ROMs, DVDs, flash memories, programmable ROM (Programmable ROM, PROM), random access memory (RAM), and various other forms of volatile or non-volatile storage devices. Generally speaking, a computer-accessible storage medium includes any storage medium that is accessible by a computer during use to provide instructions and / or data to the computer. For example, computer-accessible storage media include storage media such as magnetic or optical media (e.g., disks (fixed or removable), tapes, CD-ROMs, DVD-ROMs, CD-Rs, CD-RWs, DVD-Rs, DVD-RWs, Blu-Ray (registered trademark), etc.). Storage media further include volatile or non-volatile memory media such as RAM (e.g., synchronous dynamic RAM (synchronous dynamic RAM, SDRAM), double data rate (double data rate, DDR, DDR2, DDR3, etc.) SDRAM, low-power DDR (low-power DDR, LPDDR2, etc.) SDRAM, Rambus DRAM (Rambus DRAM, RDRAM), static RAM (static RAM, SRAM), etc.), ROM, flash memory, etc., and non-volatile memory (e.g., flash memory) accessible via a peripheral interface such as a Universal Serial Bus (Universal Serial Bus, USB) interface. Storage media include microelectromechanical systems (microelectromechanical system, MEMS), as well as storage media accessible via communication media such as networks and / or wireless links.
[0052] In addition, in various embodiments, the program instructions include an operational level description of the hardware functionality or a register-transfer level (RTL) description in a high-level programming language such as C, or a design language (HDL) such as Verilog, VHDL, or a database format such as the GDSII stream format (GDSII). In some cases, the description is read by a synthesis tool that synthesizes the description to generate a netlist that includes a list of gates from a synthesis library. The netlist includes a set of gates that also represent the functionality of the hardware including the system. The netlist can then be placed and routed to generate a data set that describes the geometric shapes to be applied to the mask. The mask can then be used in various semiconductor manufacturing steps to generate a semiconductor circuit or circuits corresponding to the system. Alternatively, the instructions on the computer-accessible storage medium are, as necessary, a netlist (with or without a synthesis library) or a data set. In addition, the instructions are utilized for emulation by a hardware-based type of emulator from vendors such as Cadence®, EVE®, and Mentor Graphics®.
[0053] Although the above embodiments have been described in considerable detail, many modifications and variations will become apparent to those skilled in the art once the above disclosure is fully understood. The following claims are intended to be construed to embrace all such modifications and variations.
Claims
1. An apparatus comprising: a non-cache memory buffer; and a circuit, wherein the circuit is configured to: receive a first memory request targeted at the non-cache memory buffer; prevent data prefetching to the non-cache memory buffer based at least in part on the first memory request; receive a second memory request targeted at the non-cache memory buffer; determine that the second memory request includes an indicator specifying that data prefetching using the target address of the first memory request is permitted; permit data prefetching to the non-cache memory buffer based at least in part on the second memory request; and be configured to perform the above.
2. The apparatus according to claim 1, wherein the circuit is configured to prevent data prefetching to the non-cache memory buffer based at least in part on the first memory request in response to determining that the first memory request includes a first indicator specifying that data prefetching using the target address of the first memory request is prevented. The apparatus of claim 1.
3. The apparatus according to claim 2, wherein the circuit is configured to prevent data prefetching to a cache using the target address of the first memory request based at least in part on the determination that the first memory request includes the first indicator. The apparatus of claim 2.
4. The apparatus according to claim 1, wherein the first memory request includes an indicator indicating that the first memory request lacks one or more of temporal locality and spatial locality. The apparatus of claim 1.
5. The apparatus according to claim 1, wherein the circuit is configured to: receive a third memory request; and set a limit on one or more of the storage time and storage location in the cache for data fetched from a lower-level memory to the cache during processing of the third memory request. and be configured to perform the above.
6. The apparatus according to claim 5, wherein the circuit is configured to determine that the third memory request is a prefetch hint instruction including a third indicator specifying setting of the limit. The apparatus of claim 5.
7. A method, comprising: a circuit receiving a first memory request targeted at a non-cache memory buffer; the circuit preventing data prefetching to the non-cache memory buffer based at least in part on the first memory request; receiving a second memory request targeting the non-cache memory buffer; determining that the second memory request includes an indicator specifying that data prefetching using the target address of the first memory request is permitted; permitting data prefetching to the non-cache memory buffer based at least in part on the second memory request; and a method. **Claim 8** in response to determining that the first memory request includes a first indicator specifying that data prefetching using the target address of the first memory request is prevented, the circuit preventing data prefetching to the non-cache memory buffer based at least in part on the first memory request; The method of claim 7. **Claim 9** The method includes preventing data prefetching to a cache using the target address of the first memory request in response to determining that the first memory request includes the first indicator. The method of claim 8. **Claim 10** receiving a third memory request; and during processing of the third memory request, setting a limit on one or more of a storage time and a storage location in the cache for data fetched from a lower-level memory into the cache. The method of claim 7. **Claim 11** including determining that the third memory request is a prefetch hint instruction including a third indicator specifying setting the limit. The method of claim 10. **Claim 12** a processing unit, comprising an interface configured to communicate with a lower-level memory; a cache; and a processor core including the apparatus of any one of claims 1 to 6. a processing unit.
Citation Information
Patent Citations
Prefetching method and device
JP2000148584A
Image processor and cache memory
JP2001331793A
Cash memory device
JP2001344152A
Information processor and prefetch method
JP2004240811A
Memory interface circuit
JP2005018553A