Prefetch coordination between a first cache and a second cache of a processor
Patent Information
- Application Number
- PCT/US2025/020417
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2026-09-24
Smart Images

Figure US2025020417_24092026_PF_FP_ABST
Abstract
Description
Atty Docket No. 0120-1141WO1PREFETCH COORDINATION BETWEEN A FIRST CACHE AND A SECOND CACHE OF A PROCESSORBACKGROUND
[0001] Data caching is commonly used to improve the performance (e.g., latency performance, throughput performance, etc.) of computer processors such as central processing units (CPUs). For example, data caching may be performed by storing data that has been identified as relevant for ongoing program execution in a data store that is more easily and quickly accessible to a processor core performing the program execution. For example, relevant data may be copied or moved from a relatively large and slow main memory to a relatively small and fast memory that the processor core can access more quickly. The small and fast memory may be referred to as a data cache (or as simply a cache) and the relevant data being cached may include instruction data (representing upcoming instructions that are to be executed) or other suitable data (e.g., data referenced by such instructions) as may be predicted to be useful to the program execution in the near future. The identifying and caching of data in this way is referred to herein as prefetching the data to the cache.SUMMARY
[0002] A processor (e.g., a central processing unit (CPU)) may be configured to prefetch data from main memory to one or more caches in a multi-level caching architecture. For example, a multi-level caching architecture may refer to an architecture with at least a first cache level to which a processing core has direct access, a second cache level to which one or more caches on the first cache level have direct access, and a memory larger than caches on any of the other levels and that is at least one more level removed from the core than the second cache level. Conventional prefetch techniques used in multi-level caching architectures suffer from drawbacks such as, at least, cache pollution and coordination challenges (suffered by hierarchical prefetch techniques), and impractical request tracking and prefetch overhead due to memory latency times (suffered by flat prefetch techniques). To address the challenges with conventional techniques, methods and systems for prefetch coordination between a first cache and a second cache of a processor are described herein. As will be detailed below, a hybrid of hierarchical and flat prefetch techniques may be used toiAtty Docket No. 0120-1141WO1mitigate the drawbacks mentioned above and provide effective and efficient data prefetching.
[0003] To this end, for example, a method may be performed by a processor that has a core, a first cache (e.g., a Level- 1 or LI cache to which the core has direct access), a second cache (e.g., a Level-2 or L2 cache to which the first cache has direct access), and a connection to memory (e.g., whereby the second cache may access desired data from the memory). In this example method, the second cache may: 1) receive, from the first cache, a prefetch request for a data workload to be cached within the first cache. In response to the prefetch request and a determination that the data workload is not cached in the second cache, the method may further include: 2) indicating, by the second cache to the first cache, that the data workload is not cached in the second cache, and 3) obtaining, by the second cache, the data workload from the memory coupled to the processor. In response to obtaining the data workload, the method may further include: 4) indicating, to the first cache by the second cache, that the data workload is available for request by the first cache. In this way, as detailed below, the first cache may successfully gain access to the data workload but may do so on its own terms and in its own time so that the coordination challenges are mitigated, pollution in the second cache need not be propagated, and the timing associated with request tracking by the first cache avoids excessive latency issues.
[0004] Another example implementation described herein involves a processor that implements prefetch coordination between a first cache and a second cache in accordance with principles described herein. For example, the processor may include: 1) a first cache configured to send a prefetch request for a data workload to be cached within the first cache; and 2) a second cache coupled to the first cache and coupled to a memory. The second cache in this example may be configured to perform a process similar to the method described above. For example, the second cache may be configured to: 1) receive the prefetch request from the first cache; 2) in response to the prefetch request and a determination that the data workload is not cached in the second cache: (a) indicate, to the first cache, that the data workload is not cached in the second cache, and (b) obtain the data workload from the memory; and 3) in response to obtaining the data workload, indicate, to the first cache, that the data workload is available for request by the first cache. Here again, as detailed below, the indication to the first cache that the data workload is available for request may allow the first cache to successfully gain access to the data workload on its own terms and with its own timing, thereby mitigating coordination challenges, cache pollution, overhead-related problems caused by excessive latency delays, and other such issues described herein.
[0005] Yet another example implementation described herein involves a cache withinAtty Docket No. 0120-1141WO1a processor (e.g., an L2 cache corresponding to the second cache described in relation to implementation examples above). For instance, the cache may include: 1) a data store configured for caching data obtained from a memory coupled to the processor; 2) a first interface by which the cache communicates with an additional cache within the processor (e.g., an LI cache corresponding to the first cache described in relation to implementation examples above); 3) a second interface by which the cache communicates with the memory; and 4) a cache controller configured to: (a) receive, from the additional cache by way of the first interface, a prefetch request for a data workload to be cached within the additional cache; (b) in response to the prefetch request and a determination that the data workload is not cached in the cache: (i) indicate, to the additional cache by way of the first interface, that the data workload is not cached in the cache, and (ii) obtain, by way of the second interface, the data workload from the memory; and (c) in response to obtaining the data workload, indicate, to the additional cache by way of the first interface, that the data workload is available for request by the additional cache.
[0006] Other implementations may perform similar functions as described above and / or may use other types of hardware to perform the functions. Certain implementations may involve systems, devices, media, and / or combinations of these, that employ processors such as described herein (e.g., processors configured to implement prefetch coordination between a first cache and a second cache of the processor in accordance with principles described herein).
[0007] The details of these and other implementations are set forth in the accompanying drawings and the description below. Other features will also be made apparent from the following description, drawings, and claims.BRIEF DESCRIPTION OF THE DRAWINGS
[0008] FIG. 1 shows an illustrative implementation of prefetch coordination between a first cache and a second cache of a processor in accordance with principles described herein.
[0009] FIG. 2 shows an illustrative processor including a first cache and a second cache and configured to implement prefetch coordination between the first cache and the second cache in accordance with principles described herein.
[0010] FIG. 3 shows an illustrative cache configured for use within a processor implementing prefetch coordination between this cache and an additional cache in accordance with principles described herein.Atty Docket No. 0120-1141WO1
[0011] FIG. 4 shows an illustrative method for prefetch coordination between a first cache and a second cache of a processor in accordance with principles described herein.
[0012] FIG. 5A shows a first example scenario in which a data workload is prefetched to a first cache in accordance with principles described herein.
[0013] FIG. 5B shows a second example scenario in which a data workload is prefetched to the first cache and to a second cache in accordance with principles described herein.
[0014] FIG. 6A shows illustrative dataflow diagrams to contrast a hierarchical prefetch technique with a hybrid prefetch technique in accordance with principles described herein.
[0015] FIG. 6B shows illustrative dataflow diagrams to contrast a flat prefetch technique with the hybrid prefetch technique in accordance with principles described herein.
[0016] FIG. 7 shows an illustrative computing device that may be used to implement various devices and / or systems described herein.DETAILED DESCRIPTION
[0017] Methods and devices (e.g., computer hardware) relating to prefetch coordination between a first cache and a second cache of a processor are described herein. More particularly, as detailed below, a hybrid prefetch technique featuring benefits of both hierarchical prefetch techniques and flat prefetch techniques, but without the drawbacks of either one, is performed to cache data workloads effectively and efficiently.
[0018] Computer processors (e.g., central processing units (CPUs), as well as other types of processors) may be configured to execute instructions and process data at extremely fast rates. However, as fast as processors may be configured to operate, a technical problem arises when data that is to be processed (e.g., instructions that are to be executed, other data that is to be manipulated in accordance with the instructions, etc.) is not accessible to the processor at the same rate at which the data can be processed. If not addressed, the memory access time would create an enormous bottleneck for program execution, since, for example, it may take significantly more time (e.g., in certain cases, multiple orders of magnitude more time) for the processor to fetch a particular instruction from main memory than it takes for the processor to execute that instruction.
[0019] At least one technical solution to this technical problem is for the processor to use data caching techniques to store data that is anticipated to soon be used by the processor in small data stores, referred to as caches or data caches, that are accessible to the processorAtty Docket No. 0120-1141WO1with much less latency. For example, caches may be significantly smaller in capacity than the main memory (e.g., kilobytes or megabytes, rather than the gigabytes typical of modern computer memories), and may be located physically closer to the processing core than the memory is (e.g., fabricated on a same die as the processing core or packaged within a same chip so as to be located within millimeters of the core, rather than being fabricated on a different chip and located on a mainboard several centimeters or more away). Moreover, caches may use faster, albeit less dense and more complex and expensive, data storage technologies to decrease the data access time. For instance, data caches may be implemented using static random-access memory (SRAM) technologies, while main memory may be implemented by dynamic random-access memory (DRAM) technologies.
[0020] In some implementations, multiple levels of caches may be organized in a hierarchy, with a fastest and smallest level (level one or LI) being closest to the processing core, and progressively slower and larger levels (level two or L2, level three or L3, etc.) being located further from the core and closer to main memory (which will be understood to be even slower and larger than all the levels of caches). If it can be anticipated that program execution will require the processor to access a particular data workload, it may be desirable for the data workload to be preemptively moved from memory or from a higher-level cache to a lower-level (and faster) cache closer to the processor. A data workload as referred to herein may represent any data that may be requested (or anticipated / expected to be requested) for use by a core of a processor, (e.g., a core of a processor that includes a cache such as an LI cache associated with the core, an L2 cache shared by multiple cores, etc.). In other words, a data workload, as referred to herein, may be considered to refer to a specific instance of data expected to be required for program execution and having a certain address in memory, rather than being limited to a specific data type or data content. Examples of data workloads may include units of data such as: a group of instructions, a subroutine or function, a software application or component thereof, a file, a segment of a stream, a data structure such as an array or linked list, or the like. This prospective analysis of what data will be of use to the processor and the resulting caching of that data before the data is actually called for is referred to as prefetching.
[0021] While prefetching data within a multi-level caching architecture such as described above can be very effective in increasing processor performance (e.g., executing more instructions per unit of time, etc.), conventional prefetching techniques suffer from a variety of technical problems.
[0022] As a first example, a hierarchical prefetch technique could be used to prefetchAtty Docket No. 0120-1141WO1data to a first cache (e.g., an LI cache nearest to the processor core) by first moving the data through the other levels of cache (e.g., an L2 cache, an L3 cache, etc.) hierarchically. In this example, any data that is to be cached in the first cache (the LI cache) must first be cached, for example, in a second cache (an L2 cache). This is advantageous for the first cache, since the second cache would generally be better situated to interface with memory (or higher-level caches) and handle the relatively long amount of time it takes to access data from memory. For example, if the first cache can access the desired data workload from the second cache, there may be a much more manageable amount of overhead (e.g., tracking the outstanding request for a relatively short amount of time while waiting for the response from the second cache) than if the first cache had to wait the relatively long amount of time for memory to respond to a prefetch request. However, as convenient as this hierarchical approach may be for the first cache, a technical problem of cache pollution arises for the second cache as it, too, caches all the data that is desired not only for the first cache, but also for other LI caches associated with other processing cores. If the data workload is only needed by the first cache, it may be inefficient for the second cache to also store the data.
[0023] Additionally, a coordination problem may arise between the first cache and the second cache when the second cache initially misses on the data request (indicating that it is not currently storing the requested workload) and has to retrieve the workload from memory. Specifically, the first cache may request that the second cache retrieve the desired data but may not know exactly when to check back in to see if the data has been obtained. As a result, the first cache may resend its request to the second cache too early (before the requested data has been provided by memory) or too late (after the second cache has been using resources to store it and possibly after it has been thrashed by another LI cache or has been evicted to store different data). This coordination problem could be addressed by careful tuning of a prefetch algorithm, but this may require significant complexity in the design and would still not guarantee perfect coordination between the first and second caches.
[0024] As another example, a flat prefetch technique could be used to prefetch data to the first cache directly, without having to also cache the data workload in the other levels of the hierarchy along the way. In other words, in this example, a data workload could be obtained from memory and cached in the first cache without first being cached in the second cache and / or any other cache levels as may be used in a particular implementation. This would be advantageous for the second cache, since the data workload would now not pollute the second cache and use its limited resources when only the first cache actually needs the data. However, the overhead associated with the first cache accessing memory directly mayAtty Docket No. 0120-1141WO1be prohibitive. For example, the first cache may need to track a request that will not be satisfied for several hundred clock cycles as memory is accessed. The first cache may not be as well-situated to handle such delays as the second cache would be. For example, a request queue managed by the first cache may need to be far larger than is practical for a first-level cache in order to track requests with delays on the order of magnitude of hundreds of clock cycles.
[0025] These and other technical problems associated with conventional prefetch techniques may be elegantly solved by implementations for prefetch coordination between a first cache and a second cache of a processor described herein. For example, at least one technical solution to the problem of cache pollution in the second cache may involve a hybrid prefetch technique that brings the desired data workload directly into the first cache without any requirement that it also be cached in the second cache. At least one technical solution to the technical problems of complexity and excessive overhead, then, may involve the hybrid prefetch technique including an indication, delivered from the second cache to the first cache in response to the data workload being received from memory at the second cache, that the data workload is being held in reserve for the first cache to request when it is ready to request it. For example, a stashing snoop signal may be used for the purpose of delivering this indication. In this way, the first cache may maintain a very manageable overhead burden, only needing to track requests to the second cache that take a few clock cycles (rather than requests to memory that may take hundreds of clock cycles). Additionally, at least one technical solution to the problems of complexity, L1 / L2 coordination, and prefetch algorithm tuning is provided by the direct indication, from the second cache, that the data workload is ready for the first cache to request when ready. Rather than the first cache needing to guess when to re-request the workload after a cache miss to the second cache (e.g., based on a tuned algorithm, etc.), the first cache may receive the indication that the data is ready and may, at its own convenience (e.g., when cache space is available, when other relevant data has been obtained, etc.), obtain the data from the second cache.
[0026] Various beneficial technical effects may result from prefetch coordination implementations described herein between a first cache and a second cache of a processor. For example, processor performance (e.g., latency performance, throughput performance, etc.) may be improved in all the ways that might be expected from an effective caching scheme, due at least in part to data having been prefetched for easy and fast access when the core is ready to use the data. Unlike with conventional prefetch techniques described above (e.g., hierarchical techniques, flat techniques, etc.), however, these benefits may be obtainedAtty Docket No. 0120-1141WO1without the detrimental side effects that have been described. Moreover, the increased effectiveness and efficiency made possible by prefetch coordination implementations described herein may help the processor improve other aspects of its performance beyond the speed at which software instructions are executed (e.g., reducing the cache footprint needed on the chip, reducing power consumption, increasing battery life, reducing heat output, etc.).
[0027] Various implementations will now be described in more detail with reference to the figures. It will be understood that particular implementations described below are provided as non-limiting examples and may be applied in various situations. Additionally, it will be understood that other implementations not explicitly described herein may also fall within the scope of the claims set forth below. Prefetch coordination between a first cache and a second cache of a processor in the ways described herein may result in any or all of the technical effects mentioned above, as well as various additional and / or alternative technical effects and benefits that will be described and / or made apparent below.
[0028] FIG. 1 shows an illustrative implementation 100 of prefetch coordination between a first cache and a second cache of a processor in accordance with principles described herein. Specifically, as shown, various components of a processor 102 are shown to be in communication with one another within the processor. For example, a core 104 is shown to be communicatively coupled with a first cache 106. The first cache 106 is shown to be communicatively coupled with a second cache 108. The second cache 108 is shown to be communicatively coupled with a memory 110. While memory 110 is illustrated as being separate from processor 102 (i.e., unlike the other components, memory 110 does not fall under the bracket representing processor 102), memory 110 is shown to be coupled to processor 102. Specifically, in the case of implementation 100, memory 110 is shown to be coupled to second cache 108.
[0029] Core 104 may represent a processing component of processor 102. In other words, core 104 may be configured to access data (e.g., by loading data from first cache 106), and, based on that data, to execute instructions of various types. For any type of instruction, data representing the instruction may be loaded from memory to indicate to the core what computation is to be performed. For example, an ADD instruction stored in memory may, when loaded into core 104 for execution, cause an arithmetic logic unit (ALU) of core 104 to add certain values together. In certain cases, one or more of these values to be added together may also be stored in memory, in which case a LOAD instruction may be performed prior to the ADD to get access to the value that is to be added. A load / store unit (LSU) associated with core 104 may work closely with first cache 106 to both perform LOAD instructions (toAtty Docket No. 0120-1141WO1access data from memory), as well as to perform STORE instructions (to write processed data back to memory as appropriate.
[0030] As shown, core 104 may request a particular data workload 112 from first cache 106, which, as indicated parenthetically, may represent an LI Data Cache of processor 102. This data access, as mentioned above, may be referred to as a LOAD instruction and, assuming data workload 112 is available in first cache 106, may take on the order of about one clock cycle or a similarly small number of clock cycles (“Cycles: ~10°”). As indicated, core 104 and first cache 106 may be associated with a first clock domain (“Clock Domain 1”) that is synchronized by a first clock signal having a first clock frequency. In other words, the core 104 and the first cache 106 may communicate using the first clock signal. Additionally, as mentioned above, core 104 and first cache 106 may be physically close to one another and, indeed, may be implemented on a same chip (e.g., a processor system on a chip (SoC)) or a same die within a processor chip.
[0031] Similarly, first cache 106 may request data workload 112 from second cache 108, which, as indicated parenthetically, may represent an L2 Data Cache of processor 102. Assuming data workload 112 is available in second cache 108, this data access may may take on the order of tens of clock cycles (“Cycles: -101”), or about an order of magnitude longer than the amount of time it takes to load from the LI cache (first cache 106) to the core. As indicated, first cache 106 and second cache 108 may also be associated with the first clock domain (“Clock Domain 1”) synchronized by the first clock signal having the first clock frequency mentioned above. Additionally, as mentioned above, first cache 106 and second cache 108 may be physically close to one another, such as by being implemented on a same chip (e.g., a processor system on a chip (SoC)) or a same die.
[0032] Second cache 108 may request data workload 112 from memory 110, which may represent off-chip storage, such as by DRAM storage that is separate from processor 102 (e.g., implemented on different chips located apart from processor 102 on the mainboard) and capable of storing very large amounts of data in a dense and less expensive manner. While data workload 112 may be guaranteed to be available in memory 110 (since, unlike caches 106 and 108, memory 110 is configured to continuously store all the data used by processor 102), accessing data directly from memory 110 may take on the order of hundreds of clock cycles (“Cycles: ~102”), or about an order of magnitude longer than the amount of time it takes to load from the L2 cache (second cache 108) to the LI cache (first cache 106).Additionally, as indicated, second cache 108 and memory 110 may be associated with a second clock domain (“Clock Domain 2”) that is synchronized by a second clock signalAtty Docket No. 0120-1141WO1having a second clock frequency that is lower than the first clock frequency. In other words, second cache 108 and memory 110 may be configured to communicate using the second (slower) clock signal rather than the first clock signal. Accordingly, it will be understood that estimated numbers of clock cycles that may pass in association with certain operations described herein may refer to the (shorter) clock cycles of the first clock domain on which the processor core operates, rather than to the (longer) clock cycles of the second clock domain.
[0033] Data workload 112 may represent any suitable data that may be requested for use by core 104. For example, data workload 112 could represent a single cache line (the smallest unit of data that the caches 106 and 108 are configured to store) of a suitable size such as 32 bytes, 64 bytes, 128 bytes, or the like. As another example, data workload 112 could represent another data structure (e.g., a structure of at least two cache lines or more). For instance, as mentioned above, data workload 112 could be implemented as a software application or a component thereof (e.g., a group of instructions, a subroutine or function, etc.), a data structure (e.g., a file, an array, a linked list, etc.), a segment of a data stream (e.g., an incoming media data stream, etc.), or another suitable data instance. Data workload 112 may include data that may be loaded by core 104 to either execute (e.g., instruction data) or to use in the execution of instruction data.
[0034] In some examples, data fetching may be driven by the immediate demands of core 104 as it executes instructions. For instance, core 104 may request to load data (e.g., a particular instruction or value) included within data workload 112. If data workload 112 is not being cached by first cache 106, first cache 106 will return an indication to core 104 that the cache missed (i.e., that the requested data is not available within that cache). First cache 106 may then be configured to attempt to retrieve the data from second cache 108 if the data happens to be cached there.
[0035] As shown by increasing numbers of clock cycles for the data movement and increasingly tall memory spaces depicted for the caches and memory units in FIG. 1, each storage unit farther away from core 104 may store more data but may also take longer to access. For example, it may take on the order of one clock cycle (“Cycles: ~10°”) for data workload 112 to be obtained by core 104 from first cache 106 if data workload 112 happens to be present there, on the order of tens of clock cycles (“Cycles: -101”) for data workload 112 to be obtained by first cache 106 from second cache 108 if data workload 112 happens to be present there, and on the order of hundreds of clock cycles (“Cycles: ~102”) for data workload 112 to be obtained by second cache 108 from memory 110 if the data is not yet cached at L2. In one implementation, for instance, it could take about 2 cycles to loadAtty Docket No. 0120-1141WO1requested data from first cache 106 to core 104, about 20 cycles to load requested data from second cache 108 to first cache 106, and about 300 cycles to load requested data from memory 110 to second cache 108. However, with these longer delay times comes the advantage of larger and larger data stores. For example, while first cache 106 may be on the order of kilobytes, second cache 108 may be on the order of megabytes, and memory 110 may be on the order of gigabytes in certain implementations.
[0036] While not explicitly shown in FIG. 1, each cache will be understood to include a certain amount of data storage (e.g., fast SRAM storage) along with logic (e.g., a cache controller) configured to perform communications described herein (e.g., requesting data from other units, loading and storing data as requested, managing limited resources such as by evicting less relevant data from the cache, cache coherence operations, etc.). Similarly, memory 110 will be understood to represent not only a relatively large amount of data storage (e.g., gigabytes of DRAM storage in certain implementations), but also to represent a memory controller configured to load from and store to the storage space, as well as to otherwise maintain the stored data (e.g., refreshing the DRAM cells as appropriate, etc.)
[0037] While core 104 and each cache 106 and 108 is capable of requesting particular data (e.g., data workload 112) on demand as the data becomes necessary, processor 102 may also include components configured to analyze data usage and, based on preconfigured rules and algorithms (as well as observed patterns, fine tuning of the algorithms, etc.), configured to attempt to anticipate what data will be called for in the near future. Based on these types of determinations, data such as data workload 112 may be prefetched to the caches even before core 104 requests to load the data contained therein. As a result of such prefetching, the desired data may be ready to go the first time core 104 attempts to load it from first cache 106. More particularly, as shown, a “Prefetch to L2” refers to a loading of data from memory 110 to second cache 108, while a “Prefetch to LI” refers to a loading of data from second cache 108 to first cache 106. While algorithmic details associated with such prefetching (e.g., which data is to be prefetched to which cache, timing of when certain data is to be prefetched given the pace of execution of processor 102, etc.) are largely outside the scope of the present disclosure, it will be understood that a prefetching algorithm may attempt to balance an objective of keeping relevant data close at hand when it is to be used and an objective to conserve the limited resources of the various levels of cache. For example, if first cache 106 is capable of only storing a few kilobytes of data, a prefetch algorithm may be very careful to only store the most relevant data within the cache at any given time (and to avoid cache pollution, in which less relevant data fills the cache at the expense of more relevant data). AsAtty Docket No. 0120-1141WO1another example, if second cache 108 is capable of serving a plurality of cores, each with their own LI data cache (as will be described in more detail below), the limited resources of this cache, too, may be used very cautiously to avoid pollution and ensure efficiency.
[0038] FIG. 1 has been described to introduce certain concepts and components that will be used in the remainder of this disclosure. For example, as will be described, certain prefetch techniques described herein may be used to prefetch data from memory 110 to first cache 106 with advantages and benefits that other prefetch techniques lack. These prefetch techniques will now be described in more detail, as well as contrasted with conventional prefetch techniques that do not enjoy all the same beneficial technical effects, with reference to FIGS. 2-4, 5A-5B, and 6A-6B.
[0039] FIG. 2 shows an illustrative implementation of processor 102 that includes multiple implementations of core 104 (e.g., a core 104-1, a core 104-2, and additional cores 104), respective implementations of first cache 106 (the LI cache) associated with each core 104 (e.g., a cache 106-1 associated with core 104-1, a cache 106-2 associated with core 104-2, and additional LI caches 106 associated with additional cores 104), and an implementation of second cache 108 (the L2 cache) configured to serve all of the LI caches 106 and all the cores 104. As mentioned above, each cache 106, as well as cache 108, may include a static random-access memory (SRAM) data cache. Processor 102 is also shown to be communicatively coupled (e.g., by way of second cache 108 in particular) to memory 110, which is illustrated in FIG. 2 as being separate from processor 102 and which, as mentioned above, may include one or more dynamic random-access memory (DRAM) module.
[0040] Processor 102 may be configured to implement prefetch coordination between LI and L2 caches in accordance with principles described herein. For example, a first cache of processor 102 (e.g., cache 106-1 in one example) may be configured to send a prefetch request for a data workload (e.g., data workload 112, not explicitly shown) to be cached within the first cache. A second cache coupled to the first cache and coupled to memory (e.g., cache 108, which is coupled both to cache 106-1 and memory 110) may then be configured to receive the prefetch request from the first cache. In response to the received prefetch request and a determination that the data workload is not cached in the second cache (i.e., a cache miss by cache 108), the second cache may be configured to indicate, to the first cache, that the data workload is not cached in the second cache (i.e., indicate the cache miss with an acknowledgement message communicated to the first cache). The prefetch request and the cache miss may also cause the second cache to perform a memory read in which the requested data workload is obtained from memory. In response to obtaining the dataAtty Docket No. 0120-1141WO1workload (which, as has been described, may not be completed until on the order of hundreds of clock cycles later), the second cache may indicate, to the first cache, that the data workload is available for request by the first cache. For example, this indication could be implemented by sending a stashing snoop signal from the second cache to the first cache, as will be described in more detail below.
[0041] In response to the indication that the data workload is available and when the first cache is ready to cache the data, the first cache may send, to the second cache, an additional prefetch request for the data workload to be cached within the first cache. Only then, when the data workload is being held at the second cache and when the first cache is ready to receive it, may the data be transferred. More specifically, based on this additional prefetch request, the first cache may receive the data workload from the second cache.Thereafter, the data workload has been successfully prefetched to the first LI cache (e.g., cache 106-1), such that data from the data workload is now readily available for use by core 104-1. Accordingly, cache 106-1 may be further configured to receive, from the core (e.g., core 104-1) and subsequent to the receiving of the data workload from second cache 108, a load request for data included within the data workload. For instance, if the data workload includes instruction data representing a plurality of instructions, the data identified by the load request may be a particular instruction or subset of instructions from the plurality of instructions. In response to the load request, the first cache may then send the requested data included within the data workload to the core.
[0042] While cache 108 was tasked, in the example above, with obtaining the data workload from memory 110 and holding the data workload temporarily for cache 106-1 while it waited for cache 106-1 to respond to the indication that the data workload is available for request, it is noted that second cache 108 did not necessarily cache the data workload in the manner required by a conventional hierarchical prefetch technique. This distinction will be detailed further below but is noted now to point out the beneficial technical effect that second cache 108 may not be polluted by the data workload if it is not determined to be useful outside the context of its use within cache 106-1 by core 104-1. It will be understood, on the other hand, that if the data workload were not considered to be polluting to second cache 108 (e.g., if core 104-2 or another one of the additional cores 104 also was anticipated to need the data such that it may soon be prefetched to cache 106-2 or one of the additional LI caches 106), the cache 108 could not only store the data workload in a temporary buffer to hold it for cache 106-1 but could also properly cache the workload for access by any of the LI caches 106.Atty Docket No. 0120-1141WO1
[0043] Another beneficial technical effect of the prefetch coordination performed by processor 102 relates to the ability of cache 106-1 to re-request the data workload only when both: 1) the data workload is known to be available from cache 108 (e.g., held in reserve specifically for use by cache 106-1), and 2) cache 106-1 is itself prepared to receive the data workload. It is noted that, in certain implementations, cache 108 could push the data along with or instead of the indication that the data workload is available for request by cache 106-1. However, this may introduce further complexities since it would require that cache 106-1 be ready at any time to receive the data (e.g., including making room for the data, etc.). By sending the indication and allowing the first cache to re-request the data workload on its own terms and with its own timing (e.g., when it is ready), a situation is advantageously avoided in which the first cache is not ready to accept the pushed data workload and the channel is stalled and a bottleneck is created as the first cache works to free up space.
[0044] Yet another beneficial technical effect of the prefetch coordination performed by processor 102 relates to overhead associated with caching requests such as the prefetch request for the data workload to be cached within the first cache. As shown in connection with each cache 106, a respective request queue 202 (e.g., a request queue 202-1 for cache 106-1, a request queue 202-2 for cache 106-2, additional request queues 202 for additional LI caches 106, etc.) may be used to track outstanding requests that have been made by the respective caches. For example, each request queue 202 may be implemented by a miss status holding register (MSHR) or other such bookkeeping or request-tracking mechanism as may serve a particular implementation. Each request queue 202 may be implemented by a table or other structure within the cache (e.g., managed by the cache controller) that keeps track of cache misses and corresponding memory requests while the requests are outstanding.
[0045] To this end, each request queue 202 may include sufficient storage to hold several entries that each represent a particular cache miss in the LI cache. An example entry could include, for instance, a memory address that was requested (i.e., the memory address of the requested data workload), a type of request (e.g., read or write), and / or other associated data (e.g., metadata, etc.). By tracking various cache misses in different entries, a request queue 202 may allow a given core 104 to handle multiple cache misses concurrently, which may help improve performance of the core by allowing the core to avoid stalling while waiting for a particular piece of data. Along with this basic tracking function, certain request queue implementations may further include advanced features such as merging entries pointing to the same target (e.g., outstanding requests for the same cache line), enabling out-of-order execution to allow additional instructions to be executed by the core while waitingAtty Docket No. 0120-1141WO1for data from a cache miss, and so forth.
[0046] As with other resources of processor 102, however, the resources associated with a given implementation of request queue 202 may be scarce and should be applied judiciously. For example, while a given request queue 202 implementation may have sufficient capacity (e.g., a sufficient number of entries) to keep track of outstanding requests for the latency required to obtain data from cache 108 (e.g., on the order of tens of clock cycles, as described above), it may not be practical or desirable for the request queue 202 implementation to be endowed with sufficient capacity to keep track of outstanding requests for the latency required for obtaining data all the way from memory 110 (e.g., on the order of hundreds of clock cycles, as described above). Accordingly, at least one beneficial technical effect of the prefetch coordination described in relation to processor 102 stems from the fact that the request queue 202 only needs to have sufficient capacity for two requests to the L2 cache (e.g., on the order of a few tens of clock cycles each) and not sufficient capacity for a request all the way to memory (e.g., on the order of a few hundreds of clock cycles).
[0047] More particularly, and as will be described in more detail below, cache 106-1 may be configured to: 1) allocate, within request queue 202-1 and in response to cache 106-1 sending the prefetch request, a first entry for the prefetch request; and 2) deallocate, within request queue 202-1 and in response to the indicating that the data workload is not cached in the second cache, the first entry for the prefetch request. Thereafter, when the indication that the data workload is available for request by cache 106-1 is received and triggers the additional prefetch request for the data workload to be prefetched to cache 106-1, cache 106- 1 may be configured to: 3) allocate, within request queue 202-1 and in response to the sending of the additional prefetch request, a second entry for the additional prefetch request; and 4) deallocate, within request queue 202-1 and in response to the receiving of the data workload from cache 108, the second entry for the additional prefetch request. In both of these cases, there would only be on the order of tens of clock cycles between the allocation and deallocation of the entry within the request queue 202-1. This is far more manageable and practical than a single entry being held in the request queue for hundreds of clock cycles while a flat prefetch request for a data workload stored in memory 110 is performed.
[0048] It is noted that the L2 cache 108 may include a request queue (not explicitly shown in FIG. 2) similar to the request queues 202 of the LI caches 106. This request queue may be built with sufficient capacity to handle the long latency time of accessing memory 110 since the L2 cache is more tailored to this purpose (e.g., utilizing a slower clock signal, etc.). Advantageously, this request queue would also not need to be duplicated to multipleAtty Docket No. 0120-1141WO1caches (as request queues 202 are).
[0049] As shown, cache 108 may be an L2 cache configured to serve a plurality of LI caches 106 (i.e., cache 106-1, cache 106-2, and additional LI caches 106). More particularly, along with the first cache 106-1 and the second cache 108 described as interacting in the example above, processor 102 is shown to further include a third cache that implements another LI cache for another core. Specifically, just as core 104-1 is associated with the first cache 106-1, core 104-2 may be associated with a third cache 106-2, and both the first cache 106-1 and the third cache 106-2 (i.e., the individual LI caches) may be configured to access memory 110 by way of the second cache 108 (i.e., the shared L2 cache).
[0050] As indicated by the dashed-line tag (“See FIG. 3”) in FIG. 2, the second cache 108 (i.e., the L2 cache) is illustrated in more detail in FIG. 3. More particularly, FIG. 3 shows an illustrative implementation of cache 108 configured for use within processor 102 to implement prefetch coordination techniques described herein between this L2 cache and a plurality of LI caches (caches 106).
[0051] As shown in this more detailed implementation, cache 108 of processor 102 may include: 1) a data store 302 configured for caching data obtained from memory 110 coupled to processor 102; 2) a first interface 304-1 by which cache 108 communicates with one or more additional caches 106 (e.g., LI caches) within processor 102; 3) a second interface 304-2 by which cache 108 communicates with memory 110; and 4) a cache controller 306 configured to perform a method 400 that, as indicated by the dashed line tag attached thereto (“See FIG. 4”), will be described in more detail in relation to FIG. 4.
[0052] As will be described in more detail below, certain operations of method 400 that may be performed by cache controller 306 include: 1) receiving, from an additional cache 106 (e.g., any one of the LI caches 106) and by way of first interface 304-1, a prefetch request for a data workload to be cached within the additional cache 106; 2) in response to the prefetch request and a determination that the data workload is not cached in cache 108 (i.e., is not being held within data store 302): (a) indicating, to the additional cache 106 by way of first interface 304-1, that the data workload is not cached in cache 108, and (b) obtaining, by way of second interface 304-2, the data workload from memory 110; and 3) in response to obtaining the data workload, indicating, to the additional cache 106 and by way of first interface 304-1, that the data workload is available for request by the additional cache 106. Once the data workload is obtained and the additional cache 106 is ready to receive the data workload (e.g., in response to the indicating that the data workload is available), additional operations may be performed by cache 108, by the additional cache 106, and by aAtty Docket No. 0120-1141WO1core associated with the additional cache 106. These operations will be further described below in relation to method 400 and FIGS. 5A-5B and 6A-6B.
[0053] Data store 302 may represent a primary store of data managed by cache 108 and which will be understood to store data that is cached by cache 108 (i.e., indexed and stored for access by any of the LI caches). In contrast, a temporary buffer 308 shown to be included in cache 108 will be understood to represent a secondary store of data managed by cache 108 that may store data that is not cached for access by any of the LI caches, but, rather, is being temporarily held in reserve for a particular LI cache that requested it. For example, after cache 108 obtains such a data workload from memory 110 (e.g., likely several hundred cycles after the original request was made, and after the LI cache has likely moved on to other tasks) and after cache 108 has indicated to the requesting LI cache that the data workload is now available for request (e.g., using a stashing snoop signal, or the like), the data workload could be stored in temporary buffer 308. More particularly, the data workload could be held in reserve for the LI cache within temporary buffer 308 until an additional prefetch request is received from the LI cache (e.g., in response to the stashing snoop signal) to indicate that the LI cache is again ready to receive the data workload. At that point (once the LI cache is ready for it), the data workload being stored in temporary buffer 308 may be sent to the LI cache.
[0054] As will be described and illustrated in more detail below with reference to method 400 (in FIG. 4) and certain example scenarios (in FIGS. 5A-5B), there are multiple possibilities for how data from memory 110 may be handled once cache 108 obtains it. For example, the data workload could be: 1) held in reserve within temporary buffer 308 until it is requested by an LI cache such as described above; 2) cached within data store 302 where any of the LI caches may request access to it (e.g., at any time until it is evicted and no longer being cached); 3) stored in both data store 302 and temporary buffer 308; 4) pushed directly to the LI cache that originally requested it (potentially several hundred clock cycles ago) without waiting for the additional prefetch request from the LI cache; or 5) handled using a combination of the above or in other suitable ways. The way that a memory-obtained data workload is handled by cache 108 may be determined based on the prefetch request itself (e.g., whether the request is only for the data workload to be cached in the LI cache or whether it also requests for the data workload to be cached in the L2 cache), based on certain characteristics of cache 108 itself (e.g., whether the L2 cache is configured to be inclusive or non-inclusive of the LI caches it serves), or based on other suitable factors.
[0055] As used herein, an L2 cache is inclusive of an LI cache if all the data cachedAtty Docket No. 0120-1141WO1by the LI cache is also cached by the L2 cache. In other words, an inclusive L2 cache is configured to maintain a copy of all the data stored in each of the LI caches that share use of the L2 cache. If L2 cache 108 is configured to be inclusive of LI caches 106, temporary buffer 308 may not be implemented, or the data it stores may be duplicative of data also cached in data store 302. However, if L2 cache 108 is configured to be non-inclusive of LI caches 106, it will be understood that temporary buffer 308 may be used to reduce cache pollution within L2 cache 108. For example, if it is determined that a requested workload is to be delivered to a particular LI cache (for use by a particular core) but that no other LI cache (or corresponding core) is likely to load or prefetch the data workload in the near term, it may be considered wasteful or inefficient to cache the data workload at the L2 cache 108 (i.e., within data store 302) as well. For example, the data workload could be unwanted pollution within the L2 cache in this scenario. Accordingly, implementing temporary buffer 308 may help reduce this pollution and allow the data workload to be omitted from data store 302 to conserve caching resources for other data that may be more relevant.
[0056] FIG. 4 shows an illustrative method 400 for prefetch coordination between a first cache and a second cache of a processor in accordance with principles described herein. Method 400 may be performed by a processor (e.g., an implementation of processor 102) or, more particularly, by certain caches (e.g., cache controllers) included therein. For example, as shown by dashed line boxes labeled “LI” and “L2” in FIG. 4, certain operations of method 400 (i.e., operations 402-414) may be performed by an LI cache (e.g., an implementation of one of caches 106 described above) while other operations of method 400 (i.e., operations 416-426) may be performed by an L2 cache (e.g., an implementation of cache 108 described above). While FIG. 4 shows illustrative operations according to a specific implementation, it will be understood that other implementations of this method may omit, add to, reorder, and / or modify any of operations 402-426 that are explicitly represented in FIG. 4.Additionally, while operations 402-426 are illustrated with arrows suggestive of a sequential order of operation, it will be understood that one or more of the operations of method 400 may be performed concurrently (e.g., in parallel) with one another.
[0057] Each of the operations of method 400 will now be described in more detail. More particularly, the operations 402-414 that are performed by an LI cache (e.g., first cache 106) to accomplish prefetch coordination described herein will be described first, followed by the operations 416-426 that are performed by the L2 cache (e.g., second cache 108). As shown in method 400, arrows between different operations are shown to suggest the movement of data and / or possible dependencies between operations being performed. TheseAtty Docket No. 0120-1141WO1data movements and dependencies will be further described below in connection with the associated operations.
[0058] At operation 402, the LI cache may send a prefetch request to the L2 cache. For example, an algorithm used by a cache controller of the LI cache or by a prefetch unit associated with the LI cache and / or with a load / store unit (LSU) of the processor may determine that access to a certain data workload is likely to soon be needed by the core, such that it would be desirable for the data workload to be cached in the LI cache. Accordingly, the prefetch request sent at operation 402 may be for the data workload to be cached within the LI cache (“Data workload to LI”).
[0059] At operation 404, the LI cache may allocate an entry for the prefetch request within a request queue associated with the LI cache. For example, in response to the LI cache sending the prefetch request to the L2 cache, the LI cache may allocate an entry within a request queue such as a request queue 202 that is associated with the LI cache (e.g., request queue 202-1 of for cache 106-1, etc.). As has been described, the allocation of this entry may allow the LI cache to track the outstanding prefetch request that was sent at operation 402 until either the requested data workload is received (in the event of an L2 cache hit), or, alternatively, until an acknowledgement is provided by the L2 cache that the prefetch request is received and will be carried out in due course (in the event of an L2 cache miss).
[0060] At operation 406, the LI cache may deallocate the entry for the prefetch request within the request queue. For method 400, it will be understood that an L2 cache miss was identified (i.e., a determination was made by the L2 cache that the requested data workload is not currently cached within the L2 cache) and an indication of the miss was provided. Accordingly, in response to the indicating to the LI cache that the data workload is not cached in the L2 cache, the entry for the prefetch request may be deallocated so as to not use resources within the request queue for the relatively long period it will take for the L2 cache to obtain the data workload from memory. As mentioned above, this deallocation of the entry provides an efficiency advantage over a flat prefetch technique in which the entry is maintained for (potentially) hundreds of clock cycles as main memory (e.g., DRAM of memory 110) is accessed.
[0061] At operation 408, the LI cache may send an additional prefetch request to the L2 cache. For example, a relatively long time (e.g., several hundred clock cycles in one example) after the indication of the L2 cache miss was received and the entry for the first prefetch request was deallocated at operation 406, an indication (e.g., a stashing snoop signal such as described in more detail below) may be received from the L2 cache. In response toAtty Docket No. 0120-1141WO1this indicating (by the L2 cache) that the data workload is available (e.g., by the stashing snoop signal or other suitable communication), the LI cache may send the additional prefetch request to the L2 cache. The additional prefetch request may, again (like the original prefetch request sent at operation 402), be for the data workload to be cached within the LI cache. The difference is that now the LI cache has been notified that the data is ready (e.g., being cached in data store 302 and / or being held in temporary buffer 308) so that there will not be another L2 cache miss.
[0062] While not explicitly shown in FIG. 4, it will be understood that, in response to the sending of the additional prefetch request at operation 408, the LI cache may again (similarly to operation 404) allocate an entry within the request queue associated with the LI cache. In this case, the entry is for the additional prefetch request. As in the example of operation 404, the entry is not expected to tie up request queue resources for longer than a few tens of clock cycles while a response is awaited from the L2 cache (as opposed to the conventional flat prefetch technique in which the entry could tie up those resources for at least an order of magnitude longer while memory is accessed).
[0063] At operation 410, the LI cache may receive the data workload from the L2 cache. For example, the data workload may be provided (e.g., either from the data store of the L2 cache or from a temporary buffer holding the data workload in reserve) to the LI cache based on (i.e., in response to) the additional prefetch request. Similarly as described above in relation to operation 408, it will be understood that, in response to the receiving of the data workload at operation 410, the LI cache may again (similarly to operation 406) deallocate the entry for the additional prefetch request within the request queue. Again, this deallocation may occur relatively shortly after the allocation at operation 408 (e.g., within a few tens of clock cycles and sooner than several hundreds of clock cycles), thereby using the request queue resources efficiently.
[0064] At operation 412, sometime subsequent to the receiving of the data workload at operation 410, the LI cache may receive, from a core of the processor (e.g., a core 104 to which the particular LI cache corresponds, a load request for data included within the data workload. The load request may identify the entire data workload or a particular portion of the data workload at this stage, in accordance with what is needed for program execution. For example, if the data workload includes instruction data representing a set of instructions, the data identified in the load request could be a single instruction or subset of instructions from the set of instructions. As another example, if the data workload includes a data structure such as an array or linked list, the data identified in the load request could be one or severalAtty Docket No. 0120-1141WO1entries from the data structure.
[0065] At operation 414, the LI cache may send the data included within the data workload (i.e., the data requested at operation 412) to the core that requested it. For instance, this providing of requested data may be performed in response to the load request received at operation 412.
[0066] As mentioned above, operations 402-414 of method 400 may be performed by the LI cache, while operations 416-426, which are also shown to be included within method 400, are performed by the L2 cache. As indicated by some of the arrows between the LI cache and the L2 cache in FIG. 4, certain operations performed by the LI cache may relate to operations performed by the L2 cache. For example, two operations may refer to the same transaction from two different perspectives (e.g., one cache sending data that the other cache is receiving, etc.), one operation performed by one cache may trigger or allow for the performance of another operation by the other cache, or the like. Each of operations 416-426 of method 400 will now be described as these operations may be performed by the L2 cache.
[0067] At operation 416, the L2 cache may receive a prefetch request for a data workload to be cached within the LI cache. Specifically, as indicated by the arrow from operation 402 to operation 416, the L2 cache may receive at operation 416, the prefetch request described above to be sent by the LI cache.
[0068] At operation 418, the L2 cache may indicate to the LI cache that the data workload (i.e., the data workload requested to be cached in the LI cache by the prefetch request received at operation 416) is not cached in the L2 cache. Operation 418 may be performed in response to the prefetch request (received at operation 416) and further in response to a determination that the data workload is not cached in the L2 cache (i.e., that there is an L2 cache miss). While this determination is not explicitly shown in method 400, it will be understood that part of operation 416 or 418 (or a separate operation performed between these two) may involve determining whether the requested data workload is cached in the second cache (i.e., if there is an L2 cache hit or an L2 cache miss on the prefetch request). As shown by the arrow from operation 418 to operation 406, the indication of the L2 cache miss at operation 418 may trigger the deallocation of the entry from the request queue described above in relation to operation 406.
[0069] At operation 420, the L2 cache may obtain the data workload from a memory coupled to the processor (e.g., memory 110, as described above). As with operation 418, operation 420 may be performed in response to the prefetch request (received at operation 416) and further in response to the determination (described above in relation to operationAtty Docket No. 0120-1141WO1418) that the data workload is not cached in the L2 cache (i.e., the determination of the L2 cache miss). In some examples, operations 418 and 420 may be performed concurrently, such as by operation 420 being initiated (e.g., the data workload being requested from memory) prior to the indication of operation 418 being sent. As operation 420 may take several hundred clock cycles of the clock domain in which the LI cache operates (as has been described), however, operation 418 may at least be completed (if not started and completed) prior to the completion of operation 420 as the data workload is received from memory.
[0070] As indicated by a dashed-line tag attached to operation 420 in FIG. 4, FIGS.5 A and 5B illustrate additional aspects of the obtaining of the data workload performed at operation 420. More particularly, FIG. 5A shows a first example scenario 500-A in which a data workload is prefetched to a first cache (i.e., the LI cache 106-1), while FIG. 5B shows a second example scenario 500-B in which a data workload is prefetched to both the first cache (LI cache 106-1) and to a second cache (L2 cache 108) in accordance with principles described herein. Each of FIGS. 5 A and 5B will now be described before returning to the description of method 400 in FIG. 4.
[0071] In FIG. 5A, a requested data workload 112 (i.e., an implementation of the data workload 112 described above) is shown to be provided from memory 110 (“From Memory”) and to be passed through cache 108 (the L2 cache) to ultimately be prefetched to, and cached at, cache 106-1 (one of the LI caches served by the L2 cache). As described above in relation to FIG. 3, there are various options or scenarios for how an L2 cache such as cache 108 may handle a data workload obtained from memory before the workload is sent on to the LI cache (e.g., cache 106-1 in this example).
[0072] One such option is illustrated by scenario 500-A in FIG. 5A. In this example, the obtaining of the data workload from the memory at operation 420 includes: 1) requesting, by the L2 cache (i.e., cache 108), the data workload 112 from the memory; 2) receiving, by the L2 cache, the data workload 112 from the memory (e.g., several hundred clock cycles later in certain implementations); and 3) in response to receiving the data workload, storing the data workload in temporary buffer 308 of second cache 108. As described above, temporary buffer 308 may be configured to hold data workload 112 in reserve for the first cache (i.e., cache 106-1) and may be separate from the data store 302 of second cache 108 (which may be configured for caching data obtained from the memory).
[0073] As shown in scenario 500-A, the data workload 112 goes straight from memory to temporary buffer 308 of cache 108 to cache 106-1 without ever being cached in cache 108 (i.e., indexed and stored in data store 302) where it could also be accessed byAtty Docket No. 0120-1141WO1another LI data cache corresponding to another core of the processor. As such, if first cache 106-1 is associated with a first core of the processor (e.g., core 104-1, not shown in FIGS. 5A-5B) and a third cache 106-2 of the processor is associated with a second core of the processor (e.g., core 104-2, also not shown in FIGS. 5A-5B), both the first cache and the third cache (i.e., LI caches 106-1 and 106-2) may be configured to access the memory by way of the second cache (i.e., L2 cache 108). Accordingly, in scenario 500-A, if an additional prefetch request from the third cache (i.e., LI cache 106-2) to the second cache (i.e., L2 cache 108) for the data workload 112 is sent and received while the data workload 112 is stored in temporary buffer 308 of cache 108, cache 108 may indicate to cache 106-2 that the data workload is not cached in cache 108. In other words, even though there is a copy of data workload 112 in temporary buffer 308, that copy of the data workload is, in this example, being held in reserve specifically for LI cache 106-1. It is thus not available to other LI caches (such as cache 106-2) and an L2 cache miss would be indicated to cache 106-2 in this scenario.
[0074] In contrast, alternative options for how the obtaining of data workload 112 at operation 420 may be handled are illustrated by scenario 500-B in FIG. 5B. In this example, the obtaining of the data workload is performed in a similar way as described above in relation to scenario 500-A. For instance, as in scenario 500-A, the obtaining may start by: 1) requesting, by the L2 cache (i.e., cache 108), the data workload 112 from the memory; and 2) receiving, by the L2 cache, the data workload 112 from the memory (e.g., several hundred clock cycles later in certain implementations). However, whereas scenario 500-A proceeded by storing data workload 112 in temporary buffer 308 in response to receiving data workload 112, scenario 500-B is shown to proceed by: 3) in response to receiving data workload 112, caching data workload 112 in data store 302 of cache 108 (the data store 302 being configured, as described above, for caching data obtained from the memory).
[0075] As shown in scenario 500-B, the data workload 112 may, in addition or as an alternative to being held at temporary buffer 308 of cache 108, be cached within cache 108 (i.e., indexed and stored in data store 302) where it could also be accessed by another LI data cache corresponding to another core of the processor. As such, if first cache 106-1 is associated with a first core of the processor (e.g., core 104-1, not shown in FIGS. 5A-5B) and a third cache 106-2 of the processor is associated with a second core of the processor (e.g., core 104-2, also not shown in FIGS. 5A-5B), both the first cache and the third cache (i.e., LI caches 106-1 and 106-2) may be configured to access the memory by way of the second cache (i.e., L2 cache 108). Accordingly, in scenario 500-B, if an additional prefetch requestAtty Docket No. 0120-1141WO1from the third cache (i.e., LI cache 106-2) to the second cache (i.e., L2 cache 108) for the data workload 112 is sent and received while the data workload 112 is stored in data store 302 of cache 108, cache 108 may send the data workload 112 to cache 106-2 in response to the additional prefetch request. In other words, in contrast to scenario 500-A, there may be an L2 cache hit (rather than miss) when cache 106-2 sends the additional prefetch request, and the cached data workload 112 may be provided accordingly.
[0076] It will be understood that the data handling may be performed in a variety of ways to ultimately accomplish the outcomes described above in relation to scenario 500-B. For example, one arrow in FIG. 5B shows a possibility that data workload 112 may be transferred directly from memory to data store 302. In this example, there may be no temporary buffer 308 or the temporary buffer 308 may not be used to store data workload 112. As another possibility illustrated by another arrow in FIG. 5B, there is a possibility that data workload 112 may be transferred from memory to temporary buffer 308 and then copied or moved to data store 302 thereafter. In this case, cache 106-1 could still receive the copy of data workload 112 from temporary buffer 308 while cache 106-2 could receive the cached data workload 112 from data store 302. As still another possibility, caches 106-1 and 106-2 could both receive the cached data workload 112 from data store 302 (again, making the use of temporary buffer 308 optional for this scenario).
[0077] Cache 108 may determine how to handle data obtained from memory based on static characteristics (e.g., which may vary from implementation to implementation) or based on more dynamic considerations. For instance, in one dynamic scenario, the prefetch request for data workload 112 to be cached within the first cache (i.e., cache 106-1) may further be a prefetch request for data workload 112 to be cached within the second cache (i.e., cache 108). In other words, rather than the prefetch request merely indicating that the data workload 112 is to be cached at the LI level (e.g., within cache 106-1), the prefetch request may further indicate that the data workload 112 is to be cached at both the LI level (e.g., within cache 106-1) and at the L2 level (e.g., within cache 108). The caching of the data workload 112 in data store 302 of cache 108 may then be performed (i.e., dynamically) based on the prefetch request for the data workload to be cached within the second cache. In this way, when another cache at the LI level (e.g., cache 106-2) needs access to the data workload 112, the data will be cached and available for request from data store 302.
[0078] Returning to FIG. 4, operation 422 is shown to be performed by the L2 cache by indicating, to the LI cache and in response to obtaining the data workload as described in connection with operation 420, that the data workload is available for request by the LIAtty Docket No. 0120-1141WO1cache. Notably, while the indication that the data workload was not cached (i.e., the L2 cache miss) in operation 418 was provided in response to a prefetch request sent by the LI cache, this indication in operation 422 that the data workload is available for request by the LI cache may be provided in response to the data workload being received from memory (e.g., after likely waiting hundreds of clock cycles for the data, as has been described) rather than in response to anything the LI cache has done. In other words, after receiving an indication at operation 418 that there was an L2 cache miss and that the L2 cache is working on obtaining the data, the LI cache may no longer track, prepare for, or otherwise take action to accommodate the reception of the data workload. Indeed, as described above, the LI cache may, at operation 406, deallocate the entry that was used to track the prefetch request in response to the indication from the L2 cache (as part of operation 418), and may move on to other tasks.
[0079] Rather than the indication of operation 422, an alternative option is that the L2 cache could push the data workload to the LI cache at this stage. However, assuming that the LI cache is no longer necessarily prepared to receive the data workload for the reasons described above (i.e., because the entry was deallocated and the LI cache moved onto other tasks without continuing to track and prepare to accommodate the data workload) this alternative option may present certain risks. For example, if the LI cache is not able to instantly receive the data workload whenever the L2 cache happens to be ready to send it (e.g., several hundred clock cycles after the LI received the L2 cache miss indication and moved on to other tasks) the LI cache may need to stall the data channel on which the data workload is being sent, thereby creating a bottleneck that could impact performance.Consequently, while pushing the data without warning may be performed in certain implementations and / or circumstances, the risk may be mitigated and advantageous technical benefits may be obtained by sending the indication of operation 422 and thereafter allowing the LI cache to request the data workload on its own terms and on its own timeline (i.e., when the LI cache is prepared to receive the data workload, as described above in relation to operation 408).
[0080] To this end, the indication at operation 422 may be configured to indicate certain aspects of the identity of the data workload (e.g., the memory address of the workload, a data type or other metadata characterizing the workload, etc.) and to indicate the readiness of the data in any suitable way. For example, one convenient and advantageous way to communicate both the workload identity and the fact that it is now received and available for the LI cache to request is by sending a stashing snoop signal to the LI cache. Generally, aAtty Docket No. 0120-1141WO1stashing snoop signal is sent by one cache to another for purposes of cache coherence. For example, if one cache (e.g., an L2 cache) plans to stash data (i.e., push a copy of data) to another cache (e.g., an LI cache), a snoop signal prior to the stash (referred to herein as a stashing snoop signal) may indicate to the latter cache that the stash is coming and that it should prepare accordingly. In some cases, the stashing snoop signal further serves as a request for the cache to indicate if it has changed the data that is to be stashed, so that any changes can be captured and saved prior to the new version of the data being stashed.
[0081] In this context, the stashing snoop signal may be a convenient way to accomplish the indication of operation 422, but it will be understood that the stashing snoop signal is, in this example, being appropriated for a different purpose than the standard one that it would generally serve. Specifically, as described above, the stashing snoop signal in this case may be repurposed to communicate to the LI cache that a previously requested data workload has now been obtained and is being held in reserve for the LI cache to request and receive when it is ready to do so. While a stashing snoop signal is not the only way for the L2 cache to achieve this communication, there may be various advantages to co-opting or repurposing an established communication protocol such as the stashing snoop to communicate this information from one cache to the other.
[0082] At operation 424, the L2 cache may receive an additional prefetch request from the LI cache. For example, as suggested by an arrow in FIG. 4, the additional prefetch request received at operation 424 may be the additional prefetch request sent by the LI cache at operation 408, described above). As mentioned, this additional prefetch request, like the original prefetch request sent at operation 402 and received at operation 416, may be for the data workload to be cached within the LI cache. Unlike with the previous prefetch request, however, this additional prefetch request is sent in response to receiving the indication that the data workload is available for request (e.g., being cached in data store 302 and / or being held in temporary buffer 308). As such, by the point the additional prefetch request is sent and received, the L2 cache will have the data workload ready to send and the LI cache will be prepared to receive it.
[0083] At operation 426, the L2 cache may therefore send the data workload from the L2 cache to the LI cache in response to the additional prefetch request. In this way, the LI cache will have the data workload cached and ready to go when, later at operation 412 (described above), the core sends the load request to the LI cache for the data within the data workload.
[0084] FIGS. 6A and 6B show illustrative dataflow diagrams that represent dataflowAtty Docket No. 0120-1141WO1between a particular core 104, an LI cache 106, an L2 cache 108, and memory 110. More particularly, each of FIGS. 6 A and 6B show two dataflow diagrams to contrast different types of prefetch techniques. First, FIG. 6A shows a dataflow diagram 600-H that illustrates a hierarchical prefetch technique and a dataflow diagram 600 that illustrates a hybrid prefetch technique that contrasts with the hierarchical prefetch technique. FIG. 6B then shows a dataflow diagram 600-F that illustrates a flat prefetch technique next to the same dataflow diagram 600 illustrating the hybrid prefetch technique to show how it also contrasts with the flat prefetch technique.
[0085] In both dataflow diagrams of both FIGS. 6A and 6B, the four components are represented (i.e., the core 104, the LI cache 106, the L2 cache 108, and the memory 110) as they perform actions and intercommunicate with one another in accordance with the description provided below. These components have been described above and will be understood to operate accordingly. For example, communications may be sent and received by way of input and output interfaces, data processing may be performed using controllers of the components, data may be stored in buffers or data stores associated with the components, and so forth in accordance with principles described above.
[0086] In dataflow diagram 600-H, the hierarchical prefetch technique is shown to be performed in three phases separated by dashed lines: Phase 1, Phase 2, and Phase 3. In Phase 1, LI cache 106 sends a communication 602-1 to L2 cache 108 representing a prefetch request for a particular data workload to be cached within L2 cache 108. Ultimately, the objective may be for the data workload to become cached within LI cache 106. However, because of the hierarchical nature of the caching strategy in this example, the data workload must first be cached at the L2 level prior to being cached at the LI level in this example. In response to communication 602-1, L2 cache 108 may determine that the data workload is not yet cached in the L2 cache 108 (an L2 cache miss) and may send thus two communications based on the prefetch request and the determination of the L2 cache miss. These two communications may be sent in any order or simultaneously. First, a communication 602-2 from L2 cache 108 to memory 110 may request the data workload from memory 110.Second, a communication 602-3 from L2 cache 108 back to LI cache 106 may acknowledge the prefetch request and indicate that the data workload is not cached in cache 108 (i.e., indicate that there was an L2 cache miss).
[0087] A time period 604-1 is shown in dataflow diagram 600-H to indicate an amount of time that LI cache 106 waits to hear back from L2 cache 108 after sending the prefetch request of communication 602-1. This time period 604-1 may represent the amountAtty Docket No. 0120-1141WO1of time that the outstanding prefetch request is tracked in a request queue (e.g., a request queue 202 such as described above) of LI cache 106. Because L2 cache 108 responds immediately to acknowledge the request and indicate there has been an L2 cache miss (i.e., communication 602-3), time period 604-1 may be a reasonably short amount of time. For example, as has been described, time period 604-1 may represent a few tens of clock cycles (e.g., approximately 20 clock cycles in one example), which the request queue 202 may be well-equipped to handle.
[0088] Once the data workload can be accessed by the memory controller (e.g., obtained from DRAM by memory 110), memory 110 may send a communication 602-4 to L2 cache 108 that includes the requested data workload. A time period 606 is shown to indicate an amount of time that memory 110 takes to access the data workload. As has been described, this time period 606 may be a relatively long amount of time. For example, time period 606 may represent a few hundreds of clock cycles (e.g., approximately 300 clock cycles in one example). An omission symbol (lightning bolt) is illustrated in connection with time period 606 to indicate that time period 604-1 and time period 606 are not drawn to scale, but that time period 606 may be much longer (e.g., at least an order of magnitude longer). It will be understood that, at the end of Phase 1 of dataflow diagram 600-H, L2 cache 108 has received the data workload from memory 110 (communication 602-4) and has cached the data workload so that LI caches below it in the hierarchy (including the LI cache 106 that originally requested the data) will be able to access it.
[0089] In Phase 2 of dataflow diagram 600-H, the LI cache 106 sends a communication 602-5 that now represents a prefetch request for the same data workload to be cached within the LI cache 106. LI cache 106 tracks this request for a time period 604-2 that is similar to time period 604-1 (e.g., a few tens of clock cycles). However, whereas there was previously an L2 cache miss (so that communication 602-3 acknowledged the request and indicated the miss), L2 cache 108 now has the requested data workload and there is an L2 cache hit. Accordingly, at communication 602-6, the requested data workload is sent to LI cache 106 and the data workload is successfully prefetched to the LI cache.
[0090] In Phase 3 of dataflow diagram 600-H, core 104 is shown to send a communication 602-7 that represents a load request for data from the data workload that was previously prefetched to LI cache 106. Because the data workload was successfully prefetched and is cached in LI cache 106, there is an LI cache hit and the requested data is loaded to core 104 by a communication 602-8. As has been mentioned (though not drawn to scale here), this load may happen on a relatively very short time scale (e.g., within one clockAtty Docket No. 0120-1141WO1cycle or just a few clock cycles).
[0091] The hierarchical prefetch technique illustrated in dataflow diagram 600-H ultimately accomplishes the objective of prefetching the data workload to both the L2 cache 108 and to the LI cache 106, thereby benefiting processor performance when core 104 loads the data from the data workload and finds it readily available from its LI cache. However, the hierarchical prefetch technique has several shortcomings and introduces several technical problems. First, the operations of Phase 1 and the operations of Phase 2 are each performed independently without any hints or indications as to when one is complete and the other can begin. As a result, there is a need for a prefetch coordinator (e.g., associated with a cache controller of LI cache 106, associated with a cache controller of L2 cache 108, etc.) that can perform highly-sophisticated coordination between the prefetch to the L2 cache (in Phase 1) and the prefetch to the LI cache (in Phase 2). For example, the prefetch coordination would need to be configured to perfectly time the transactions to follow one another in the order shown, based on estimated memory latency and L2 hit latency. While this may be possible to implement, this type of coordination makes for a complex design that could be expensive in terms of power, hardware footprint, manufacturing, design time, and so forth.
[0092] Another technical problem of the hierarchical prefetch technique of dataflow diagram 600-H is that there is no guarantee that even when the data workload has been cached to L2 cache 108 (in Phase 1) that the later prefetch to the LI cache (in Phase 2) will occur as intended. For example, since L2 cache 108 may be shared among many cores and LI caches (as described above, though not shown in FIG. 6A), other cores may thrash the cache line (e.g., access and change the data of the data workload cached in the L2 cache) before the prefetch to LI cache 106 has been performed. This thrashing of the data workload would result in a loss of performance. Additionally, another potential shortcoming of the hierarchical prefetch technique of dataflow diagram 600-H is that any data workload that is desired for prefetch into the LI cache also has to be cached to the L2 cache first (even though the data workload may be able to be fully contained within LI cache 106). This requirement may lead to significant L2 cache pollution for the shared L2 cache, which also will tend to reduce processor performance.
[0093] Turning to FIG. 6B, dataflow diagram 600-F shows the flat prefetch technique that is distinct from the hierarchical prefetch techniques but has its own set of shortcomings (as mentioned above and detailed further below). The flat prefetch technique of dataflow diagram 600-F is shown to be performed in two phases separated by dashed lines: Phase 1 and Phase 2. Similarly to dataflow diagram 600-H described above, dataflow diagram 600-FAtty Docket No. 0120-1141WO1is shown to begin, in Phase 1, with LI cache 106 sending a communication 602-1 to L2 cache 108. However, while the hierarchical structure of the earlier example required the prefetch request to first request the data workload to be cached within L2 cache 108, the flat structure of this example allows communication 602-1 to represent a direct prefetch request for the data workload to be cached within LI cache 106. Again in this example, L2 cache 108 may determine that the data workload is not cached in the L2 cache 108 (an L2 cache miss).However, instead of sending communication 602-3 to LI cache 106 to acknowledge the request and indicate the L2 cache miss, dataflow diagram 600-F shows that L2 cache 108 may proceed (while LI cache 106 continues to wait) to send communication 602-2 to memory 110 to request the data workload from memory 110.
[0094] One consequence of using the flat prefetch technique (in which data can be prefetched directly to a lower level cache without also necessarily being cached at the higher levels) rather than the hierarchical prefetch technique (in which all data moves hierarchically from memory, to higher level caches, to lower level caches) is illustrated by a time period 604-3 that indicates the amount of time that LI cache 106 has to track the outstanding prefetch request in its request queue (e.g., a request queue 202 such as described above). Whereas the hierarchical use of the L2 cache to first cache the data workload itself allowed for time periods 604-1 and 604-2 to be limited reasonably to a few tens of clock cycles, the flat prefetch technique shows that the LI cache 106 now keeps track of the outstanding request all the way from when the prefetch request is sent at communication 602-1 until the data is received from memory 110, via L2 cache 108 (communication 602-4), to be cached at LI cache 106 (communication 602-6). If time period 606 is on the order of hundreds of clock cycles (e.g., approximately 300 clock cycles in one example), this means that time period 604-3 is even longer (e.g., also on the order of hundreds of clock cycles, such as approximately 350 clock cycles in one example).
[0095] Similar to Phase 3 of dataflow diagram 600-H, in Phase 2 of dataflow diagram 600-F, core 104 is shown to send a communication 602-7 that represents a load request for data from the data workload that was previously prefetched to LI cache 106. Because the data workload was successfully prefetched and is cached in LI cache 106, there is again an LI cache hit and the requested data is loaded to core 104 by a communication 602-8. As has been mentioned (though not drawn to scale here), this load may happen on a relatively very short time scale (e.g., within one clock cycle or just a few clock cycles).
[0096] As with the hierarchical prefetch technique, the flat prefetch technique illustrated in dataflow diagram 600-F ultimately accomplishes the objective of prefetchingAtty Docket No. 0120-1141WO1the data workload to LI cache 106 so that processor performance is benefitted when core 104 loads the data from the data workload and finds it readily available from the LI cache. Like the hierarchical prefetch technique, however, the flat prefetch technique also has its own shortcomings and introduces its own technical problems (e.g., different problems than those introduced by the hierarchical prefetch technique). Foremost among these problems is the length of time period 604-3. It would be extremely inefficient and likely prohibitive for an LI cache to implement a request queue large enough to hold entries for a time period such as required by the flat prefetch technique illustrated in dataflow diagram 600-F (i.e., a time period such as time period 604-3). An LI request queue configured to track outstanding requests for hundreds of clock cycles such as required by this technique would create challenges for power, performance, area (hardware footprint), and other aspects of the design, and would therefore likely be prohibitive.
[0097] To address and remedy all the shortcomings and technical problems of the hierarchical and flat prefetch techniques that have been described, a hybrid prefetch technique is illustrated by dataflow diagram 600 in both FIG. 6A (to contrast with dataflow diagram 600-H) and FIG. 6B (to contrast with dataflow diagram 600-F). As suggested by the term “hybrid,” this prefetch technique has certain similarities with both the hierarchical and the flat prefetch techniques so that it provides all the same benefits that they each provide. However, novel features of the hybrid prefetch technique allow it to offer these benefits while avoiding the shortcomings that have been described. Accordingly, dataflow diagram 600 illustrates technical solutions provided by the hybrid prefetch technique to the various technical problems laid out for the hierarchical and flat prefetch techniques described above.
[0098] The hybrid prefetch technique of dataflow diagram 600 is shown to be performed in two phases separated by dashed lines: Phase 1 and Phase 2. Similarly to dataflow diagrams 600-H and 600-F described above, dataflow diagram 600 is shown to begin, in Phase 1, with LI cache 106 sending a communication 602-1 to L2 cache 108. As with the flat prefetch technique (and in contrast with the hierarchical prefetch technique) this communication 602-1 in dataflow diagram 600 will be understood to represent a prefetch request for the data workload to be cached directly to the LI cache 106 (without necessarily also caching it in the L2 cache, though, as has been described, caching to both levels may be performed in certain instances). As in the hierarchical example above, L2 cache 108 may determine that the data workload is not cached in the L2 cache 108 (an L2 cache miss) and, in response, may send the same two communications described above: 1) the communication 602-2 to memory 110 to request the data workload from memory 110, and 2) theAtty Docket No. 0120-1141WO1communication 602-3 to LI cache 106 to acknowledge the request and indicate the L2 cache miss. The sending of communication 602-3 allows LI cache 106 to deallocate the entry from the request queue, such that, as shown, time period 604-1 is reasonably short (similar to the hierarchical prefetch example and avoiding the extremely long time period illustrated by time period 604-3 in the flat prefetch example).
[0099] Meanwhile, like the hierarchical example, memory 110 may take a time period 606 to access the requested data workload, followed by sending the data in a communication 602-4. While the hierarchical prefetch technique concluded Phase 1 with this communication 602-4 and relied on complex coordination from a prefetch controller to trigger another prefetch request in Phase 2, however, dataflow diagram 600 shows that a new communication 610 may be used to avoid the need for this complex coordination. Specifically, as described above, this communication 610 may indicate, to LI cache 106, that the data workload is available for request by the LI cache (e.g., is being held in reserve in a temporary buffer such as illustrated in FIG. 5A and / or is being cached at the L2 level as illustrated in FIG. 5B). For example, communication 610 could be implemented by a stashing snoop signal or another suitable communication between the caches.
[0100] As shown, the result of communication 610 in dataflow diagram 600 is that the receiving of the data by L2 cache 108 (by communication 602-4) is now coupled to the additional prefetch request for the workload to be cached at LI cache 106 provided within communication 602-5. In other words, while a new phase (Phase 2) was entered for the hierarchical prefetch technique and complex coordination was relied on (with a possible risk of thrashing by another LI cache), dataflow diagram 600 shows that the additional prefetch request of communication 602-5 is still part of Phase 1 and is triggered by communication 610. When the indication of communication 610 is received and LI cache 106 is ready, the additional prefetch request of communication 602-5 is sent and, as described in the other example, the L2 cache hit is determined so that the data workload is sent to and cached by LI cache 106 with communication 602-6. Here again, the time period 604-2 associated with tracking the additional prefetch request is a relatively short amount of time that may be easily handled by the request queue of LI cache 106.
[0101] Similar to Phase 3 of dataflow diagram 600-H and Phase 2 of dataflow diagram 600-F, Phase 2 of dataflow diagram 600 shows that core 104 may send a communication 602-7 that represents a load request for data from the data workload that was previously prefetched to LI cache 106. Because the data workload was successfully prefetched and is cached in LI cache 106, there is an LI cache hit and the requested data isAtty Docket No. 0120-1141WO1loaded to core 104 by a communication 602-8. As has been mentioned (though not drawn to scale here), this load may happen on a relatively very short time scale (e.g., within one clock cycle or just a few clock cycles).
[0102] As has been described, the limitations of the hierarchical prefetch technique (complex coordination, risk of thrashing, unwanted L2 cache pollution, etc.), as well as the limitations of the flat prefetch technique (unreasonably large request queue need to track an outstanding LI prefetch request for hundreds of clock cycles), are all addressed and solved by the hybrid prefetch technique illustrated in dataflow diagram 600. At the same time, all the benefits of prefetching and caching (e.g., improved processor performance, etc.) are still enjoyed by the processor using the hybrid technique.
[0103] As has been mentioned, various methods and processes described herein may be implemented at least in part as instructions embodied in a non-transitory computer-readable medium and executable by one or more computing devices. In general, a processor (e.g., a microprocessor) receives instructions, from a non-transitory computer-readable medium (e.g., a memory, etc.), and executes those instructions, thereby performing one or more operations such as the operations described herein. Such instructions may be stored and / or transmitted using any of a variety of known computer-readable media.
[0104] A computer-readable medium (also referred to as a processor-readable medium) includes any non-transitory medium that participates in providing data (e.g., instructions) that may be read by a computer (e.g., by a processor of a computer). Such a medium may take many forms, including, but not limited to, non-volatile media, and / or volatile media. Non-volatile media may include, for example, optical or magnetic disks and other persistent memory. Volatile media may include, for example, dynamic random-access memory (DRAM), which typically constitutes a main memory.
[0105] FIG. 7 shows an illustrative computing device 700 that may use or implement apparatuses (e.g., LI caches, L2 caches, cores, memories, processors, CPUs, etc.) described herein. As shown in FIG. 7, computing device 700 may include a communication interface 702, a processor 704 (which may represent an implementation of processor 102 with one or more cores 104, one or more caches 106 and 108, etc.), a storage device 706 (which may represent a memory such as memory 110), and an input / output (I / O) module 708 communicatively connected via a communication infrastructure 710. While an illustrative computing device 700 is shown in FIG. 7, the components illustrated in FIG. 7 are not intended to be limiting. Additional or alternative components may be used in other embodiments. Components of computing device 700 shown in FIG. 7 will now be describedAtty Docket No. 0120-1141WO1in additional detail.
[0106] Communication interface 702 may be configured to communicate with one or more computing devices. Examples of communication interface 702 include, without limitation, a wired network interface (such as a network interface card), a wireless network interface (such as a wireless network interface card), a modem, an audio / video connection, and any other suitable interface.
[0107] Processor 704 generally represents any type or form of processing unit capable of processing data or interpreting, executing, and / or directing execution of one or more of the instructions, processes, and / or operations described herein. Processor 704 may direct execution of operations in accordance with one or more applications 712 or other computerexecutable instructions such as may be stored in storage device 706 or another computer-readable medium.
[0108] Storage device 706 may include one or more data storage media, devices, or configurations and may employ any type, form, and combination of data storage media and / or device. For example, storage device 706 may include, but is not limited to, a hard drive, network drive, flash drive, magnetic disc, optical disc, RAM, dynamic RAM, other non-volatile and / or volatile data storage units, or a combination or sub-combination thereof. Electronic data, including data described herein, may be temporarily and / or permanently stored in storage device 706. For example, data representative of one or more executable applications 712 configured to direct processor 704 to perform any of the operations described herein may be stored within storage device 706. In some examples, data may be arranged in one or more databases residing within storage device 706.
[0109] I / O module 708 may include one or more VO modules configured to receive user input and provide user output. One or more VO modules may be used to receive input for a single virtual experience. I / O module 708 may include any hardware, firmware, software, or combination thereof supportive of input and output capabilities. For example, I / O module 708 may include hardware and / or software for capturing user input, including, but not limited to, a keyboard or keypad, a touchscreen component (e.g., touchscreen display), a receiver (e.g., an RF or infrared receiver), motion sensors, and / or one or more input buttons.
[0110] I / O module 708 may include one or more devices for presenting output to a user, including, but not limited to, a graphics engine, a display (e.g., a display screen), one or more output drivers (e.g., display drivers), one or more audio speakers, and one or more audio drivers. In certain embodiments, I / O module 708 is configured to provide graphical data to a display for presentation to a user. The graphical data may be representative of one orAtty Docket No. 0120-1141WO1more graphical user interfaces and / or any other graphical content as may serve a particular implementation.
[0111] The following clauses describe implementations of prefetch coordination between a first cache and a second cache of a processor in accordance with principles described herein.
[0112] Clause 1. A method comprising: receiving, from a first cache within a processor and by a second cache within the processor, a prefetch request for a data workload to be cached within the first cache; in response to the prefetch request and a determination that the data workload is not cached in the second cache: indicating, by the second cache to the first cache, that the data workload is not cached in the second cache, and obtaining, by the second cache, the data workload from a memory coupled to the processor; and in response to obtaining the data workload, indicating, to the first cache by the second cache, that the data workload is available for request by the first cache.
[0113] Clause 2. The method of clause 1, further comprising: allocating, within a request queue associated with the first cache and in response to the first cache sending the prefetch request to the second cache, an entry for the prefetch request; and deallocating, within the request queue and in response to the indicating to the first cache that the data workload is not cached in the second cache, the entry for the prefetch request.
[0114] Clause 3. The method of any of clauses 1 to 2, further comprising: sending, to the second cache from the first cache and in response to the indicating that the data workload is available, an additional prefetch request for the data workload to be cached within the first cache; and based on the additional prefetch request, receiving the data workload by the first cache from the second cache.
[0115] Clause 4. The method of clause 3, further comprising: subsequent to the receiving of the data workload, receiving, by the first cache from a core of the processor, a load request for data included within the data workload; and in response to the load request, sending, by the first cache to the core, the data included within the data workload.
[0116] Clause 5. The method of clause 3, further comprising: allocating, within a request queue associated with the first cache and in response to the sending of the additional prefetch request, an entry for the additional prefetch request; and deallocating, within the request queue and in response to the receiving of the data workload, the entry for the additional prefetch request.
[0117] Clause 6. The method of any of clauses 1 to 5, wherein the obtaining of the data workload from the memory includes: requesting and receiving, by the second cache, theAtty Docket No. 0120-1141WO1data workload from the memory; and in response to receiving the data workload, storing the data workload in a temporary buffer of the second cache, the temporary buffer being configured to hold the data workload in reserve for the first cache and being separate from a data store of the second cache configured for caching data obtained from the memory.
[0118] Clause 7. The method of clause 6, wherein: the first cache is associated with a first core of the processor; a third cache of the processor is associated with a second core of the processor; the first cache and the third cache are both configured to access the memory by way of the second cache; and the method further comprises, in response to an additional prefetch request from the third cache to the second cache for the data workload while the data workload is stored in the temporary buffer of the second cache, indicating, by the second cache to the third cache, that the data workload is not cached in the second cache.
[0119] Clause 8. The method of any of clauses 1 to 5, wherein the obtaining of the data workload from the memory includes: requesting and receiving, by the second cache, the data workload from the memory; and in response to receiving the data workload, caching the data workload in a data store of the second cache configured for caching data obtained from the memory.
[0120] Clause 9. The method of clause 8, wherein: the first cache is associated with a first core of the processor; a third cache of the processor is associated with a second core of the processor; the first cache and the third cache are both configured to access the memory by way of the second cache; and the method further comprises, in response to an additional prefetch request from the third cache to the second cache for the data workload while the data workload is cached in the data store of the second cache, sending, by the second cache to the third cache, the data workload.
[0121] Clause 10. The method of clause 8, wherein: the prefetch request is for the data workload to be cached within the first cache and is further for the data workload to be cached within the second cache; and the caching the data workload in the data store of the second cache is performed based on the prefetch request for the data workload to be cached within the second cache.
[0122] Clause 11. The method of any of clauses 1 to 10, wherein the indicating that the data workload is available for request by the first cache includes sending, to the first cache by the second cache, a stashing snoop signal.
[0123] Clause 12. The method of any of clauses 1 to 11, wherein: the first cache is associated with a core of the processor; the first cache and the core are associated with a first clock domain synchronized by a first clock signal having a first clock frequency; and theAtty Docket No. 0120-1141WO1second cache and the memory are associated with a second clock domain synchronized by a second clock signal having a second clock frequency, the second clock frequency being lower than the first clock frequency.
[0124] Clause 13. A processor comprising: a first cache configured to send a prefetch request for a data workload to be cached within the first cache; and a second cache coupled to the first cache and coupled to a memory, the second cache being configured to: receive the prefetch request from the first cache and, in response to the prefetch request and a determination that the data workload is not cached in the second cache: indicate, to the first cache, that the data workload is not cached in the second cache, and obtain the data workload from the memory; and in response to obtaining the data workload, indicate, to the first cache, that the data workload is available for request by the first cache.
[0125] Clause 14. The processor of clause 13, further comprising: a third cache; a first core associated with the first cache; and a second core associated with the third cache; wherein the first cache and the third cache are both configured to access the memory by way of the second cache.
[0126] Clause 15. The processor of clause 13, further comprising a core and wherein the first cache is further configured to: send, to the second cache in response to the indicating that the data workload is available, an additional prefetch request for the data workload to be cached within the first cache; receive, based on the additional prefetch request, the data workload from the second cache; receive, from the core and subsequent to the receiving of the data workload, a load request for data included within the data workload; and send, to the core and in response to the load request, the data included within the data workload.
[0127] Clause 16. The processor of clause 15, further comprising a request queue associated with the first cache and wherein the first cache is further configured to: allocating, within the request queue and in response to the first cache sending the prefetch request, a first entry for the prefetch request; deallocating, within the request queue and in response to the indicating that the data workload is not cached in the second cache, the first entry for the prefetch request; allocating, within the request queue and in response to the sending of the additional prefetch request, a second entry for the additional prefetch request; and deallocating, within the request queue and in response to the receiving of the data workload from the second cache, the second entry for the additional prefetch request.
[0128] Clause 17. The processor of any of clauses 13 to 16, wherein: the first cache includes a first static random-access memory (SRAM) data cache; the second cache includesAtty Docket No. 0120-1141WO1a second SRAM data cache; and the memory includes a dynamic random-access memory (DRAM) module.
[0129] Clause 18. A cache controller within a cache of a processor, the cache controller configured to: receive, from an additional cache of the processor, a prefetch request for a data workload to be cached within the additional cache; in response to the prefetch request and a determination that the data workload is not cached in the cache: indicate, to the additional cache, that the data workload is not cached in the cache, and obtain the data workload from the memory; and in response to obtaining the data workload, indicate, to the additional cache, that the data workload is available for request by the additional cache.
[0130] Clause 19. The cache controller of clause 18, wherein the cache controller obtains the data workload from the memory by: requesting and receiving the data workload from the memory; and in response to receiving the data workload, storing the data workload in a temporary buffer of the cache that is configured to hold the data workload in reserve for the additional cache and is separate from a data store configured for caching data obtained from memory.
[0131] Clause 20. The cache controller of clause 18, wherein the cache controller obtains the data workload from the memory by: requesting and receiving the data workload from the memory; and in response to receiving the data workload, caching the data workload in a data store of the cache that is configured for caching data obtained from memory.
[0132] Clause 21. The cache controller of clause 20, wherein: the prefetch request is for the data workload to be cached within the additional cache and is further for the data workload to be cached within the cache; and the caching the data workload in the data store is performed based on the prefetch request for the data workload to be cached within the cache.
[0133] Clause 22. A cache comprising the cache controller of clause 18.
[0134] Clause 23. A processor including the cache of clause 22.
[0135] Various implementations of the systems and techniques described herein can be realized in digital electronic circuitry, integrated circuitry, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.Atty Docket No. 0120-1141WO1
[0136] A number of implementations have been described. Nevertheless, it will be understood that various modifications may be made without departing from the spirit and scope of the description and claims. In addition, the logic flows depicted in the figures do not require the particular order shown, or sequential order, to achieve desirable results. In addition, other steps may be provided, or steps may be eliminated, from the described flows, and other components may be added to, or removed from, the described systems.Accordingly, other implementations are within the scope of the following claims.
[0137] Specific structural and functional details disclosed herein are merely representative for purposes of describing example implementations. Example implementations, however, may be embodied in many alternate forms and should not be construed as limited to only the implementations set forth herein.
[0138] It will be understood that when the terms first, second, etc. are used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. A first element could be termed a second element, and, similarly, a second element could be termed a first element, without departing from the scope of the implementations of the disclosure. As used herein, the term and / or includes any and all combinations of one or more of the associated listed items.
[0139] The terminology used herein is for the purpose of describing particular implementations only and is not intended to be limiting of the implementations. As used herein, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises,” “comprising,” “includes,” and / or “including,” when used in this specification, specify the presence of the stated features, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or groups thereof.
[0140] It will be understood that when an element is referred to as being “coupled,” “connected,” or “responsive” to, or “on,” another element, it can be directly coupled, connected, or responsive to, or on, the other element, or intervening elements may also be present. In contrast, when an element is referred to as being “directly coupled,” “directly connected,” or “directly responsive” to, or “directly on,” another element, there are no intervening elements present. As used herein the term “and / or” includes any and all combinations of one or more of the associated listed items.
[0141] Spatially relative terms, such as “beneath,” “below,” “lower,” “above,” “upper,” and the like, may be used herein for ease of description to describe one element orAtty Docket No. 0120-1141WO1feature in relationship to another element(s) or feature(s) as illustrated in the figures. It will be understood that the spatially relative terms are intended to encompass different orientations of the device in use or operation in addition to the orientation depicted in the figures. For example, if the device in the figures is turned over, elements described as “below” or “beneath” other elements or features would then be oriented “above” the other elements or features. Thus, the term “below” can encompass both an orientation of above and below. The device may be otherwise oriented (rotated 130 degrees or at other orientations) and the spatially relative descriptors used herein may be interpreted accordingly.
[0142] Unless otherwise defined, the terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which these concepts belong. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and / or the present specification and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein.
[0143] Further to the descriptions above, a user may be provided with controls allowing the user to make an election as to both if and when systems, programs, or features described herein may enable collection of user information (e.g., information about a user's social network, social actions, or activities, profession, a user's preferences, or a user's current location), and if the user is sent content or communications from a server. In addition, certain data may be treated in one or more ways before it is stored or used, so that personally identifiable information is removed. For example, a user's identity may be treated so that no personally identifiable information can be determined for the user, or a user's geographic location may be generalized, or location information may be obtained (such as to a city, zip code, or state level), so that a particular location of a user cannot be determined. Thus, the user may have control over what information is collected about the user, how that information is used, and what information is provided to the user.
[0144] While certain features of the described implementations have been illustrated as described herein, many modifications, substitutions, changes, and equivalents may occur to those skilled in the art. It is therefore to be understood that the appended claims are intended to cover such modifications and changes as fall within the scope of the implementations. It will be understood that they have been presented by way of example only, not limitation, and various changes in form and details may be made. Any portion of the apparatus and / or methods described herein may be combined in any combination, exceptAtty Docket No. 0120-1141WO1mutually exclusive combinations. The implementations described herein can include various combinations and / or sub-combinations of the functions, components, and / or features of the different implementations described. As such, the scope of the present disclosure is not limited to the particular combinations hereafter claimed, but instead extends to encompass any combination of features or example implementations described herein irrespective of whether or not that particular combination has been specifically enumerated in the accompanying claims at this time.
Claims
Atty Docket No. 0120-1141WO1WHAT IS CLAIMED IS:
1. A method compri sing :receiving, from a first cache within a processor and by a second cache within the processor, a prefetch request for a data workload to be cached within the first cache;in response to the prefetch request and a determination that the data workload is not cached in the second cache:indicating, by the second cache to the first cache, that the data workload is not cached in the second cache, andobtaining, by the second cache, the data workload from a memory coupled to the processor; andin response to obtaining the data workload, indicating, to the first cache by the second cache, that the data workload is available for request by the first cache.
2. The method of claim 1, further comprising:allocating, within a request queue associated with the first cache and in response to the first cache sending the prefetch request to the second cache, an entry for the prefetch request; anddeallocating, within the request queue and in response to the indicating to the first cache that the data workload is not cached in the second cache, the entry for the prefetch request.
3. The method of any of claims 1 to 2, further comprising:sending, to the second cache from the first cache and in response to the indicating that the data workload is available, an additional prefetch request for the data workload to be cached within the first cache; andbased on the additional prefetch request, receiving the data workload by the first cache from the second cache.
4. The method of claim 3, further comprising:subsequent to the receiving of the data workload, receiving, by the first cache from a core of the processor, a load request for data included within the data workload; andin response to the load request, sending, by the first cache to the core, the data included within the data workload.Atty Docket No. 0120-1141WO15. The method of claim 3, further comprising:allocating, within a request queue associated with the first cache and in response to the sending of the additional prefetch request, an entry for the additional prefetch request; and deallocating, within the request queue and in response to the receiving of the data workload, the entry for the additional prefetch request.
6. The method of any of claims 1 to 5, wherein the obtaining of the data workload from the memory includes:requesting and receiving, by the second cache, the data workload from the memory; andin response to receiving the data workload, storing the data workload in a temporary buffer of the second cache, the temporary buffer being configured to hold the data workload in reserve for the first cache and being separate from a data store of the second cache configured for caching data obtained from the memory.
7. The method of claim 6, wherein:the first cache is associated with a first core of the processor;a third cache of the processor is associated with a second core of the processor; the first cache and the third cache are both configured to access the memory by way of the second cache; andthe method further comprises, in response to an additional prefetch request from the third cache to the second cache for the data workload while the data workload is stored in the temporary buffer of the second cache, indicating, by the second cache to the third cache, that the data workload is not cached in the second cache.
8. The method of any of claims 1 to 5, wherein the obtaining of the data workload from the memory includes:requesting and receiving, by the second cache, the data workload from the memory; andin response to receiving the data workload, caching the data workload in a data store of the second cache configured for caching data obtained from the memory.
9. The method of claim 8, wherein:Atty Docket No. 0120-1141WO1the first cache is associated with a first core of the processor;a third cache of the processor is associated with a second core of the processor; the first cache and the third cache are both configured to access the memory by way of the second cache; andthe method further comprises, in response to an additional prefetch request from the third cache to the second cache for the data workload while the data workload is cached in the data store of the second cache, sending, by the second cache to the third cache, the data workload.
10. The method of claim 8, wherein:the prefetch request is for the data workload to be cached within the first cache and is further for the data workload to be cached within the second cache; andthe caching the data workload in the data store of the second cache is performed based on the prefetch request for the data workload to be cached within the second cache.
11. The method of any of claims 1 to 10, wherein the indicating that the data workload is available for request by the first cache includes sending, to the first cache by the second cache, a stashing snoop signal.
12. The method of any of claims 1 to 11, wherein:the first cache is associated with a core of the processor;the first cache and the core are associated with a first clock domain synchronized by a first clock signal having a first clock frequency; andthe second cache and the memory are associated with a second clock domain synchronized by a second clock signal having a second clock frequency, the second clock frequency being lower than the first clock frequency.
13. A processor comprising:a first cache configured to send a prefetch request for a data workload to be cached within the first cache; anda second cache coupled to the first cache and coupled to a memory, the second cache being configured to:receive the prefetch request from the first cache and, in response to the prefetch request and a determination that the data workload is not cached in the second cache:Atty Docket No. 0120-1141WO1indicate, to the first cache, that the data workload is not cached in the second cache, andobtain the data workload from the memory; and in response to obtaining the data workload, indicate, to the first cache, that the data workload is available for request by the first cache.
14. The processor of claim 13, further comprising:a third cache;a first core associated with the first cache; anda second core associated with the third cache;wherein the first cache and the third cache are both configured to access the memory by way of the second cache.
15. The processor of claim 13, further comprising a core and wherein the first cache is further configured to:send, to the second cache in response to the indicating that the data workload is available, an additional prefetch request for the data workload to be cached within the first cache;receive, based on the additional prefetch request, the data workload from the second cache;receive, from the core and subsequent to the receiving of the data workload, a load request for data included within the data workload; andsend, to the core and in response to the load request, the data included within the data workload.
16. The processor of claim 15, further comprising a request queue associated with the first cache and wherein the first cache is further configured to:allocating, within the request queue and in response to the first cache sending the prefetch request, a first entry for the prefetch request;deallocating, within the request queue and in response to the indicating that the data workload is not cached in the second cache, the first entry for the prefetch request;allocating, within the request queue and in response to the sending of the additional prefetch request, a second entry for the additional prefetch request; andAtty Docket No. 0120-1141WO1deallocating, within the request queue and in response to the receiving of the data workload from the second cache, the second entry for the additional prefetch request.
17. The processor of any of claims 13 to 16, wherein:the first cache includes a first static random-access memory (SRAM) data cache; the second cache includes a second SRAM data cache; andthe memory includes a dynamic random-access memory (DRAM) module.
18. A cache controller within a cache of a processor, the cache controller configured to:receive, from an additional cache of the processor, a prefetch request for a data workload to be cached within the additional cache;in response to the prefetch request and a determination that the data workload is not cached in the cache:indicate, to the additional cache, that the data workload is not cached in the cache, andobtain the data workload from memory; andin response to obtaining the data workload, indicate, to the additional cache, that the data workload is available for request by the additional cache.
19. The cache controller of claim 18, wherein the cache controller obtains the data workload from the memory by:requesting and receiving the data workload from the memory; andin response to receiving the data workload, storing the data workload in a temporary buffer of the cache that is configured to hold the data workload in reserve for the additional cache and is separate from a data store configured for caching data obtained from memory.
20. The cache controller of claim 18, wherein the cache controller obtains the data workload from the memory by:requesting and receiving the data workload from the memory; andin response to receiving the data workload, caching the data workload in a data store of the cache that is configured for caching data obtained from memory.
21. The cache controller of claim 20, wherein:Atty Docket No. 0120-1141WO1the prefetch request is for the data workload to be cached within the additional cache and is further for the data workload to be cached within the cache; andthe caching the data workload in the data store is performed based on the prefetch request for the data workload to be cached within the cache.