Cache Allocation Control

By dynamically restricting cache allocation based on workload monitoring and set sampling, the system addresses cache thrashing, enhancing performance and reducing memory access latency.

JP2025520668APending Publication Date: 2025-07-03ADVANCED MICRO DEVICES INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024575314
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-06-28
Filing Date
2023-05-12
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Cache thrashing occurs due to increased competition for the last-level cache, leading to reduced performance as data does not stay in the cache long enough, thereby increasing memory access latency.

Method used

Implementing a system where the cache allocation is restricted based on workload monitoring and set sampling to determine which clients are permitted to allocate cache entries, adjusting the allocation policy dynamically in response to changes in workload.

Benefits of technology

Reduces cache thrashing by optimizing cache allocation, improving performance and reducing memory access latency by ensuring that the most relevant clients have access to the cache entries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025520668000001_ABST
    Figure 2025520668000001_ABST
Patent Text Reader

Abstract

Techniques for operating a cache are disclosed. The techniques include identifying a first allocation permission policy based on a change in the workload, operating the cache according to the first allocation permission policy, identifying a second allocation permission policy based on set sampling, and operating the cache according to the second allocation permission policy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] (Cross - Reference to Related Applications) This application claims the benefit of U.S. Patent Application No. 17 / 852,296, filed on June 28, 2022, the entire disclosure of which is incorporated herein by reference.

Background Art

[0002] Caches improve performance by storing copies of data that are likely to be accessed again in the future in low - latency cache memory. Improvements to cache technology are constantly being made.

[0003] A more detailed understanding can be obtained from the following description given as an example together with the accompanying drawings.

Brief Description of the Drawings

[0004]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Modes for Carrying Out the Invention

[0005] Techniques for operating a cache are disclosed. The techniques include identifying a first allocation permission policy based on a change in a workload, operating the cache according to the first allocation permission policy, identifying a second allocation permission policy based on set sampling, and operating the cache according to the second allocation permission policy.

[0006] FIG. 1 is a block diagram of an exemplary computing device 100 that can implement one or more features of the present disclosure. In various examples, the computing device 100 can be, for example, a computer, a gaming device, a handheld device, a set-top box, a television, a cellular phone, a tablet computer, or any other computing device, but is not limited thereto. The device 100 includes, without limitation, one or more processors 102, a memory 104, one or more auxiliary devices 106, a storage device 108, and a last level cache (LLC) 110. An interconnect 112, which can be a bus, a combination of buses, and / or any other communication component, communicatively links the one or more processors 102, the memory 104, the one or more auxiliary devices 106, the storage device 108, and the last level cache 110.

[0007] In various alternative forms, one or more processors 102 include a central processing unit (CPU), a graphics processing unit (GPU), a CPU and a GPU located on the same die, or one or more processor cores, and each processor core can be a CPU, a GPU, or a neural processor. In various alternative forms, at least a portion of the memory 104 is located on the same die as one or more of the processors 102, such as on the same chip or within an interposer configuration, and / or at least a portion of the memory 104 is located separately from the one or more processors 102. The memory 104 includes volatile memory or non-volatile memory (e.g., random access memory (RAM), dynamic RAM, cache).

[0008] The storage device 108 includes a fixed or removable storage device (e.g., but not limited to, a hard disk drive, a solid state drive, an optical disk, a flash drive). The one or more auxiliary devices 106 include, but are not limited to, one or more auxiliary processors 114 and / or one or more input / output (IO) devices. The auxiliary processor 114 includes, but is not limited to, a processing unit capable of executing instructions, such as a central processing unit, a graphics processing unit, a parallel processing unit capable of executing compute shader operations in a single instruction multiple data format, a multimedia accelerator such as a video encoding or decoding accelerator, or any other processor. Any auxiliary processor 114 can be implemented as a programmable processor that executes instructions, a fixed-function processor that processes data according to fixed hardware circuits, a combination thereof, or any other type of processor.

[0009] One or more I / O devices 116 include one or more input devices such as a keyboard, keypad, touch screen, touch pad, detector, microphone, accelerometer, gyroscope, biometric scanner, or a network connection (e.g., a wireless local area network card for transmitting and / or receiving wireless IEEE 802 signals), and / or one or more output devices such as a display, speaker, printer, tactile feedback device, one or more lights, antenna, or a network connection (e.g., a wireless local area network card for transmitting and / or receiving wireless IEEE 802 signals).

[0010] The final level cache 110 functions as a shared cache for various components of the device 100, such as the processor 102 and various auxiliary devices 106. In some embodiments, there are other caches within the device 100. For example, in some instances, the processor 102 includes a cache hierarchy with different levels such as level 1 and level 2. In some examples, each cache level is specific to a particular logical partition of the processor 102, such as a processor core, or a processor chip, die or package. In some examples, the hierarchy also includes other types of caches. In various examples, one or more of the auxiliary devices 106 include one or more caches.

[0011] In some examples, the last-level cache 110 is "last-level" in the sense that it is the last cache that the device 100 attempts to service a memory access request before such cache services a request from the memory 104 itself. For example, when the processor 102 accesses data that is not stored at any of the cache levels of the processor 102, the processor exports the memory access request to be satisfied by the last-level cache 110. The last-level cache 110 determines whether the requested data is stored in the last-level cache 110. If the data is within the last-level cache 110, the last-level cache 110 services the request by providing the requested data from the last-level cache 110. If the data is not within the last-level cache 110, the device 100 services the request from the memory 104. As can be seen, in some embodiments, the last-level cache 110 functions as the last cache level before the memory 104, which helps to reduce the total amount of memory access latency for accesses to the memory 104. Although techniques for operations with the last-level cache 110 are described herein, it should be understood that the techniques may alternatively be used in other types of caches or memories.

[0012] FIG. 2 shows, by way of example, the elements of the device 100 associated with the last-level cache 110. The elements include the last-level cache 110, the client 212, the cache controller 202, and the workload monitor 204.

[0013] The last-level cache 110 is shared among many clients 212 of the device 100. As used herein, the term "client" refers to any element that requests access to the last-level cache 110, such as an element of the device 100. In various examples, the client 212 includes one or more software elements (e.g., an operating system, a driver, an application, a thread, a process, or firmware) executed on a processor such as the processor 102, one or more hardware elements such as the processor 102 or the auxiliary device 106, or a combination of software and hardware.

[0014] The last-level cache 110 has a limited size. As competition for the last-level cache 110 increases, cache thrashing may occur, which can reduce the performance of the client 212. Thus, in some situations, it can be beneficial to grant a particular type of client 212, rather than other types of clients 212, the ability to allocate entries within the cache 110. The allocation is performed in response to a miss by the client 212. Specifically, in some situations, when a memory access request to the client 212 does not hit within the cache 110, the cache 110 allocates an entry within the cache 110, fetches the data targeted by the memory access request from the memory 104, and places that data in the allocated entry. If there are no empty entries within the cache 110, the cache 110 evicts data from an entry to the memory 104 and allocates the entry to the new data. Cache thrashing occurs when there is too much competition for the cache, leading to too many misses, such that the data does not stay in the cache very long, reducing the effectiveness of the cache as a means of reducing memory access latency.

[0015] Accordingly, techniques for reducing cache thrashing are provided herein by restricting which clients 212 are permitted to allocate to cache 110 based on the operating conditions of device 100. Here, allocation to the cache means designating an entry in cache 110 to store the missing data fetched from memory 104 in response to a miss. Allocation results in eviction when there is no free (e.g., invalid) entry to store the requested data. A client 212 that is not permitted to allocate to cache 110 may still be permitted in some embodiments or operating modes to access the data in the cache in other ways, such as fetching data already stored in the cache or modifying data already in the cache. However, such a client 212 does not bring new data into cache 110 in the event of a miss.

[0016] Techniques for restricting which clients 212 are permitted to allocate include determining which clients 212 are permitted to allocate to cache 110 and, accordingly, permitting or denying allocation to those clients 212. In some embodiments, the determination of which clients 212 are permitted to allocate is made according to inputs from workload monitor 204 and cache controller 202 that performs set sampling. Workload monitor 204 is one or more of software, hardware (e.g., circuitry), or a combination thereof. In some examples, at least a portion of workload monitor 204 is a driver or part of the operating system executed on processor 102. In some examples, workload monitor 204 is alternatively a hardware circuit or includes a hardware circuit. Cache controller 202 is likewise a hardware circuit, a software entity, or a combination thereof.

[0017] The workload monitor 204 monitors that the workload is being executed within the device 100. In some examples, each different workload is defined by which application is being executed on the processor 102 and / or which client 212 is "active". The client 212 is active when the client is powered on and performing at least a threshold amount of work, and the threshold can be predefined and / or adjusted dynamically. In some examples, a first type of workload that is a gaming workload includes a game application being executed on the processor 102, and the graphics processing unit (a client 212 of the auxiliary processor 114 and the LLC 110) is active. In another example, a second type of workload that is a video playback workload includes video player software being executed on the processor 102, and the video decoder (a client 212 of the auxiliary processor 114 and the LLC 110) is active. In some exemplary workloads, multiple different clients 212 are active and thus compete for the LLC 110.

[0018] The workload monitor 204 maintains permission client data 206 that indicates, for each of a plurality of workloads, which client 212 should reject the assignment of an entry in the LLC 212 while the device 100 is executing that workload. For example, in the case of a workload where the processor 102 is running a game and a client 212 including a graphics processing unit and an audio hardware device is active, the permission client data 206 indicates that the graphics processing unit and the processor 102 are permitted to be assigned to the LLC 110, but the audio hardware device is not permitted to be assigned to the LLC 110. In another example, if audio playback software is being executed on the processor 102, the audio hardware device is active, no game is being executed, but the graphics processing unit is being intermittently used (and thus is active), the permission client data 206 indicates that the audio hardware is permitted to be assigned to the LLC 110. In this case as well, the permission client data 206 indicates, for each of the plurality of workloads, which of one or more clients 212 is permitted to be assigned to the LLC 110. The permission client data 206 is stored, in various examples, in memory specifically associated with the workload monitor 204 (e.g., within the workload monitor 204 in some embodiments where the workload monitor 204 is a hardware unit), or in a different memory such as the system memory 104 or a different memory.

[0019] The cache controller 202 performs set sampling within the cache 110. The cache controller 202 permits or does not permit the client 212 based on the monitored workload and set sampling. More specifically, the workload monitor 204 determines when a workload switch occurs. In response to the workload switch, the workload monitor 204 examines a set of data (permitted client data 206) indicating which clients 212 are permitted to allocate to the cache 110 for the current workload. Then, the workload monitor 204 permits or rejects the allocation to the client 212 accordingly. During any specific period when the workload has not changed, the cache controller 202 performs set sampling to identify the set of clients 212 for which the allocation is permitted and / or the set of clients 212 for which the allocation is not permitted.

[0020] Generally, set sampling includes reserving a small portion of the cache 110 to test different configurations and operating the cache 110 according to the configuration considered to be optimal for testing. As is generally known, a set-associative cache is divided into sets, and each set has one or more ways. Set sampling uses a small portion of these sets (test sets) to test different allocation techniques and periodically adjusts the non-test sets (i.e., the sets of the cache 110 other than the test sets) to use the allocation technique considered to be optimal. "Allocation technique" refers to which clients 212 are permitted to allocate to the LLC 110 and which clients are not permitted to allocate to the LLC 110. Set sampling provides the advantage of adjusting the operation of the cache 110 to take into account the current operating conditions, but there may be a delay in that it may take some time for the cache controller 202 to "recognize" that a particular allocation technique is more optimal than the technique currently being used in the non-test test.

[0021] Figure 3 shows a set sampling operation according to an example. LLC 110 includes a plurality of non-test sets 304 and a plurality of test sets 306. Set 302 is a set in a set associative caching scheme. Such a scheme is a way in which data placed in the cache (i.e., in response to a miss) is placed in one of the ways within a particular set based on the address of the data. In one example, in the case of a miss in a cache line having an address, the cache fetches the cache line from memory 104 and places that cache line in one of the ways within the set identified by some bits of that address.

[0022] For non-test set 304, the cache controller 202 operates those sets according to the current allocation permission policy. The allocation permission policy indicates which client 212 is permitted to allocate to the last-level cache 110. For test set 306, the cache controller 202 operates those sets according to the candidate allocation permission policy. The cache controller 202 operates different test sets 306 according to different candidate allocation permission policies. Based on the performance measured in each test set 306, the cache controller 202 selects an allocation permission policy. In one example, the test set 306 selects an allocation permission policy for the test set 306 that is considered to exhibit the best performance. In one example, the test set 306 is considered to exhibit the best performance when the test set experiences the lowest rate of misses or the highest rate of hits among all test sets 306. Here, the miss rate means the rate of misses with respect to the total number of accesses within a predetermined time, and the hit rate means the rate of hits with respect to the total number of accesses. In some examples, the cache controller 202 accumulates the hit rate or miss rate over time, and after a certain period has elapsed, selects a new allocation permission policy for operating the last-level cache 110. The cache controller 202 then operates the last-level cache 110 according to that policy. Operating according to the allocation permission policy means not allowing an allocation to the client 212 specified by the allocation permission policy.

[0023] Figure 4 shows the operation of the system 100 according to an example. The last-level cache 110 operates according to the current allocation permission policy at any given point in time. As described above, the allocation permission policy indicates which client 212 is allowed to allocate an entry to the LLC 110. Operating according to the current allocation permission policy means allowing or denying an allocation to the client 212 according to the current allocation permission policy.

[0024] The cache controller 202 performs set sampling on the test set of the LLC 110. The cache controller 202 configures different test sets to operate different allocation permission policies. The cache controller 202 measures the performance of the test set over time. At various times, the cache controller 202 selects the test set with the best performance and applies the allocation permission policy of that test set to the non-test sets.

[0025] Set sampling alone may inaccurately capture the operating mode of the device 100. For example, when the workload switches the device 100 on, the newly active client 212 and / or the newly executed software may utilize the cache 110 in a different way than before the switch. However, set sampling alone cannot capture that new usage method immediately or quickly. Therefore, the workload monitor 204 controls the allocation permission policy within the LLC 110 according to the monitored workload and the permitted client data 206. Therefore, the workload monitor 204 and the cache controller 202 operate together to select an allocation permission policy for operating the LLC 110. When the workload monitor 204 detects a change in the workload that results in a different allocation permission policy, the workload monitor 204 causes the cache controller 202 to operate the LLC 110 based on that allocation permission policy. When the cache controller 202 determines, based on set sampling, that the LLC 110 should operate according to the new allocation permission policy, the cache controller 202 operates the LLC 110 according to that new permission policy.

[0026] FIG. 5 is a flowchart of a method 500 for operating a cache according to an example. Although described with respect to the systems of FIGS. 1-4, one of ordinary skill in the art will recognize that any system configured to perform the steps of method 500 in any technically feasible order is within the scope of this disclosure.

[0027] In step 502, a workload monitor 204 configured to monitor changes in a workload identifies a change in the workload. Based on this change in the workload, the workload monitor 204 identifies a new allocation permission policy. In some examples, the workload monitor 204 accesses the permission client data 206 to identify an allocation policy associated with the new workload. In some examples, the permission client data 206 includes an entry for each of a set of different workloads. Each entry indicates which allocation permission policy to use for a particular workload. In some examples, the workload monitor 204 communicates with hardware and / or software (e.g., an operating system or driver) to determine the current workload. In step 504, in response to the workload changing, the workload monitor 204 changes the allocation permission policy based on the new allocation permission policy.

[0028] In step 506, cache controller 202 identifies a new allocation permission policy based on set sampling. In various examples, this identification is performed at various timing intervals, such as irregular or regular timing intervals. At each timing interval, cache controller 202 collects test data indicating the performance of a predetermined allocation permission policy in some test sets 306 and identifies the allocation permission policy of the test set that is considered to be executed optimally. In some examples, a test set 306 having the highest hit rate (e.g., the rate of hits relative to the total memory access requests) or the lowest miss rate (e.g., the rate of misses relative to the total memory access requests) is considered to be executed optimally. In step 508, cache controller 202 operates LLC 110 according to the selected allocation permission policy.

[0029] It should be understood that the order of the steps in FIG. 5 can be reversed or rearranged in any way. Generally, method 500 changes the current allocation permission policy in response to the occurrence of a new workload. This is because the new workload serves to identify what the appropriate allocation permission policy should be. Next, since the actual operating conditions may indicate that different allocation permission policies are required based on set sampling, cache controller 202 monitors the actual cache performance using set sampling to change the allocation permission policy as needed. By changing the allocation permission policy in response to both actual performance and workload monitoring, the cache operation becomes flexible and can respond to changes in operating conditions.

[0030] Elements in the figures are embodied, where appropriate, as software executed on a processor, a fixed function processor, a programmable processor, or combinations thereof. Processor 102, final level cache 110, interconnect 112, memory 104, storage device 108, various auxiliary devices 106, client 212, cache controller 202, and workload monitor 204 include at least some hardware circuitry and, in some embodiments, include software executed on a processor within those components or within another component.

[0031] It should be understood that many variations are possible based on the disclosure herein. Although features and elements are described above in specific combinations, each feature or element can be used alone, without other features and elements, or in various combinations with or without other features and elements.

[0032] The provided method can be implemented in a general-purpose computer, processor, or processor core. Suitable processors include, by way of example, general-purpose processors, dedicated processors, conventional processors, digital signal processors (DSPs), multiple microprocessors, one or more microprocessors associated with a DSP core, controllers, microcontrollers, application specific integrated circuits (ASICs), field programmable gate array (FPGA) circuits, any other type of integrated circuit (IC), and / or state machines. Such processors can be manufactured by configuring a manufacturing process using the results of processed hardware description language (HDL) instructions and other intermediate data (such instructions that can be stored on a computer-readable medium), including netlists. The result of such processing can be a mask work, which can then be used in a subsequent semiconductor manufacturing process to manufacture a processor implementing the features of the present disclosure.

[0033] The methods or flowcharts provided herein can be implemented in a computer program, software, or firmware incorporated into a non-transitory computer-readable storage medium for execution by a general-purpose computer or processor. Examples of non-transitory computer-readable storage media include magnetic media such as read only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, internal hard disks, and removable disks, magneto-optical media, and optical media such as CD-ROM disks and digital versatile disks (DVDs).

Claims

1. A method for operating a cache of a device, comprising: identifying a first allocation permission policy based on a change in workload; operating the cache according to the first allocation permission policy; identifying a second allocation permission policy based on set sampling; operating the cache according to the second allocation permission policy. A method.

2. The method of claim 1, wherein the change in workload includes a change from the device operating according to a first workload to operating according to a second workload. The method of claim 1.

3. The method of claim 2, wherein the first workload includes a first combination of a client being active and software being executed on the device, and the second workload includes a second combination of a client being active and software being executed on the device. The method of claim 2.

4. The method of claim 1, wherein identifying the first allocation permission policy includes identifying the first allocation permission policy based on the currently active workload by referring to a set of permission client data. The method of claim 1.

5. The method of claim 1, wherein the set sampling includes operating different test sets of the cache according to different allocation permission policies. The method of claim 1.

6. The method of claim 5, wherein identifying the second allocation permission policy includes selecting an allocation permission policy in which the performance regarded as optimal is observed. The method of claim 5.

7. The method of claim 5, wherein the different test sets include sets of the cache in a set associativity scheme. The method of claim 5.

8. The method of claim 1, wherein the first allocation permission policy indicates which clients are permitted to allocate entries in the cache and which clients are not permitted to allocate entries in the cache. The method of claim 1.

9. The method of claim 1, wherein the cache is a last-level cache. The method of claim 1.

10. A system, comprising: a cache; a cache controller, wherein the cache controller is configured to: identify a first allocation permission policy based on a change in workload; Operating the cache according to the first allocation permission policy; Identifying a second allocation permission policy based on set sampling; Operating the cache according to the second allocation permission policy; configured to perform: System. **Claim 11** The system of claim 10, wherein the change in the workload includes a change from the device operating according to a first workload to operating according to a second workload. The system of claim 10. **Claim 12** The system of claim 11, wherein the first workload includes a first combination of a client being active and software being executed on the device, and the second workload includes a second combination of a client being active and software being executed on the device. The system of claim 11. **Claim 13** The system of claim 10, wherein identifying the first allocation permission policy includes identifying the first allocation permission policy based on a currently active workload by referring to a set of permission client data. The system of claim 10. **Claim 14** The system of claim 10, wherein the set sampling includes operating different test sets of the cache according to different allocation permission policies. The system of claim 10. **Claim 15** The system of claim 14, wherein identifying the second allocation permission policy includes selecting an allocation permission policy in which optimal performance is observed. The system of claim 14. **Claim 16** The system of claim 14, wherein the different test sets include sets of the cache in a set associativity scheme. The system of claim 14. **Claim 17** The system of claim 10, wherein the first allocation permission policy indicates which clients are permitted to allocate entries in the cache and which clients are not permitted to allocate entries in the cache. The system of claim 10. **Claim 18** The system of claim 10, wherein the cache is a last level cache. The system of claim 10. **Claim 19** A system comprising: a processor; a cache configured to service requests of the processor; a cache controller, wherein the cache controller is configured to identify a first allocation permission policy based on a change in workload; ​ Operating the cache in accordance with the first allocation permission policy; Identifying a second allocation permission policy based on set sampling; Operating the cache in accordance with the second allocation permission policy; configured to perform; system. **Claim 20** The change in the workload includes a change from the device operating according to a first workload to operating according to a second workload. The system of claim 19.