Dynamic cache bypass for power savings
By powering down the final level cache in response to low effectiveness triggers, the technique addresses the issue of unnecessary power consumption, enhancing energy efficiency and operational effectiveness in computing devices.
Patent Information
- Application Number
- JP2024561877
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-04-26
- Filing Date
- 2023-04-20
- Publication Date
- 2025-05-02
AI Technical Summary
The final level cache in computing devices often consumes power unnecessarily when its effectiveness is low, leading to wasted energy and reduced efficiency.
Powering down the final level cache in response to a power-down trigger, such as when cache effectiveness is deemed low, and redirecting memory access transactions directly to main memory.
This approach reduces power consumption and improves energy efficiency by disabling the final level cache when it provides little benefit, while ensuring seamless operation by flushing data to main memory.
Smart Images

Figure 2025514071000001_ABST
Abstract
Description
[Technical field]
[0001] (CROSS REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of U.S. Patent Application No. 17 / 730,041, filed April 26, 2022, which is incorporated by reference as if fully set forth herein. [Background technology]
[0002] Caching improves performance by storing copies of data that are deemed likely to be accessed again in the future in a low latency cache memory. Improvements to cache technology are constantly being made.
[0003] A more detailed understanding may be had from the following description, given by way of example in conjunction with the accompanying drawings, in which: [Brief description of the drawings]
[0004] [Figure 1] FIG. 1 is a block diagram of an example computing device capable of implementing one or more features of the present disclosure. [Diagram 2] FIG. 2 illustrates operations for powering down a last level cache in response to a power-down trigger, according to an example. [Diagram 3] FIG. 1 illustrates operations for powering down a last level cache in response to a power-down trigger, according to an example. [Figure 4A] 1 illustrates the operation of a computing device when a last level cache is powered down and when a last level cache is powered up, according to an example. [Figure 4B] 1 illustrates the operation of a computing device when a last level cache is powered down and when a last level cache is powered up, according to an example. [Figure 5A] 1 illustrates a technique for powering down a last level cache, according to an example. [Figure 5B] 1 illustrates a technique for powering down a last level cache, according to an example. [Figure 5C] 1 illustrates a technique for powering down a last level cache, according to an example. [Figure 5D] 1 illustrates a technique for powering down a last level cache, according to an example. [Figure 6] FIG. 2 illustrates a method for operating a last level cache, according to an example. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0005] Techniques are disclosed for operating a cache that include powering down the cache in response to a power-down trigger that indicates that the cache validity is deemed low.
[0006] 1 is a block diagram of an example computing device 100 capable of implementing one or more features of the present disclosure. In various examples, the computing device 100 may be, for example, but not limited to, a computer, a gaming device, a handheld device, a set-top box, a television, a mobile phone, a tablet computer, or any other computing device. The device 100 includes, but is not limited to, one or more processors 102, a memory 104, one or more auxiliary devices 106, a storage device 108, and a last level cache (LLC) 110. An interconnect 112, which may be a bus, a combination of buses, and / or any other communication component, communicatively links the one or more processors 102, the memory 104, the one or more auxiliary devices 106, the storage device 108, and the last level cache 110.
[0007] In various alternatives, the one or more processors 102 may include a central processing unit (CPU), a graphics processing unit (GPU), a CPU and a GPU located on the same die, or one or more processor cores, each of which may be a CPU, a GPU, or a neural processor. In various alternatives, at least a portion of the memory 104 may be located on the same die as one or more of the one or more processors 102, such as on the same chip or in an interposer configuration, and / or at least a portion of the memory 104 may be located separately from the one or more processors 102. The memory 104 may include volatile or non-volatile memory (e.g., random access memory (RAM), dynamic random access memory, cache).
[0008] The storage 108 includes fixed or removable storage devices (e.g., but not limited to, hard disk drives, solid state drives, optical disks, flash drives). The one or more auxiliary devices 106 include, but are not limited to, one or more auxiliary processors 114 and / or one or more input / output (IO) devices. The auxiliary processor 114 includes, but is not limited to, a processing unit capable of executing instructions, such as a central processing unit, a graphics processing unit, a parallel processing unit capable of performing compute shader operations on a single instruction multiple data format, a multimedia accelerator such as a video encoding or decoding accelerator, or any other processor. Any auxiliary processor 114 can be implemented as a programmable processor that executes instructions, a fixed function processor that processes data according to fixed hardware circuitry, a combination thereof, or any other type of processor.
[0009] The one or more IO devices 116 include one or more input devices, such as a keyboard, keypad, touch screen, touchpad, detector, microphone, accelerometer, gyroscope, biometric scanner, or network connection (e.g., a wireless local area network card for transmitting and / or receiving wireless IEEE 802 signals), and / or one or more output devices, such as a display, speaker, printer, haptic feedback device, one or more lights, antenna, or network connection (e.g., a wireless local area network card for transmitting and / or receiving wireless IEEE 802 signals).
[0010] The last level cache 110 serves as a shared cache for various components of the device 100, such as the processor 102 and the various auxiliary devices 106. In some embodiments, there are other caches in the device 100. For example, in some examples, the processor 102 includes a cache hierarchy including different levels, such as levels 1 and 2. In some examples, each such cache level is specific to a particular logical division of the processor 102, such as a processor core or a processor chip, die, or package, etc. In some examples, the hierarchy also includes other types of caches. In various examples, one or more of the auxiliary devices 106 includes one or more caches.
[0011] The last level cache 110 is “last level” in the sense that such cache is the last cache that the device 100 attempts to service a memory access request before servicing the request from the memory 104 itself. For example, if the processor 102 accesses data that is not stored in any of the cache levels of the processor 102, the processor exports the memory access request to be satisfied by the last level cache 110. The last level cache 110 determines whether the requested data is stored in the last level cache 110. If the data is in the last level cache 110, the last level cache 110 services the request by providing the requested data from the last level cache 110. If the data is not in the last level cache 110, the device 100 services the request from the memory 104. As can be seen, in some embodiments, the last level cache 110 serves as a last cache level before the memory 104, which helps to reduce the total amount of memory access latency for accesses to the memory 104.
[0012] Although last level cache 110 may provide certain benefits, such as improving access latency for frequently accessed or densely populated data, there are situations in which last level cache 110 provides no benefit or little benefit, and in such situations, the power consumed by operating last level cache 110 may be considered wasted.
[0013] Thus, the present disclosure provides techniques for powering down the last level cache 110 in situations where the last level cache 110 provides little or no benefit (the last level cache 110 has "low cache effectiveness"). Generally, according to these techniques, in response to a power-down trigger, the last level cache 110 is powered down. Memory access transactions serviced from data in the last level cache 110 are instead serviced directly by the memory 104. For example, when the last level cache 110 is powered up, if a memory access request fails in a cache at a lower level than the last level cache 110 (such as in the processor 102 or another component), the last level cache 110 checks whether the appropriate data is in the last level cache 110. If the data is present, the last level cache 110 provides the data from the last level cache 110 to the requestor. When the last level cache 110 is powered down and another unit requests access to data that is not stored in a cache at a level lower than the last level cache 110, the last level cache 110 fetches the data from the memory 104 and provides the data to the requestor. The last level cache 110 does not check whether such data is stored in the last level cache 110 when the last level cache 110 is powered down.
[0014] It should be noted that the statement that the last level cache 110 is "powered down" may mean that some or all of the data banks are powered down. However, the cache controller of the last level cache 110 may still perform operations such as communicating data and memory access requests between the last level cache 110 and the memory 104. In other examples, the last level cache 110 being powered down may mean that all circuitry of the last level cache 110 is powered down. In such cases, actions described herein as being performed by the last level cache 110 may be performed by other entities, such as a portion of the memory 104 (e.g., a memory controller), or by a separate unit not within the last level cache 110.
[0015] As described above, the last level cache 110 powers down in response to a power-down trigger. In some examples, the power-down trigger includes a sufficient number of entities serviced by the last level cache 110 being powered down. In some examples, the sufficient number is at least a threshold number of entities. In some examples, the power-down trigger is the device 100 entering a low power mode. In some examples, the power-down trigger occurs when a sufficient number of devices are powered down, including the processor 102 and one or more auxiliary devices 106. In some examples, the sufficient number of devices includes all clients of the last level cache 110 (e.g., the processor 102 and all auxiliary devices 106). Thus, in such examples, the power-down trigger occurs when all clients of the last level cache 110 are powered down. In some examples, "power down" means inactive or idle. In some examples, an operating system, application, or other software executes one or more instructions that instruct the device to enter a low power mode. In some such examples, entering this low power mode is the power-down trigger. In some examples, the power-down trigger includes any action that causes device 100 or a system-on-chip of device 100 to be placed into a low power state. In different embodiments, the power-down trigger is any combination of low power states of any of the elements of device 100.
[0016] In other examples, the power down trigger occurs when an access pattern in the last level cache 110 indicates that the last level cache 110 provides a low benefit to the device 100. In some examples, the last level cache 110 or another entity, such as the processor 102, a memory controller, or some other entity coupled to the last level cache 110, tracks misses and / or hits in the last level cache 110. In some examples, the power down trigger occurs when a miss rate in the last level cache 110 is too high or a hit rate in the last level cache 110 is too low. In some examples, the miss rate is a percentage of misses in the last level cache 110 compared to a total number of access requests to the last level cache 110. In some examples, the hit rate is a percentage of hits in the last level cache 110 compared to a total number of access requests to the last level cache 110. A miss is an access request to the last level cache 110 that results in the requested data not being in the last level cache 110. A hit is an access request where the requested data is in the last level cache 110. In some examples, a miss rate that is too high occurs when the miss rate exceeds a threshold at which the miss rate is considered too high. In some examples, a hit rate that is too low occurs when the hit rate falls below a threshold at which the hit rate is considered too low. An access request is a request from another entity (e.g., processor 102 or auxiliary device 106) to access memory. In some examples, an access request is a request to read from or write to a memory address. In some examples, the request occurs because a miss occurs in all caches at levels lower than last level cache 110 in the cache hierarchy.
[0017] 2 and 3 illustrate operations for powering down the last level cache 110 in response to a power-down trigger. In FIG. 2, the power-down trigger is that a sufficient number of devices (e.g., the processor 102 and / or one or more auxiliary devices 106) are powered down. As described above, any or all of these devices may be powered down for any reason, such as in response to a request by software, such as an operating system, an application, a driver, or other software running on the processor 102 or other devices. In FIG. 3, the power-down trigger is that the access patterns of the last level cache 110 indicate that there is little or no benefit to powering on the last level cache 110. There are various possible workload types that would not benefit from the last level cache 100, at least some of the time. Some examples include workloads that do not exhibit a large degree of spatial and / or temporal memory access locality. Spatial memory access locality means that memory accesses are made to memory addresses that are close to each other. Such access patterns would benefit from the last level cache 110 because misses result in cache lines being placed in the last level cache 110. Therefore, if a subsequent access request is close to the first access request that causes a miss, the subsequent access is likely to be to the same cache line that is already in the cache. Temporal memory access locality means that memory accesses are made close together in time. If two accesses are made close together in time to the same cache line, the subsequent access is likely to hit in the cache. The longer it takes to reaccess the same cache line, the more likely it is that the cache line has been removed from the cache, for example due to an eviction.
[0018] In some embodiments, the last level cache 110 uses the power-down triggers of FIG. 2 (powering down the processor 102 and / or auxiliary device 106) but does not use the power-down triggers of FIG. 3 (where the access pattern to the last level cache 110 causes a power-down). In other embodiments, the last level cache 110 uses the power-down triggers of FIG. 3 but does not use the power-down triggers of FIG. 2. In other embodiments, the last level cache 110 uses both the power-down triggers of FIG. 2 and the power-down triggers of FIG. 3. In some embodiments, the last level cache 110 operates in a first mode in which the last level cache 110 uses the power-down triggers of FIG. 2 but does not use the power-down triggers of FIG. 3, and then switches to a second mode in which the last level cache 110 uses the power-down triggers of FIG. 3 but does not use the power-down triggers of FIG. 2. In some embodiments, the last level cache 110 operates in a third mode in which the last level cache 110 uses both the power-down triggers of FIG. 2 and the power-down triggers of FIG. 3. In some embodiments, last level cache 110 operates in either the first mode, the second mode, or the third mode. In some examples, last level cache 110 switches between the first mode, the second mode, and the third mode.
[0019] 4A and 4B illustrate the operation of the device 100 when the last level cache 110 is powered down (FIG. 4A) and when the last level cache 110 is powered up (FIG. 4B). In FIG. 4A, in the situation where the last level cache 110 is powered down, the last level cache 110 does not provide a response to a memory access request from data in the last level cache 110. The last level cache 110 is bypassed and transactions that would be serviced by the last level cache 110 are instead serviced directly by the memory 104. In some examples, the last level cache 110 is completely powered down and no components of the last level cache 110 process memory access transactions. In other examples, a portion of the last level cache 110 is powered down while a portion of the last level cache 110 remains powered up. In one example, one or more data banks of the last level cache 110 are powered down and one or more control circuits (such as a cache controller) remain powered. In some examples, the one or more control circuits treat all received memory access requests as misses and forward such requests to memory 104. In such examples, the one or more control circuits do not access the data banks of last level cache 110 because those banks are powered down.
[0020] 4B illustrates the device 100 in a mode in which the last level cache 110 is powered. In this mode, the last level cache 110 services transactions received from other components. Servicing some such transactions may require communication between the last level cache 110 and the memory 104. For example, if a miss occurs in the last level cache 110, the last level cache 110 fetches the cache line on which the miss occurred from the memory 104. In another example, if a dirty cache line is evicted from the last level cache 110, the last level cache 110 sends the cache line to the memory 104 to be written.
[0021] In response to a power-down trigger, the device 100 powers down the last level cache 110, as described elsewhere herein. Figures 5A-5D illustrate a technique for powering down the last level cache 110, according to one example. In these figures, the last level cache 110 includes a number of cache lines 508, each of which includes at least data 504 and state 506. For any cache line whose state is dirty, the last level cache 110 flushes the cache line to memory 104. This flush copies the contents of the cache line to memory. A cache line is dirty if it has been modified compared to the copy stored in memory 104.
[0022] Because flushing the entire last level cache 110 takes a significant amount of time, the last level cache 110 may continue to service memory requests while it is powered down. The power-down walker 502 walks through the cache lines and flushes the cache lines to the memory 104. For cache lines that have already been flushed, such cache lines have an invalid state and therefore cannot be used to service memory access requests. However, cache lines that have not yet been flushed may be used to service memory access requests. For read requests that hit on cache lines that have not yet been reached by the power-down walker 502 (and therefore are not invalid), such requests are serviced by providing the data in the cache line as a response to the request. For write requests that hit on cache lines that have not yet been reached by the power-down walker 502 (and therefore are not invalid), such requests are serviced by modifying the data in the cache line. If the cache line has not been marked as dirty before servicing the write request, the cache line is marked as dirty. In some examples, the last level cache 110 flushes the cache line to the memory 104 when the power down walker 502 arrives at that cache line 508. In other examples, the last level cache 110 flushes the cache line as soon as it performs a write, regardless of the location of the power down walker 502, resulting in the cache line becoming invalid in the last level cache 110. If a miss occurs for a memory access request while the last level cache 110 is powered down, the last level cache 110 passes the memory access request to the memory 104. The last level cache 110 does not allocate cache lines to the last level cache 110 during power down.In other words, if a memory access request misses in the last level cache 110 while the last level cache 110 is powered down, the last level cache 110 does not fetch the associated cache line from the memory 104 to the last level cache 110. Instead, the last level cache 110 passes the request to the memory 104.
[0023] 5A-5D show an exemplary sequence of operations for powering down last level cache 110. In FIG. 5A, power down walker 502 is flushing cache line 508(1). The lines past cache line 508(1), 508(2), 508(3), and 508(4), are valid. Any such lines can service incoming requests. In FIG. 5B, power down walker 502 is invalidating cache line 508(1) and flushing cache line 508(2). Cache lines 508(3) and 508(4) can service incoming requests, but cache line 508(1) cannot service incoming requests. In FIG. 5C, power down walker 502 is proceeding to cache line 508(3). Power down walker 502 flushes the contents of cache line 508(3) to memory 104 and invalidates cache line 508(3). 5D, power-down walker 502 walks to cache line 508(4), but because the cache line is not dirty, power-down walker 502 does not flush the cache line. However, power-down walker 502 invalidates the cache line.
[0024] When the last level cache 110 is operating in a power-down mode and the device 100 detects a power-up trigger, the device powers up the last level cache 110 again. In some examples, such as when the power-down trigger that powered down the last level cache 110 is that a sufficient number of entities are powered down, the power-up trigger is that at least some of those entities are powered up. In some examples, the power-up trigger is that all of those entities are powered up. In some examples, when the power-down trigger is that the access patterns of the last level cache 110 indicate that the last level cache 110 is not providing sufficient benefit (such as when the miss rate of the last level cache 110 is too high or the hit rate of the last level cache 110 is too low), the power-up trigger is that the device 100 is detecting that it is running with substantially different workload characteristics than when the last level cache 110 was originally powered down. In other words, in some examples, in a situation where the device 100 powers down the last level cache 110 as a result of the last level cache 110 not providing sufficient benefit, the device 100 powers up the last level cache 110 again in response to detecting that the workload characteristics have changed. The change in the workload characteristics indicates that the last level cache 110 is again able to provide substantial benefit to the operation of the device 100. For example, if the last level cache 110 is powered down due to not providing sufficient benefit for a first workload, when the workload characteristics change, the last level cache 110 may be able to provide sufficient benefit to the operation of the device resulting in the new workload characteristics. In more simple terms, the last level cache 110 may provide benefit to a second workload even if the last level cache 110 is not providing benefit to a first workload.The change in workload characteristics is a hint that a new workload may be running on the device 100 and that the last level cache 110 should be powered on again to benefit the new workload. If the last level cache 110 does not provide sufficient benefit for this new workload, the last level cache 110 can be powered down again.
[0025] There are various ways in which the device 100 can detect a change in workload as a power-up trigger for the last level cache 110. In one example, a change in resource utilization is a power-up trigger. In various examples, the change in resource utilization is a change in the amount of memory used or a change in the load on the processor 102. In another example, a power-up trigger is that there is a change in which application has user focus. An application has user focus if the user is actively interacting with the application (e.g., a web browser has user focus if the user is browsing a web page and providing mouse clicks and keyboard input to the web browser). A change in which application has user focus indicates that the workload has likely changed. Another power-up trigger is a particular auxiliary device 106 going from low utilization to high utilization or from high utilization to low utilization. In some examples, going from low utilization to high utilization means that the utilization of the auxiliary device 106 is above a threshold. In some examples, going from high utilization to low utilization means that the utilization of the auxiliary device 106 is below a threshold. In some examples, the utilization of the auxiliary device refers to the amount of work being performed on the auxiliary device compared to the total amount of work that could be performed on the auxiliary device. In some examples, the auxiliary device 106 is a video decoder that can handle decoding four different streams of video in parallel at very high resolution with high picture settings, and an example of low utilization is decoding one stream at low resolution with low picture settings. In another example, low utilization means that the video decoder is not active at all. A change in the utilization of the auxiliary device is a hint that the device 100 is running a different workload and therefore the last level cache 110 may be useful for such a new workload. Any other change in the operating characteristics of the device 100 may be a power-up trigger.
[0026] When a power-up trigger is detected, device 100 powers up last level cache 110. In some examples, an initialization sequence is required after powering up the cache before the cache can begin to be used. In some examples, powering up last level cache 110 includes powering up powered-down banks and beginning to operate normally as a cache (e.g., servicing requests from other entities within device 100, including fetching cache lines from memory 104 in response to misses and placing those cache lines in last level cache 110).
[0027] 6 illustrates a method 600 for operating the last level cache 110, according to one example. Although described with respect to the systems of FIGS. 1-5D, one skilled in the art will recognize that any system configured to perform the steps of method 600 in any technically feasible order is within the scope of the present disclosure.
[0028] In step 602, device 100 detects a power-down trigger, which may be any of the power-down triggers described herein, such as a sufficient number of components of device 100 being powered down, or detecting an access pattern of last level cache 110 that indicates operation of last level cache 110 is not providing significant benefit.
[0029] In response to the power-down trigger, device 100 powers down the last level cache at step 604. In some examples, this power-down occurs as described elsewhere herein, including with respect to Figures 5A-5D.
[0030] In step 606, device 100 detects a power-up trigger. In various examples, the power-up trigger is any of the power-up triggers described elsewhere herein. In step 608, device 100 powers up last level cache 110 in response to the power-up trigger. In various examples, powering up last level cache 110 includes placing components of last level cache 110 in an operational state, allowing last level cache 110 to again service memory access requests.
[0031] Elements in the figures may be embodied as software running on a processor, fixed function processor, programmable processor, or combination thereof, where appropriate. The processor 102, last level cache 110, interconnect 112, memory 104, storage 108, and various auxiliary devices 106 include at least some hardware circuitry and, in some embodiments, software running on a processor within that component or within another component. Certain elements of the last level cache 110 are illustrated in Figures 5A-5D. Cache line 508 represents a portion of the memory bank of the last level cache 110 sized to store a cache line. Data 504 and state 506 are portions of the cache line 508 designed to store data or state.
[0032] It should be understood that many variations are possible based on the disclosure herein, and although features and elements are described above in particular combinations, each feature or element can be used alone without other features and elements, or in various combinations, with or without other features and elements.
[0033] The provided methods may be implemented in a general purpose computer, processor, or processor core. Suitable processors include, by way of example, general purpose processors, special purpose processors, conventional processors, digital signal processors (DSPs), multiple microprocessors, one or more microprocessors associated with a DSP core, controllers, microcontrollers, Application Specific Integrated Circuits (ASICs), Field Programmable Gate Array (FPGA) circuits, any other type of integrated circuit (IC), and / or state machines. Such processors may be manufactured by configuring a manufacturing process using the results of processed hardware description language (HDL) instructions and other intermediate data, including netlists (instructions that may be stored on a computer readable medium). The results of such processing may be a mask work that is used in a subsequent semiconductor manufacturing process to manufacture a processor implementing features of the present disclosure.
[0034] The methods or flow charts provided herein may be implemented in a computer program, software, or firmware embodied in a non-transitory computer-readable storage medium for execution by a general purpose computer or processor. Examples of non-transitory computer-readable storage media include read only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM disks and digital versatile disks (DVDs).
Claims
1. 1. A method of operating a cache, comprising: powering down the cache in response to a power-down trigger indicating that the cache is deemed to be low in validity. method.
2. and powering the cache back on in response to a power-up trigger.
2. The method of claim 1.
3. the power-down trigger includes one or more components of a device including the cache being powered down; 2. The method of claim 1.
4. the power down trigger comprises an access pattern to the cache indicative of low cache utilization; 2. The method of claim 1.
5. the access pattern includes a miss rate in the cache above a threshold; The method of claim 4.
6. the access pattern includes a hit rate in the cache being below a threshold; The method of claim 4.
7. the power-up trigger includes one or more components of a device including the cache being powered up; The method of claim 2.
8. the power-up trigger comprises a device including the cache having a substantial change to an operating characteristic; The method of claim 2.
9. powering down the cache includes flushing lines of the cache that are indicated as dirty.
2. The method of claim 1.
10. 1. A system comprising: Cache and a processor configured to execute instructions to store data in the cache; The cache is configured to be powered down in response to a power down trigger indicating that the cache is deemed to be low in validity. system.
11. the cache is configured to power on in response to a power-up trigger; The system of claim 10.
12. the power-down trigger comprises one or more components of a device including the cache being powered down; The system of claim 10.
13. the power down trigger comprises an access pattern to the cache indicative of low cache utilization; The system of claim 10.
14. the access pattern includes a miss rate in the cache above a threshold; The system of claim 13.
15. the access pattern includes a hit rate in the cache being below a threshold; The system of claim 13.
16. the power-up trigger includes one or more components of a device including the cache being powered up; The system of claim 11.
17. the power-up trigger comprises a device including the cache having a substantial change to an operating characteristic; The system of claim 11.
18. powering down the cache includes flushing lines of the cache that are indicated as dirty. The system of claim 10.
19. 1. A system comprising: A processor; A plurality of auxiliary devices; a cache; the processor and the plurality of auxiliary devices are clients of the cache; The cache is configured to be powered down in response to a power down trigger indicating that the cache is deemed to be low in validity. system.
20. The cache is configured to power on again in response to a power-up trigger.
20. The system of claim 19.