Live Profile Driven Cache Aging Policy
The dynamic cache policy controller optimizes cache management by leveraging access history to adapt cache policies for recurring access patterns, improving performance by enhancing hit rates and reducing latency.
Patent Information
- Application Number
- JP2024575312
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-06-27
- Filing Date
- 2023-05-23
- Publication Date
- 2025-07-03
AI Technical Summary
Existing cache technologies suffer from inaccuracies in cache replacement policies due to overly general rules that do not account for specific data access patterns, leading to suboptimal performance.
A dynamic cache policy controller that records access data for a first frame, identifies parameters based on access patterns, and applies these parameters to subsequent frames to optimize cache management policies, including how to age cache lines and handle hits/misses based on observed access history.
Improves cache performance by dynamically adapting to recurring access patterns, enhancing hit rates and reducing memory access latency through targeted cache line management.
Smart Images

Figure 2025520667000001_ABST
Abstract
Description
Technical Field
[0001] (Cross - Reference to Related Applications) This application claims the benefit of U.S. Patent Application No. 17 / 850,905, filed on June 27, 2022, the entire disclosure of which is incorporated herein by reference.
Background Art
[0002] Caches improve performance by storing copies of data that is likely to be accessed again in the future in low - latency cache memory. Improvements to cache technology are constantly being made.
[0003] A more detailed understanding can be obtained from the following description given by way of example in conjunction with the accompanying drawings.
Brief Description of the Drawings
[0004]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Modes for Carrying Out the Invention
[0005] Techniques for operating a cache are disclosed. The techniques include recording access data for a first set of memory accesses of a first frame, identifying parameters for a second set of memory accesses of a second frame subsequent to the first frame based on the access data, and applying the parameters to the second set of memory accesses.
[0006] FIG. 1 is a block diagram of an exemplary computing device 100 that can implement one or more features of the present disclosure. In various examples, the computing device 100 can be, for example, but not limited to, any one of a computer, a gaming device, a handheld device, a set-top box, a television, a cellular phone, a tablet computer, or other computing device. The device 100 includes, without limitation, one or more processors 102, a memory 104, one or more auxiliary devices 106, a storage device 108, and a last level cache (LLC) 110. An interconnect 112, which can be a bus, a combination of buses, and / or any other communication component, communicatively links the one or more processors 102, the memory 104, the one or more auxiliary devices 106, the storage device 108, and the last level cache 110.
[0007] In various alternative forms, one or more processors 102 include a central processing unit (CPU), a graphics processing unit (GPU), a CPU and a GPU on the same die, or one or more processor cores, where each processor core can be a CPU, a GPU, or a neural processor. In various alternative forms, at least a portion of the memory 104 is on the same die as one or more of the one or more processors 102, such as on the same chip or within an interposer configuration, and / or at least a portion of the memory 104 is mounted independently of the one or more processors 102. The memory 104 includes volatile or non-volatile memory (e.g., random access memory (RAM), dynamic RAM, cache).
[0008] The storage device 108 includes a fixed or removable storage device (e.g., but not limited to, a hard disk drive, a solid state drive, an optical disk, a flash drive). The one or more auxiliary devices 106 include, but are not limited to, one or more auxiliary processors 114 and / or one or more input / output (I / O) devices. The auxiliary processor 114 includes, but is not limited to, a processing unit capable of executing instructions, such as a central processing unit, a graphics processing unit, a parallel processing unit capable of performing compute shader operations in a single instruction multiple data format, a multimedia accelerator such as a video encoding or decoding accelerator, or any other processor. Any auxiliary processor 114 can be implemented as a programmable processor that executes instructions, a fixed function processor that processes data according to fixed hardware circuitry, a combination thereof, or any other type of processor. The auxiliary processor 114 includes an accelerated processing device (APD) 116.
[0009] One or more IO devices 118 include one or more input devices such as a keyboard, keypad, touch screen, touch pad, detector, microphone, accelerometer, gyroscope, biometric scanner, or a network connection (e.g., a wireless local area network card for transmitting and / or receiving wireless IEEE802 signals), and / or one or more output devices such as a display, speaker, printer, tactile feedback device, one or more lights, antenna, or a network connection (e.g., a wireless local area network card for transmitting and / or receiving wireless IEEE802 signals).
[0010] The last-level cache 110 functions as a shared cache for various components of the device 100 such as the processor 102, APD 116, and various auxiliary devices 106. In some embodiments, there are other caches within the device 100. For example, in some instances, the processor 102 includes a cache hierarchy that includes different levels such as level 1 and level 2. In some instances, each such cache level is specific to a particular logical partition of the processor 102, such as a processor core or a chip, die, or package of the processor. In some instances, the hierarchy includes other types of caches. In various instances, one or more of the auxiliary devices 106 include one or more caches.
[0011] In some examples, the last-level cache 110 is "last-level" in the sense that such a cache is the last cache that the device 100 attempts to service memory access requests before servicing requests from the memory 104 itself. For example, when the processor 102 accesses data that is not stored in any of the cache levels of the processor 102, the processor exports a memory access request to be satisfied by the last-level cache 110. The last-level cache 110 determines whether the requested data is stored in the last-level cache 110. If the data is within the last-level cache 110, the last-level cache 110 services the request by providing the requested data from the last-level cache 110. If the data is not within the last-level cache 110, the device 100 services the request from the memory 104. As can be seen, in some embodiments, the last-level cache 110 acts as the last cache level before the memory 104, which helps reduce the total amount of memory access latency for accesses to the memory 104. Although techniques for operations with the last-level cache 110 are described herein, it should be understood that the techniques can alternatively be used in other types of caches or memories as well.
[0012] FIG. 2 is a block diagram of a device 100 showing additional details regarding the execution of processing tasks on an APD 116, according to an example. A processor 102 maintains, within a system memory 104, one or more control logic modules for execution by the processor 102. The control logic modules include an operating system 120, a driver 122, and an application 126, and may optionally include other modules not shown. These control logic modules control various aspects of the operation of the processor 102 and the APD 116. For example, the operating system 120 communicates directly with the hardware and provides an interface to the hardware for other software executed on the processor 102. The driver 122 controls the operation of the APD 116, for example, by providing an application programming interface (“API”) to software (such as application 126) executed on the processor 102 to access various functions of the APD 116. The driver 122 also includes a just-in-time compiler that compiles shader code into a shader program for execution by processing components of the APD 116 (such as the SIMD unit 138, described in more detail below).
[0013] APD116 executes commands and programs for selected functions such as graphics operations and non-graphics operations that may be suitable for parallel processing. Based on the commands received from processor 102, APD116 can be used to perform graphics pipeline operations such as pixel operations, geometric calculations, and rendering of images to the IO device 118. Also, based on the commands received from processor 102, APD116 performs computational processing operations not directly related to graphics operations, such as operations related to video, physical simulation, computational fluid dynamics, or other tasks, or computational processing operations that are not part of the "normal" information flow of the graphics processing pipeline, or computational processing operations that are not at all related to graphics operations (sometimes referred to as "GPGPU" or "general-purpose graphics processing unit").
[0014] APD116 includes a computing unit 132 (sometimes collectively referred to herein as a "programmable processing unit") that includes one or more SIMD units 138 configured to perform operations in a parallel fashion according to the SIMD paradigm. The SIMD paradigm is one in which multiple processing elements share a single program control flow unit and program counter, and thus can execute the same program but with different data. In one example, each SIMD unit 138 includes 16 lanes, and each lane can execute the same instruction simultaneously with other lanes within the SIMD unit 138, but with different data for that instruction. The lanes can be switched off with predication (prediction) if not all lanes need to execute a given instruction. Also, predication can be used to execute programs with a branch control flow. More specifically, for a program with conditional branch instructions, or other instructions where the control flow is based on calculations performed by individual lanes, predication of the lanes corresponding to the currently non-executed control flow path and serial execution of different control flow paths allow following any control flow.
[0015] The basic unit of execution in the computing unit 132 is a work item. Each work item represents a single instantiation of a shader program that is executed in parallel in a specific lane of a wavefront. Work items can be executed simultaneously as a "wavefront" on a single SIMD unit 138. Multiple wavefronts can be included in a "work group", which includes a collection of work items designated to execute the same program. A work group can be executed by executing each of the wavefronts that make up the work group. Wavefronts can be executed continuously on a single SIMD unit 138, or partially or fully in parallel on various SIMD units 138. A wavefront can be considered an instance that executes a shader program in parallel, but each wavefront includes multiple work items that are executed simultaneously on a single SIMD unit 138 according to the SIMD paradigm (e.g., one instruction control unit that executes a stream of the same instructions for multiple data). The command processor 137 exists within the computing unit 132 and launches wavefronts based on work (e.g., execution tasks) that are waiting for completion. The scheduler 136 is configured to perform operations related to the scheduling of various wavefronts on various computing units 132 and SIMD units 138.
[0016] The parallelism provided by the computing unit 132 is suitable for graphics-related operations such as pixel value calculation, vertex transformation, tessellation, geometry shading operations, and other graphics operations. Therefore, the graphics processing pipeline 134 that receives graphics processing commands from the processor 102 provides computing tasks to the computing unit 132 for parallel execution.
[0017] In addition, the computing unit 132 is also used to perform computational tasks that are not related to graphics or not part of the "normal" operations of the graphics processing pipeline 134 (e.g., custom operations performed to supplement the operations of the graphics processing pipeline 134). An application 126 or other software executed on the processor 102 transmits a program (often referred to as a "compute shader program" and which can be compiled by the driver 122) that defines such a computational task to the APD 116 for execution. Although the APD 116 having the graphics processing pipeline 134 is shown, the teachings of the present disclosure are also applicable to the APD 116 that does not have the graphics processing pipeline 134. Various entities of the APD 116, such as the computing unit 132, can access the last-level cache 110.
[0018] FIG. 3 shows a dynamic cache policy system 300 for a cache 302 according to an example. The system 300 includes an APD 116, a cache 302, and a dynamic cache policy controller 304. In some examples, the cache 302 is the last-level cache 110 of FIG. 1. In other examples, the cache 302 is any other cache that services memory access requests by the APD 116. In some examples, the dynamic cache policy controller 304 is a hardware circuit, software executed on a processor, or a combination thereof. In some examples, the dynamic cache policy controller 304 is included in the cache 302 or is independent of the cache.
[0019] The APD 116 processes a series of frames 306 of data. Each frame 306 is a full image rendered by the APD 116. In order to generate graphical data over time, the APD 116 renders a series of frames, each of which includes graphical objects at a particular instant.
[0020] To generate each frame 306, various components of the APD 116 access the data stored in the memory. These accesses are serviced by a cache 302 that stores a copy of the accessed data. The accessed data includes a variety of data types for rendering graphics, such as geometry data (e.g., vertex coordinates and attributes), texture data, pixel data, or a wide variety of other types of data. Aspects of cache performance, such as the hit rate (the ratio of hits in the cache to the total number of accesses), depend in part on how the cache 302 holds data in the case of eviction. The cache replacement policy indicates which cache line to evict when a new cache line is brought into the cache. A variety of replacement policies and other techniques are known. However, many of these techniques suffer from inaccuracies due to overly general rules that may not apply to specific data items. For example, a policy of replacing the least recently used one is not appropriate for all data access patterns.
[0021] For at least these reasons, the dynamic cache policy controller 304 dynamically applies a cache management policy to access requests from the APD 116 based on the observed access history. The dynamic cache policy controller 304 maintains cache usage data that records memory access information across frames. The memory access information is related to aspects of memory access such as the re-reference interval and re-reference intensity for a memory page. The re-reference interval is the "time" (e.g., number of clock cycles, number of instructions, or other measure of "time") between accesses to a particular memory page. The re-reference intensity is the frequency of access to a particular memory page within a "time" of a given length (e.g., number of clock cycles, number of instructions, or other measure of time).
[0022] Based on the memory access information from the previous frame, the dynamic cache policy controller 304 applies a memory access policy to the memory access requests from the APD 116 in the current frame. More specifically, the pattern of memory access is generally very similar between frames. Therefore, it is possible to observe the memory accesses in a certain frame and use those observations to control the memory access policy for those accesses in subsequent frames. In other words, based on the observations made for a specific memory access in a specific frame, the dynamic cache policy controller 304 controls the memory access policy for the same memory access in one or more subsequent frames. Memory accesses are considered "the same" between frames when they occur at approximately the same location in the order of frame memory accesses. More specifically, generally, adjacent frames in time render very similar graphical information, and as a result, have very similar patterns of memory access. The term "approximately" means that the frames can be somewhat different and the order of memory access between frames is not the same, allowing for the possibility that some memory accesses occurring in one frame may not occur in another frame.
[0023] The memory access policy is a policy that controls how the cache 302 manages information about the cache lines targeted by memory access requests. In one example, the memory access policy controls three aspects, namely, the age (e.g., as a result of a miss) when a new cache line is brought into the cache, the manner in which the age of the cache line is updated when a miss occurs (e.g., how cache lines other than the accessed cache line are aged), and the age at which the cache line is updated when a hit occurs in the cache line (e.g., what age the cache line is set to when accessed). These features are described in more detail elsewhere in this specification.
[0024] Note that the age to be referenced is the age in the policy of replacing the cache with the least recent usage. More specifically, cache 302 is a set-associative cache that includes several sets each having multiple ways. Any given cache line is mapped to at least one set and can be placed in any way within that set, but cannot be placed within any set to which the cache line is not mapped. A cache line is defined as part of the memory address space. In one example, each cache line has a cache line address that includes a specific number of upper bits of the address, with the remaining lower bits specifying the offset within that cache line. Typically, but not necessarily, a cache line is the minimum unit of data that is read into the cache or written back from the cache. If a cache line is brought into a set of cache 302 and there is no empty way for the cache line, cache 302 evicts one of the cache lines within the set. The replacement algorithm selects, as the cache line to evict, the oldest cache line based on the age of the cache lines within that set. Also, cache 302 inserts a new cache line into cache 302 at the "start age". In the case of a hit, cache 302 updates the age of the cache line where the hit occurred based on the memory access policy. Further details are provided elsewhere in this specification.
[0025] FIG. 4 shows additional details of the dynamic cache policy controller 304 according to an example. The dynamic cache policy controller 304 stores and maintains cache usage data 402. The cache usage data 402 includes memory access information regarding a "set" of memory accesses in a frame. Each access data item 406 includes information characterizing a particular set of one or more memory accesses. In some examples, each set of memory accesses is one or more memory accesses that occur within a particular length of "time" and within a certain memory address range (such as a page). In some examples, the time is represented by real time, cycle count, instruction count, or any other measure of time. Thus, each data item 406 includes information regarding memory accesses that occur within a particular period within the frame and to the same memory address range. In other words, in different examples, accesses that occur within a threshold number of instructions, a threshold number of computer clock cycles, or a threshold length of time are grouped together into a set corresponding to the access data item 406.
[0026] In some examples, the dynamic cache policy controller 304 records the access data item 406 when the corresponding access occurs. In some examples, the dynamic cache policy controller 304 calculates additional information for each access data item 406 in response to the end of a frame or at some other point in time (e.g., when the "time" for the access data item 406 has ended). In one example, the dynamic cache policy controller 304 determines the reference interval and reference strength of the access corresponding to the access data item 406. Then, the dynamic cache policy controller 304 determines the policy corresponding to the access data item 406 according to the reference interval and reference strength of the access corresponding to the access data item 406. In some examples, the dynamic cache policy controller 304 determines the policy based on the reference intervals and reference strengths from multiple frames. In some examples, the dynamic cache policy controller 304 records this policy in the access data item 406 for use in subsequent frames. Thus, the access data item 406 indicates which policy to use for a particular page and a particular "time" in the frame.
[0027] Based on this policy information, the dynamic cache policy controller 304 can "condition" accesses made in frames subsequent to the frame in which the access data item 406 is recorded. "Conditioning" a memory access means setting parameters for the memory access, and the parameters correspond to the determined re-reference interval value and the determined re-reference intensity value. For example, in a first frame, the dynamic cache policy controller 304 records a first access data item 406 for a set of memory accesses that access the same page. The first access data item 406 indicates a specific policy for that set of memory accesses. In subsequent frames, the dynamic cache policy controller 304 identifies the same access (i.e., the access of the first access data item 406) and "conditions" those accesses according to the policy stored in the first access data item 406. More specifically, the dynamic cache policy controller 304 configures those accesses to use the policy corresponding to the first access data item 406. Such a policy indicates that these accesses in subsequent frames should be made according to a specific set of configurations. In some examples, the policy indicates how to age cache lines within the same set as the cache line accessed in case of a miss, at what age to insert a new cache line into the cache lines, and what age to set for the cache line in case of a hit.
[0028] Next, the aging of cache lines in case of a miss will be described. When a miss occurs in cache 302, cache 302 identifies a cache line and brings it into the cache. Also, cache 302 determines into which set the identified cache line should be brought. If any cache line within that set has an age equal to or greater than a threshold value (e.g., 3 when the age counter is 2 bits), cache 302 selects that cache line for eviction. If there is no cache line with an age higher than the threshold value, cache 302 ages the cache lines of the set. The setting of how cache lines are aged in case of a miss indicates which cache line of the set will be aged when a miss occurs in that set and there is no cache line with an age equal to or greater than the threshold value. In some examples, this setting indicates for each age which cache line will be aged in case of a miss. In other words, the setting indicates the age of the cache line that will be aged in case of a miss. In an example of the setting, the setting indicates that cache lines of all ages less than the above threshold value ("eviction threshold") will be aged. In another example, the setting indicates that cache lines with an age exceeding an age trigger threshold value that may be lower than the eviction threshold will be aged. In this situation, cache lines with an age below the lower threshold value will not be aged in case of a miss. In short, in some examples, the reuse interval and reuse intensity of a set of memory accesses for a certain frame indicate how cache lines will be aged (specifically, the age of the cache line that will be aged) in case of a miss for a set of memory accesses in a subsequent frame.
[0029] Next, the setting regarding at what age to insert a new cache line into the cache will be described. When a cache line is brought into the cache as a result of a miss, a specific age is first given to the cache line. In some examples, this "starting" age is the minimum possible age, some intermediate age above the minimum possible age, or the maximum age. Also in this case, this setting depends on the memory access reflected in the access data item 406. Therefore, the access data item 406 corresponding to the memory page indicates the starting age of the cache line when a miss occurs and the cache line is copied into the cache.
[0030] Next, what age to set for the new cache line when a hit occurs will be described. When a hit occurs in a cache line, the cache 302 updates the age of that cache line (e.g., to indicate that the cache line is "newer"). In some examples, this setting indicates that the cache line will have a specific age (such as 0) when a hit occurs for that cache line. In other examples, the setting indicates that the cache line will have different ages such as 1, 2, or 3 (in the case of a 2-bit age counter) when a hit occurs. In other examples, this setting indicates that the age of the cache line is changed in a specific manner, such as by decrementing the age by a number like 1. In short, the access data item 406 corresponding to the memory page indicates how the age of the cache line is changed when a hit occurs in that cache line.
[0031] It should be understood that conditioning a particular cache access according to a policy means causing a cache access by that policy. In the case of an insertion age policy, this policy is applied to the access conditioned according to the policy if the access ends up as a miss. The cache line to be brought in is brought in according to the age specified by the policy. In the case of an aging policy, this occurs for the access conditioned according to the policy if the access ends up as a miss. In this situation, the aging policy ages other cache lines for the same set as specified by the policy. In the case of a policy that defines what age a cache line is set to in the case of a hit, when an access conditioned according to the policy occurs and that access ends up as a hit, the policy causes the age of the cache line being accessed to be set according to the policy.
[0032] Figure 5 shows, by way of example, an operation for recording an access pattern. Figure 5 shows a single frame, frame 1 508(1), in which several accesses 512 are occurring. The first set of accesses, accesses 512(1) to 512(4), are for the first page P1 and occur within a particular period. The dynamic cache policy controller 304 identifies a set of three accesses 502 and generates an access data item 506 corresponding to each of the set of accesses 502. These data items indicate the access time and the address being accessed (including or represented as a page). For the first set 502(1), the dynamic cache policy controller 304 notes that the set has a high reference intensity and a low re-reference interval. Since there are many accesses to the same page, there is a high re-reference intensity, and since the accesses occur relatively close to each other "in time", there is a short reference interval. Therefore, the dynamic cache policy controller 304 records in the access data item 506(1) an access data item 506(1) indicating the policy associated with this re-reference interval and re-reference intensity.
[0033] For the second set of accesses 512(2) that includes access 502(5) and is performed on page P2, there are a low re-reference intensity and a short re-reference interval. Therefore, the dynamic cache policy controller 304 records the policy associated with this combination in the access data item 506(2). For the third set of accesses 502(3) that includes accesses having a low re-reference intensity and a long re-reference interval performed on page P3, as reflected in accesses 502(6) to 502(7), the dynamic cache policy controller 304 records an access data item 506(3) that records the policy reflecting the re-reference intensity and re-reference interval of the set 502(3).
[0034] FIG. 6 is a block diagram illustrating the use of access data items to condition access in a second frame 508(2) according to an example. Three sets 602 of data access 612 are shown. The dynamic cache policy controller 304 conditions these accesses according to the access data items 506 generated in frame 1 508(1) of FIG. 5. For the first set 602 of data access 612, the dynamic cache policy controller 304 identifies a first access data item 506(1). This access data item 506 indicates a particular manner of conditioning the set 602(1) of accesses 612. As described elsewhere herein, the access data item 506 indicates a policy that indicates one or more of how to age cache lines in the event of a miss, what age to set new cache lines to, and what age to set cache lines to in the event of a hit. In addition, this policy depends on the data recorded in frame 1 508(1) for the same access from that frame. The dynamic cache policy controller 304 causes the policy applied to an access of a given set 602 to be applied based on the access data item 506 recorded for the same access in the previous frame. Similar activity is applied to set 602(2) and set 602(3).
[0035] The dynamic cache policy controller 304 determines, in the following manner, whether any access in a particular frame is "the same" as an access in the previous frame. In each frame, memory accesses occur in a particular sequence. This sequence is, in most cases, repeated between frames. Thus, the dynamic cache policy controller 304 tracks the sequence of memory accesses to identify which accesses are "the same" between frames. It is true that some accesses may differ between frames, and thus the dynamic cache policy controller 304 takes such differences into account. For example, the dynamic cache policy controller 304 can be notified of omitted accesses, added accesses, or other changes and can take them into consideration. In some examples, memory accesses that occur at the same "time" in different frames to the same page are considered to be "the same" memory access. In the example, "time" is defined based on the access order. For example, the first 100 accesses are the first "time", the 101st to 200th accesses are the second "time", and so on. Any other technically feasible means for determining time are possible.
[0036] It should be understood that the manner in which an access is "conditioned" is based on the access data item 506 recorded in the previous frame. Thus, for a particular access data item 506 indicating a particular combination of reference interval and reference intensity, a first set of parameters including the manner of aging cache lines, the age at which new cache lines are inserted, and the age at which a hit cache line is updated is used for the corresponding access. For another particular access data item 506 indicating a different combination of reference interval and reference intensity, the corresponding access is performed using a second set of parameters including the manner of aging cache lines, the age at which new cache lines are inserted, and the age at which a hit cache line is updated. At least one of the parameters in the second set is different from at least one of the parameters in the first set.
[0037] FIG. 7 is a flowchart of a method 700 for managing memory access according to an example. Although described with respect to the systems of FIGS. 1 - 6, one of ordinary skill in the art will understand that any system configured to perform the steps of method 700 in any technically feasible order is within the scope of the present disclosure.
[0038] In step 702, the dynamic cache policy controller 304 records an access data item 406 for a memory access of a first frame. Each access data item 406 corresponds to a set of one or more memory accesses. In some examples, each set of memory accesses shares a memory page or shares different subsets of memory. Each access data item 406 includes information characterizing the set of memory accesses corresponding to that access data item 406. In some examples, this information is associated with one or both of a reuse interval (e.g., the "time" between references to the same address or page) or a reuse intensity (e.g., the number of times accessed again within a particular window of "time" meaning number of instructions, clock time, clock cycles, or other measure of time). In some examples, the dynamic cache policy controller 304 records access data items 406 for multiple sets of accesses within a frame.
[0039] In step 704, for a second frame subsequent to the first frame, the dynamic cache policy controller 304 identifies parameters for a memory access corresponding to the access for which the access data is recorded in the first frame. The parameters include information indicating how such a memory access should be conditioned if provided to the cache 302 to be satisfied. In some examples, the memory access for which the parameters are identified is the same memory access for which the corresponding access data is stored in the first frame. In other words, in some examples, in step 704, the dynamic cache policy controller 304 identifies how to condition the accesses of the second frame based on the access data recorded for those accesses in the first frame.
[0040] In step 706, the dynamic cache policy controller 304 applies the identified parameters to the memory accesses in the second frame. In some examples, the parameters indicate how to age cache lines in the event of a miss, as described elsewhere in this specification. In some examples, the parameters indicate the age at which to insert a new cache line into the cache 302 in the event of a miss. In some examples, the parameters indicate what age to set the cache line to in the event of a hit. Applying these parameters to the memory accesses is described elsewhere in this specification.
[0041] The elements in the figures are embodied as software, where appropriate, running on a processor, a fixed function processor, a programmable processor, or a combination thereof. The processor 102, the last level cache 110, the interconnect 112, the memory 104, the storage device 108, the various auxiliary devices 106, the APD 116 and their elements, and the dynamic cache policy controller 304 include at least some hardware circuitry and, in some embodiments, include software running on a processor within those components or another component.
[0042] It should be understood that many variations are possible based on the disclosure of this specification. Although features and elements have been described above in specific combinations, each feature or element can be used alone without using other features and elements, or in various combinations with or without other features and elements.
[0043] The provided method can be implemented on a general-purpose computer, a processor, or a processor core. Suitable processors include, by way of example, general-purpose processors, dedicated processors, conventional processors, digital signal processors (DSPs), multiple microprocessors, one or more microprocessors associated with a DSP core, controllers, microcontrollers, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) circuits, any other type of integrated circuit (IC), and / or state machines. Such processors can be manufactured by configuring a manufacturing process using the results of processed hardware description language (HDL) instructions and other intermediate data (instructions that can be stored on a computer-readable medium), including netlists. The result of such processing may be a mask work, which can then be used in a subsequent semiconductor manufacturing process to manufacture a processor that implements the features of this disclosure.
[0044] The methods or flowcharts provided herein may be implemented in a computer program, software, or firmware incorporated into a non-transitory computer-readable storage medium for execution by a general-purpose computer or processor. Examples of non-transitory computer-readable storage media include magnetic media such as read only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, internal hard disks, and removable disks, magneto-optical media, and optical media such as CD-ROM disks and digital versatile disks (DVDs).
Claims
1. A method for operating a cache, comprising: recording access data for a set of first memory accesses of a first frame; identifying parameters for a set of second memory accesses of a second frame subsequent to the first frame based on the access data; applying the parameters to the set of second memory accesses; The method.
2. The method of claim 1, wherein the access data includes memory accesses characterized by one or both of a reuse interval and a reuse intensity. The method of claim 1.
3. The method of claim 2, wherein the reuse interval indicates the time between accesses, and the reuse intensity indicates the number of accesses to the same address within a time window. The method of claim 2.
4. The method of claim 1, wherein the parameters include one or more of an indication of how to age a cache line in case of a miss, an indication of at what age to insert a cache line in case of a miss, and an indication of how to change the age of a cache line in case of a hit. The method of claim 1.
5. The method of claim 1, wherein the set of second memory accesses is considered to be the same as the set of first memory accesses based on the order of the first memory accesses within the first frame and the order of the second memory accesses within the second frame. The method of claim 1.
6. The method of claim 1, wherein applying the parameters to the set of second memory accesses includes causing the set of second memory accesses as specified by the parameters. The method of claim 1.
7. The method of claim 1, wherein the set of first memory accesses and the set of second memory accesses include accesses to a cache. The method of claim 1.
8. The method of claim 1, wherein the first frame and the second frame are frames of graphics rendering. The method of claim 1.
9. The method of claim 1, wherein each memory access of the set of first memory accesses is to the same memory page. The method of claim 1.
10. A system, comprising: a cache; a dynamic cache policy controller; wherein the dynamic cache policy controller is configured to: record access data for a set of first memory accesses of a first frame; Based on the access data, identifying parameters for a set of second memory accesses in a second frame subsequent to the first frame; Applying the parameters to the set of second memory accesses; Is configured to perform, A system.
11. The access data includes memory accesses characterized by one or both of a reuse interval and a reuse intensity, The system of claim 10.
12. The reuse interval indicates the time between accesses, and the reuse intensity indicates the number of accesses to the same address within a time window, The system of claim 10.
13. The parameters include one or more of an indicator of how to age a cache line in case of a miss, an indicator of at what age to insert a cache line in case of a miss, and an indicator of how to change the age of a cache line in case of a hit, The system of claim 10.
14. The set of second memory accesses is considered to be the same as the set of first memory accesses based on the order of the first memory accesses in the first frame and the order of the second memory accesses in the second frame, The system of claim 10.
15. Applying the parameters to the set of second memory accesses includes generating the set of second memory accesses as specified by the parameters, The system of claim 10.
16. The set of first memory accesses and the set of second memory accesses include accesses to a cache, The system of claim 10.
17. The first frame and the second frame are frames of graphics rendering, The system of claim 10.
18. Each memory access in the set of first memory accesses is to the same memory page, The system of claim 10.
19. Comprising a graphics processing pipeline configured to render the first frame and the second frame, The system of claim 10.
20. A computer-readable storage medium storing instructions, The instructions, when executed by a processor, Record access data for a set of first memory accesses of a first frame; Identifying parameters for a set of second memory accesses of a second frame subsequent to the first frame based on the access data; Applying the parameters to the set of second memory accesses; Causing the processor to perform operations including; A computer-readable storage medium.