Cache policy informing criticality

By introducing critical values into the cache and modifying the LRU policy, and preserving critical cache rows is solved, the problem that the LRU policy cannot reflect the critical differences in cache rows is solved, and processor performance and memory utilization efficiency are improved.

CN120104520APending Publication Date: 2025-06-06APPLE INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510258695.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-04-22
Filing Date
2022-07-28
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Existing LRU replacement policies cannot effectively reflect critical differences in cache lines, resulting in frequent cache lines ejection affecting processor performance, especially when multiple operations rely on cache lines.

Method used

By introducing criticality values in cache lines, classifying cache lines according to criticality levels, and preserving critical cache lines in replacement policies, modifying the LRU policy to take into account the criticality of cache lines, dynamically adjusting their position and replacement order in cache.

Benefits of technology

Improve processor performance, reduce performance losses due to cache line ejection, especially for critical cache line access latency, and optimize memory latency and bandwidth utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104520A_ABST
    Figure CN120104520A_ABST
Patent Text Reader

Abstract

The invention relates to a cache policy for notifying criticality. A cache may store critical and non-critical cache lines, and may attempt to retain a critical cache line in the cache by, for example, supporting the critical cache line in replacement data updates, retaining the critical cache line at a particular probability when a sacrificial cache block is selected, etc. Critical values may be retained at various levels of the cache hierarchy. Additionally, an accelerated eviction may be employed if a thread that previously accesses a critical cache block is considered to be disappeared.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of invention patent application No. 202280058339.X, filed on July 28, 2022, and named “Cache Strategy for Notification Criticality”. Technical Field

[0002] Embodiments described herein relate to caches in computer systems, and more particularly to cache strategies. Background Art

[0003] Caches have long been employed in digital systems to reduce effective memory latency by capturing copies of data that has been accessed by a processor, coprocessor, or other digital device in a cache local to the device. A cache can be smaller than a main memory system and can be optimized for low latency (whereas a main memory system is typically optimized for storage density at some expense of latency). Thus, cache memory itself can reduce latency. Additionally, cache memory can be local to the device and thus can reduce latency because it does not cause transmission delays to the memory controller / main memory system and back to the device. Furthermore, a cache can be dedicated to a device or a small number of devices (e.g., a processor / coprocessor cluster) and thus can reduce contention for bandwidth to the cache compared to main memory.

[0004] Although caches reduce effective memory latency, they are finite storage devices and are therefore subject to misses (in addition to providing data to the requesting device in the case of a miss for a read request or an update in the case of a miss for a write request, which also causes a fill from memory to the cache to obtain the data). The fill is allocated to a storage device (e.g., a cache line or cache block) in the cache. This allocation can cause other data to be replaced in the cache (also known as evicting a cache line from the cache). There are multiple replacement strategies for selecting cache lines to be evicted based on cache geometry. For example, a set associative cache has a memory arranged as a two-dimensional array of cache lines: a "row" (called a group) is selected based on a subset of the memory address of the cache line, and the row includes multiple cache lines as "columns" (called ways) of the array. When a cache miss is detected and a fill is initiated, one of the ways is allocated for the fill. A popular replacement strategy for set associative caches is the least recently used (LRU) strategy. With LRU, accesses to cache lines in a group are tracked from the most recently accessed (most recently used or MRU) to the least recently accessed (least recently used or LRU). Typically, when a cache line is accessed, it is updated to the MRU, and cache lines between the previous order of cache lines and the previous MRU are adjusted. When a cache miss occurs, the LRU cache line may be selected for replacement. BRIEF DESCRIPTION OF THE DRAWINGS

[0005] The following detailed description refers to the accompanying drawings, which are now briefly described.

[0006] Figure 1 is a block diagram of one embodiment of a portion of a system.

[0007] Figure 2 is a flow chart illustrating criticality determination for one embodiment.

[0008] Figure 3 It is shown Figure 1 A table of LRU insertions and updates in one embodiment of a last level cache (LLC) is shown.

[0009] Figure 4 is a flow chart illustrating one embodiment of victim selection in LLC.

[0010] Figure 5 is a flow chart illustrating criticality determination for another embodiment.

[0011] Figure 6 is a flow chart illustrating LRU insertion in LLC for one embodiment.

[0012] Figure 7 is a flow chart illustrating LRU promotion in LLC for one embodiment.

[0013] Figure 8 is a flow chart illustrating sacrifice selection for one embodiment.

[0014] Fig. 9 is a flow chart illustrating sacrifice selection for another embodiment.

[0015] Fig.10 is a flow chart illustrating eviction acceleration for cache lines marked as critical.

[0016] Fig.11 is a flow diagram illustrating LRU insertion for a cache line at a memory cache for one embodiment.

[0017] Fig.12 is a block diagram of one embodiment of a system on a chip (SOC).

[0018] Fig.13 is a block diagram of various embodiments of a computer system.

[0019] Fig.14 is a block diagram of one embodiment of a computer-accessible storage medium.

[0020] Although the embodiments described in the present disclosure may be subject to various modifications and alternatives, specific embodiments thereof are shown by way of example in the drawings and will be described in detail herein. However, it should be understood that the drawings and specific embodiments thereof are not intended to limit the embodiments to the particular forms disclosed, but on the contrary, the present invention is intended to cover all modifications, equivalents and alternatives falling within the spirit and scope of the appended claims. The headings used herein are for organizational purposes only and are not intended to limit the scope of the specification. DETAILED DESCRIPTION

[0021] Although the LRU replacement strategy generally provides good performance (e.g., cache hit rates remain high and memory latency is effectively reduced), there are situations where performance may be limited. For example, when competition for cache lines is high and evictions occur frequently, some cache lines may be evicted, which cause a higher loss in performance of the requesting device than other cache lines when accessed again. For example, if multiple operations in the requesting device depend directly or indirectly through other operations on data in the cache line, the requesting device may stop waiting for the data. Other cache lines with lesser dependencies may be less critical to performance. The LRU strategy cannot reflect differences in the criticality of cache lines.

[0022] In one embodiment, a system comprising one or more processors and a cache memory coupled to the one or more processors can classify cache memory according to one or more criticality levels based on one or more criteria measured at the time when the cache memory is filled into the cache memory. Criteria can be selected to attempt to identify such cache memory, which has a greater impact on the performance of the processor than other cache memory when it is a miss in the cache memory. Each cache memory can have a criticality value specifying its criticality level. For example, a critical value can indicate a non-critical state or a critical state. In one embodiment, a critical state also can have a plurality of criticality levels, as described in detail below. In another embodiment, relative to a non-critical state, a critical state can be a critical single level.

[0023] High-speed cache can realize the replacement strategy that uses the criticality value of cache line as factor.For example, LRU strategy can be used, but this strategy can be modified to consider the criticality of various cache lines.The cache line (" critical cache line ") with the criticality value indicating critical state can be inserted into LRU replacement data at MRU position, and the cache line (" non-critical cache line ") with the criticality value indicating non-critical state can be inserted at the lower position (for example, closer to LRU position) in the data.In one embodiment, the criticality value also can affect the updating of LRU replacement data.Although LRU is used as an exemplary replacement strategy, other embodiments can realize other replacement strategies.For example, multiple pseudo-LRU strategies can be used, which approximate LRU operation by simplifying to make this strategy easier to realize, especially in wide set associative cache.Also can use random replacement strategy, and criticality can be used to reduce the possibility that critical line is selected.The least frequently used strategy can be used, and critical line can be selectively retained in a similar manner to the manner described below for LRU. A last-in-first-out or first-in-first-out strategy may be used, and critical cache lines may be at least partially exempted from LIFO or FIFO replacement.Any of these strategies may be modified to take criticality into account.

[0024] In one embodiment, the system may include one or more additional cache levels between the above-mentioned cache and the system memory. For example, a memory cache implemented at a memory controller of the control system memory may be used. When a cache line is evicted and revisited, the criticality value of the cache line may be exchanged between caches, thereby retaining the criticality value, while the cache line remains cached in the cache hierarchy. Once the cache line is removed from the cache hierarchy (and therefore the data only exists in the system memory), the criticality value may be lost.

[0025] Figure 1 1 is a block diagram of one embodiment of a system including a plurality of processors 10A to 10N, a coprocessor 12, a last level cache (LLC) 14, a memory controller 16, and a memory 18. The processors 10A to 10N and the coprocessor 12 are coupled to the LLC 14, which is coupled to the memory controller 16, which is further coupled to the memory 18. The processor 10N is shown in more detail, and other processors (such as the processor 10A) may be similar. The processor 10N may include an instruction cache (ICache) 20, an instruction cache (IC) miss queue 22, an execution core 24 including a load queue (LDQ) 26, a data cache (DCache) 28, and a memory management unit (MMU) 30. The LLC 14 may include a cache 32, a criticality control circuit 34, and a memory cache (MCache) insertion lookup table (LUT) 36. The memory cache 16 may include an insertion control circuit and LUT 38, an MCache 40, and a monitoring circuit 42.

[0026] ICache 20 may store instructions fetched by processor 10N for execution by execution core 24. If a fetch misses in ICache 20, the fetch for the cache line of the instruction may be queued in IC miss queue 22 and transmitted to LLC 14 as a fill request for ICache 20. Instructions executed by execution core 24 may include load instructions (more briefly, loads). Loads may attempt to read data from DCache 28 and may be transmitted to LLC 14 as a fill request for DCache 28 if the load misses in DCache 28. Loads transmitted to LLC 14 may be held in LDQ 26 waiting for data.

[0027] The MMU 30 may provide address translation for instruction fetch addresses and load / store addresses, including a translation lookaside buffer (TLB) that may be local to the ICache 20 and the execution core 24. The MMU 30 may optionally include one or more level 2 (L2) TLBs, and table walk circuitry for performing translation table reads to obtain translations for addresses that do not hit in the TLBs. The MMU 30 may transmit table walk reads to the LLC 14. In one embodiment, the MMU 30 may access the DCache 28 for potential cache hits on table walk reads before transmitting to the LLC 14, and if the read hits in the DCache 28, the read may not be transmitted to the LLC 14. In other embodiments, page table data is not cached in the DCache 28, and the MMU 30 may transmit table walk reads to the LLC 14.

[0028] LLC 14 includes cache 32, which may have any capacity and configuration. Memory requests from processors 10A to 10N and coprocessor 12 may be checked for hits in cache 32, and data may be returned as fills to ICache 20, DCache 28, or MMU 30 if a hit occurs. If a memory request is a miss in cache 32, LLC 14 may transmit the memory request to memory controller 16, and may return fills to the requesting processor 10A to 10N or coprocessor 12 in response to memory controller 16 returning fills to LLC 14. If a miss occurs, LLC 14 may also fill data into cache 32. In general, "data" is used herein in a general sense to refer to both instructions fetched by processors 10A to 10N for execution and data (e.g., operand data and result data) read / written by the processor as a result of the execution of instructions, particularly when referring to cache lines of data.

[0029] In addition, at the time of filling the processor 10A-10N / coprocessor 12, LLC 14 may assign a criticality value to the cache line. Criticality control circuitry 34 may determine the criticality value and may update cache 32 with the criticality value. For example, a cache tag in cache 32 may include a field for a criticality value. The criticality value may indicate a non-critical state or a critical state. As described above, in some embodiments, there may be more than one criticality level. Criticality control circuitry 34 may also determine the criticality level.

[0030] Criticality control circuit 34 can consider multiple factors when assigning criticality value to cache line. For example, criticality control circuit 34 is coupled to MMU 30, IC miss queue 22 and LDQ 26. More particularly, filling for table walk request can be classified as critical. TLB miss may affect additional instruction acquisition or load / store request, because translation covers a considerable amount of data and code sequence tends to access data close to other recently accessed data. For example, a page can be 4 kilobytes in size, 16 kilobytes in size or even larger, such as 1 megabyte or 2 megabytes. Any page size can be used. In addition, if the load is at the head of LDQ 26 when filling for the load occurs, it may be the oldest unfinished load in processor 10N. Therefore, it is likely that the load stops the retirement of other completed instructions, or there are multiple instructions that are stopped due to dependence (direct or indirect) on the load data. Filling for the load at the head of LDQ 26 can be assigned a critical state. Similarly, if the fill is for an instruction fetch request and it is the oldest fetch request in the IC miss queue 22 (e.g., it is at the head of the IC miss queue 22), then instruction fetch is likely to stall waiting for instructions. Such instruction fetches may be assigned a critical state. Other embodiments may include additional factors within a given processor 10A to 10N or a subset of the above factors and other factors as needed. In one embodiment, a request from a coprocessor 12 may also be assigned a critical state. For example, an embodiment of the coprocessor 12 may not include a cache, and therefore LLC 14 is the first cache level available to the coprocessor 12. Cache lines that are not assigned a critical state may be assigned a non-critical state.

[0031] In one embodiment, the criticality value assigned to the cache line can be maintained while the cache line remains valid in the cache hierarchy. The criticality value is assigned by the criticality control circuit 34, and then when the cache line is evicted from the cache 32, it is propagated with the cache line and transmitted to the memory controller 16, wherein the criticality value can be cached in the MCache 40. If the cache line evicted after being evicted from the cache 32 is placed in the MCache 40, the criticality value can be maintained. If the cache line evicted after being evicted from the cache 32 is not placed in the MCache 40, the memory controller 16 can discard the criticality value and write data to the memory 18. There may be multiple factors that affect whether the cache line evicted is cached in the MCache 40. The MCache 40 is shared with other components of the system, and the MCache 40 may have a quota for how much data can be cached from a given component. If LLC 14 exceeds the quota, the cache line evicted may not be cached. Alternatively, the evicted cache line may be cached, and a different LLC cache line cached in MCache 40 may be evicted.

[0032] Subsequently, if a cache line previously cached by LLC 14 is re-accessed by LLC 14, MCache 40 may provide the cache line as a fill to cache 32, and may also provide the criticality value previously associated with the cache line. Criticality control circuit 34 may assign the previous criticality value provided by MCache 40 to the cache line, unless other factors from processors 10A to 10N that generate a re-access to the cache line indicate an upgrade to a critical state or to a higher criticality level. For example, a non-critical cache line from MCache 40 may be filled into LLC 14 in a non-critical state, unless it is assigned a critical state at the time of the fill for re-access (e.g., the fill is for a load at the head of LDQ 26, an instruction fetch at the head of IC miss queue 22, or an MMU table walk request). A critical cache line from MCache 40 may be filled as critical. In embodiments implementing multiple criticality levels, a critical cache line from MCache 40 that is also currently indicated as critical via the above factors (head of LDQ 26, head of IC miss queue 22, or MMU request) may be assigned a higher criticality level by criticality control circuitry 34.

[0033] In one embodiment, the evicted cache line from LLC 14 may be cached in MCache 40 and may be inserted into the replacement data of the affected group of MCache 40 at a selected location. If the evicted cache line is a critical cache line, it may be inserted at the MRU position. If the evicted cache line is a non-critical cache line, it may be inserted at a position lower than the MRU (closer to the LRU). In one embodiment, the insertion point may be dynamic for non-critical cache lines. For example, the insertion point may be based on the amount of cache capacity occupied by the cache line from LLC 14 in MCache 40. The memory controller 16 may include a monitoring circuit 42 that monitors the capacity of the MCache 40 allocated to the CPU and provides the information ("capacity_CPU") to the criticality control circuit 34. The criticality control circuit 34 may use the capacity_CPU value as an index into the MCache insertion LUT 36 and may read the insertion hint from the indexed entry when transmitting the evicted cache line to the memory controller 16. The insertion hint may be used as an index into the LUT 38 in the memory controller 16, and the associated insertion control logic may potentially adjust the insertion point (e.g., if a portion of the cache loses power, the insertion point should be within the LRU location currently in use). The MCache 40 may insert the evicted cache block at the insertion point.

[0034] Therefore, in this embodiment, for non-critical cache lines, a cooperative lookup table may be used to determine the insertion point for the evicted cache line in MCache 40. The LUT may be programmable, allowing software to tune performance as needed.

[0035] The Capacity_CPU value may be measured in any desired manner. In one embodiment, Capacity_CPU may indicate the average number of MCache ways occupied by cache lines from LLC 14. In another embodiment, an approximate percentage of cache capacity may be provided.

[0036] As previously mentioned, cache 32 may have a field for a criticality value (e.g., in the cache tags). MCache 40 may similarly include a field in the cache tags for a criticality value. In another embodiment, MCache 40 may have a data set identifier (DSID) for each cache line that identifies cache lines that belong together based on one or more criteria. Typically, cache blocks with the same DSID may come from the same source component (e.g., LLC 14 or another component of the system, such as a peripheral device component, Figure 114). The DSID may be stored in a field in the tag. The DSID may be used to distinguish between non-critical and critical cache lines (e.g., by using one DSID for non-critical cache lines and another DSID for critical cache lines, or multiple DSIDs for different criticality levels in an embodiment employing more criticality levels). When providing fills, MCache 40 may decode the DSID to determine the criticality value to be transmitted to LLC 14.

[0037] In one embodiment, processors 10A to 10N may act as a central processing unit (CPU) of the system. The CPU of the system includes one or more processors that execute the main control software of the system (such as an operating system). Generally, the software executed by the CPU during use can control other components of the system to achieve the desired functions of the system. Processors 10A to 10N may also execute other software, such as application programs. Application programs may provide user functions and may rely on the operating system for lower-level device control, scheduling, memory management, etc. Therefore, processors 10A to 10N may also be referred to as application processors.

[0038] In general, a processor may include any circuitry and / or microcode configured to execute instructions defined in an instruction set architecture implemented by the processor. The processor may encompass a processor core implemented as a system on a chip (SOC) or with other levels of integration on an integrated circuit with other components. The processor may also include a discrete microprocessor, a processor core and / or a microprocessor integrated into a multi-chip module implementation, a processor implemented as multiple integrated circuits, and the like.

[0039] In one embodiment, coprocessor 12 can be configured to accelerate certain operations. For example, an embodiment in which the coprocessor performs matrix and vector manipulation (multiple operations per instruction) on a large scale is envisioned. Coprocessor 12 can receive instructions transmitted by processors 10A to 10N. That is, the instructions executed by coprocessor 12 ("coprocessor instructions") and the instructions executed by processors 10A to 10N ("processor instructions") can be part of the same instruction set architecture and can be intertwined in the code sequence obtained by the processor. Processors 10A to 10N can decode instructions and identify coprocessor instructions for transmission to coprocessor 12, and can execute processor instructions. Coprocessor 12 can receive coprocessor instructions from processors 10A to 10N, decode coprocessor instructions, and execute coprocessor instructions. Coprocessor instructions can include load / store instructions for reading memory data for operands and writing result data to memory (in one embodiment, both of which can be completed in LLC 14).

[0040] Please note that Figure 1The number and type of various components in the system may vary from implementation to implementation. For example, there may be any number of processors 10A to 10N. There may be more than one coprocessor 12, and when multiple coprocessors are included, there may be multiple instances of the same coprocessor and / or different types of coprocessors. There may be more than one memory controller 16, and when multiple memory controllers are included, the memory space may be distributed across the memory controllers.

[0041] Note that various instructions, memory requests, etc. are referred to above as being younger or older than other instructions, requests, etc. A given operation may be younger than another operation if it is derived from an instruction that is subsequent in program order to the instruction from which the other operation is derived. Similarly, a given operation may be older than another operation if it is derived from an instruction that is prior in program order to the instruction from which the other operation is derived.

[0042] Figures 2 to 4 Embodiments are shown in which the criticality value is either critical or non-critical. Figures 5 to 9 Embodiments are shown in which a critical state has more than one level of criticality. Fig.10 A mechanism for accelerating the removal of critical cache lines from LLC 14 is shown, which, for one embodiment, is applicable to two types of criticality values. Fig.11 It is shown based on Fig.10 Flowchart of the acceleration mechanism for performing sacrifice selection from LLC 14.

[0043] Now turn to Figure 2 , a flow chart illustrating one embodiment of the criticality control circuit 34 assigning criticality values ​​to cache lines being filled into the LLC 14 is shown. However, for ease of understanding, the blocks are shown in a particular order, and other orders may be used. The blocks may be executed in parallel in the criticality control circuit 34 in combinatorial logic. The blocks, combinations of blocks, and / or the flow chart as a whole may be pipelined over multiple clock cycles. The criticality control circuit 34 may be configured to implement Figure 2 The operation shown.

[0044] If the fill is for a cache line for an MMU table walk request (decision block 50, "yes" branch), the criticality control circuitry 34 may assign a criticality status for a criticality value associated with the cache line (block 52). If the fill is for a cache line for a load operation at the head of the LDQ 26 (decision block 54, "yes" branch), the criticality control circuitry 34 may assign a criticality status for a criticality value associated with the cache line (block 52). If the fill is for a cache line for an instruction cache miss at the head of the IC miss queue 22 (decision block 56, "yes" branch), the criticality control circuitry 34 may assign a criticality status for a criticality value associated with the cache line (block 52). If the fill is for a cache line with a criticality status in the MCache (decision block 58, "yes" branch), the criticality control circuitry 34 may assign a criticality status for a criticality value associated with the cache line (block 52). If none of the above criteria apply (decision blocks 50, 54, 56, 58, and 60, "no" branch), the criticality control circuitry 34 may assign a non-critical status to the criticality value associated with the cache line. In one embodiment, a coprocessor request from coprocessor 12 may also be assigned a critical status. In another embodiment, a coprocessor request may be assigned a non-critical status.

[0045] Figure 3 62 is a table showing the operation of one embodiment of the criticality control circuit 34 for updating replacement data for a group based on the filling of cache lines into the cache 32 (insert section 64) and based on cache hits (update section 66) for processor requests from processors 10A to 10N. The replacement data update may be based on the request type, the previous state of the cache block, and the criticality value. The LRU column of the table indicates the position of the cache line filled (in the insert section 64) or the cache line hit by the request (in the update section 66) in the LRU order (from MRU to LRU). Other cache lines in the group may be updated to reflect the change. For example, if the filled / hit cache line is made MRU, the position of each other cache line from the current MRU to the previous position of the filled / hit cache line may be moved one position toward the LRU. If a filled / hit cache line is moved to a different location in the replacement data than the MRU, then each cache line having a location from the different location to the current location of the filled / hit cache line may be moved one position toward the LRU.

[0046] In insert section 64, the previous state is empty because cache lines are being filled into cache 32. For this section, request types other than non-temporal (NT) demand requests update the replacement data to make the fill MRU for critical cache lines. If the fill is for a pre-fetch request (data or instruction) and the criticality value is non-critical, then make the fill LRU position N, which is close to the LRU position but not the LRU position itself. For example, N can be about 25% above the LRU at the distance between the LRU and the MRU. For example, if cache 32 is 8-way, then 25% above the LRU would be 2 positions above the LRU. If cache 32 is 16-way, then 25% above the LRU would be 4 positions above the LRU. If the fill is for demand acquisition (instruction or data) and the criticality value is non-critical, then make the fill LRU position L (close to the middle of the replacement data range). For example, if cache 32 is 8-way, then in various embodiments, assuming the LRU position is numbered 0, L may be in the range of positions 4 to 6. If the cache is 16-way, then L may be in the range of 6 to 8. If the fill is for an NT demand fetch, then the LRU position of the fill may be position M, close to LRU but less than N.

[0047] exist Figure 3 In an embodiment, the update of replacement data for a hit for a cache line may be independent of the criticality value of the cache line. Other embodiments may take into account criticality in the update. If the hit request is a demand get (instruction or data) and the cache line is a prefetched cache line, the LRU position may be unchanged (NC), but the prefetch tracking bit may be reset for the cache line so that the next time the cache line is hit, it will be a demand get. If the hit request is a demand get (instruction or data) and the hit cache line is NT demand or demand get (instruction or data), the hit cache line may be made MRU. If the hit request is a data prefetch, the hit cache line may be placed at N (close to LRU). If the hit request is an instruction prefetch, the hit cache line is made MRU. If the hit request is an NT demand, the hit cache may be in position N.

[0048] Now turn to Figure 4 , a flow chart illustrating the operation of one embodiment of the criticality control circuit 34 for selecting a victim cache line to be evicted when a cache miss is detected is shown. However, for ease of understanding, the blocks are shown in a particular order, and other orders may be used. The blocks may be executed in parallel in combinatorial logic in the criticality control circuit 34. The blocks, combinations of blocks, and / or the flow chart as a whole may be pipelined over multiple clock cycles. The criticality control circuit 34 may be configured to implement Figure 4 The operation shown.

[0049] If there is at least one invalid cache entry in the set that is indexed by a cache miss (decision block 70, "yes" branch), the criticality control circuitry 34 may select the LRU-least invalid entry (block 72). An invalid entry may be a cache line storage location (e.g., a way) that does not currently store a cache line. The LRU-least invalid entry may be an invalid entry that is invalid and has a location that is closest to the LRU location in the replacement data when compared to the locations of the other invalid entries. The LRU-least invalid entry may be at the LRU location.

[0050] If there is no invalid entry in the group (decision box 70, "no" branch), the criticality control circuit 34 can select a valid entry as a victim. In a typical LRU strategy, an LRU entry can be selected. However, in this embodiment, the criticality control circuit 34 can retain the critical cache line with a specific probability. Therefore, a biased pseudo-random selection (e.g., based on a linear feedback shift register or LFSR and a desired probability) can be generated (box 74). Based on the pseudo-random selection, the criticality control circuit 34 can selectively mask the critical cache line to avoid being selected (box 76). For example, if the biased pseudo-random selection indicates an evaluation of a biased test (e.g., "yes"), the critical cache line can be unmasked. If the biased pseudo-random value indicates another evaluation of a biased test (e.g., "no"), the critical cache line can be masked. This type of probability-based retention can also be referred to as "biased coin flipping". The criticality control circuit 34 can select the LRU-most effective de-masking entry, and the cache block in the entry can be driven out (box 78).

[0051] Figure 5 1 is a flow chart illustrating the operation of the criticality control circuit 34 for another embodiment to assign criticality to a cache line being populated into the LLC 14. However, the blocks are shown in a particular order for ease of understanding, and other orders may be used. The blocks may be executed in parallel in the criticality control circuit 34 in combinatorial logic. The blocks, combinations of blocks, and / or the flow chart as a whole may be pipelined over multiple clock cycles. The criticality control circuit 34 may be configured to implement Figure 5 The operation shown.

[0052] Similar to Figure 2In an embodiment of the present invention, a cache line may be critical if it is being filled as a result of an MMU table walk request (decision box 80, "yes" branch), a load at the head of LDQ 26 (decision box 82, "yes" branch), or an instruction fetch at the head of IC miss queue 22 (decision box 84, "yes" branch). In this embodiment, there are multiple criticality levels. If the criticality supplied by MCache 40 indicates a critical state (decision box 86, "yes" branch), the criticality control circuit 34 may increase the level of criticality from the state provided by MCache 40 (box 88). If MCache 40 indicates non-critical (decision box 86, "no" branch), the cache line was previously non-critical or the cache line was a miss in MCache 40. In these cases, the criticality control circuit 34 may initialize the criticality value at the lowest level of criticality (box 90).

[0053] If the cache line is not critical in the current population (decision blocks 80, 82, and 84, "no" branch), but the criticality value provided by MCache 40 is in a critical state (decision block 92, "yes" branch), then criticality control circuitry 34 may retain the criticality value provided by MCache 40 (block 94). Otherwise (decision block 92, "no" branch), criticality control circuitry 34 may initialize the criticality value to a non-critical state (block 96).

[0054] Figure 6 is a flow chart illustrating the operation of one embodiment of the criticality control circuit 34 for updating replacement data for a group based on a cache line fill (cache line insertion) into the cache 32. However, for ease of understanding, the blocks are shown in a particular order, and other orders may be used. The blocks may be executed in parallel in the criticality control circuit 34 in combinatorial logic. The blocks, combinations of blocks, and / or the flow chart as a whole may be pipelined over multiple clock cycles. The criticality control circuit 34 may be configured to implement Figure 6 The operation shown.

[0055] If the cache line being filled has a high criticality (e.g., in one embodiment, a criticality other than the lowest of the criticalities) (decision box 100, "yes" branch), the cache line may be inserted at the MRU position in the replacement data (box 102). If the cache line has a criticality (e.g., the lowest criticality) (decision box 100, "no" branch, and decision box 104, "yes" branch), the criticality control circuitry 34 may be configured to insert the cache line in the replacement data as high as possible (closest to the MRU) but below the position of any high criticality cache line. Therefore, if there are one or more high criticality cache lines in the replacement data (decision box 106, "yes" branch), the criticality control circuitry 34 may insert the cache line at the highest position below the high criticality cache line (box 108). Otherwise, the cache line may be inserted at the MRU position (decision box 106, "no" branch and box 102).

[0056] If the cache line is filled non-critically (decision blocks 100 and 104, "no" branch) and the fill is due to a prefetch (instruction or data) (decision block 110, "yes" branch), then similar to the above reference Figure 3 Discussion of inserting prefetch at N near LRU (block 112). In one embodiment, instruction prefetch may be placed at a lower LRU position than data prefetch, but both may be placed near LRU. Alternatively, instruction prefetch may be placed at a higher LRU position than data prefetch, but both are near LRU, or the same LRU position may be used for both types of prefetch. If the non-critical cache line is not a prefetch but an NT request (decision block 114, "yes" branch), the cache line may be inserted at position M, which is greater than N but near LRU in this embodiment (block 116). If a non-critical cache line demand request (decision block 114, "no" branch) and there are any critical cache lines (decision block 106, "yes" branch), the non-critical cache line may be inserted below the critical cache line (block 108). If there is no critical cache line in the group (decision block 106, "no" branch), the non-critical cache line may be inserted at the MRU position (block 102).

[0057] The circuitry represented by decision block 106 and blocks 102 and 108 may provide dynamic insertion points for particular cache lines, thereby preventing "priority inversion" in replacement data where critical cache lines may be moved down the replacement data toward the LRU position by less critical cache lines.

[0058] Figure 7is a flow chart illustrating the operation of one embodiment of the criticality control circuit 34 for updating replacement data for a group based on a cache line hit (cache line promotion) into the criticality control circuit 34. However, for ease of understanding, the blocks are shown in a particular order, and other orders may be used. The blocks may be executed in parallel in combinatorial logic in the criticality control circuit 34. The blocks, combinations of blocks, and / or the flow chart as a whole may be pipelined over multiple clock cycles. The criticality control circuit 34 may be configured to implement Figure 7 The operation shown.

[0059] If the hit cache line has any criticality level (decision block 120, "yes" branch), the criticality control circuitry 34 may update the cache line to the MRU position (block 122). If the hit cache line is non-critical (decision block 120, "no" branch), and the hit cache line is not touched by the prefetch request (decision block 124, "yes" branch), the criticality control circuitry 34 may leave the replacement data position unchanged but may reset the prefetch bit (block 126). If the hit request is a demand or data prefetch (decision block 128, "yes" branch), the criticality control circuitry 34 may preserve the priority of the critical cache line by promoting the hit cache line to the highest replacement data position below the critical cache line (decision block 130, "yes" branch and block 132). If there is no critical cache line in the group, the hit cache line may be made MRU (decision block 130, "no" branch and block 122). If the hit request is an NT request (decision block 134, "yes" branch), the hit cache line may be updated to position P close to the LRU, unless the hit cache line is an untouched prefetch, in which case the position is unchanged (block 136). If the hit request is not an NT request (nor the other types of requests described above), the request may be an instruction prefetch and the hit cache line may be updated to the MRU (block 138).

[0060] Similar to the above reference Figure 6 As discussed above, the circuitry represented by decision block 130 may provide dynamic replacement data updates to prevent priority inversion between non-critical cache lines and critical cache lines. Figure 6 Embodiments may allow different critical cache line levels to be reordered in the replacement data, but may keep non-critical cache lines below critical cache lines in the replacement data.

[0061] Now turn to Figure 8, a flow chart illustrating the operation of one embodiment of the criticality control circuit 34 for selecting a victim cache line to be evicted when a cache miss is detected is shown. However, for ease of understanding, the blocks are shown in a particular order, and other orders may be used. The blocks may be executed in parallel in combinatorial logic in the criticality control circuit 34. The blocks, combinations of blocks, and / or the flow chart as a whole may be pipelined over multiple clock cycles. The criticality control circuit 34 may be configured to implement Figure 8 The operation shown.

[0062] If there is at least one invalid entry in the group (decision block 140, "yes" branch), the criticality control circuit 34 may mask all valid entries and select the LRU-least masked (invalid) entry (block 142). If all entries are valid (decision block 140, "no" branch), the criticality control circuit 34 may determine a biased pseudo-random selection, similar to the above with reference to Figure 4 144). Based on the pseudo-random selection, the criticality control circuitry 34 may selectively mask all critical cache lines (block 146). If at least one unmasked valid entry is found (decision block 148, "yes" branch), the criticality control circuitry 34 may select the LRU-least unmasked entry (block 142). If no entry is found (decision block 148, "no" branch), the criticality control circuitry 34 may horizontally unmask the lowest critical cache line while still masking higher critical cache lines (block 150). If at least one unmasked valid entry is found (decision block 152, "yes" branch), the criticality control circuitry 34 may select the LRU-least unmasked entry (block 142). If no entry is found (decision block 152, "no" branch), the criticality control circuitry 34 may horizontally unmask all critical cache lines (block 154) and may select the LRU-least unmasked entry (block 142).

[0063] Fig. 9 is a flow chart illustrating the operation of another embodiment of the criticality control circuit 34 for selecting a victim cache line to be evicted when a cache miss is detected. However, the blocks are shown in a particular order for ease of understanding, and other orders may be used. The blocks may be executed in parallel in combinatorial logic in the criticality control circuit 34. The blocks, combinations of blocks, and / or the flow chart as a whole may be pipelined over multiple clock cycles. The criticality control circuit 34 may be configured to implement Fig. 9 The operation shown.

[0064] Fig. 9 Embodiments of may employ multiple biased pseudo-random selections based on different probabilities to selectively mask or unmask various subsets of critical states before selecting a victim. Figure 8In an embodiment, if there is at least one invalid entry in the group (decision block 160, "yes" branch), the criticality control circuit 34 may mask all valid entries and select the LRU-least masked (invalid) entry (block 162). If all entries are valid (decision block 160, "no" branch), the criticality control circuit 34 may determine a first biased pseudo-random selection based on a first probability, similar to the above reference to Figure 4 164). If the selection is yes (decision block 168, "yes" branch), the criticality control circuitry 34 may mask all critical cache lines (block 168) and determine whether at least one valid unmasked entry is found (decision block 170). If yes (decision block 170, "yes" branch), the criticality control circuitry 34 may select the LRU-least unmasked entry (block 162). If no (decision block 170, "no" branch) or if the selection is no (decision block 166, "no" branch), the criticality control circuitry 34 may determine a second biased pseudo-random selection based on a second probability (block 172). If the selection is yes (decision block 174, "yes" branch), the criticality control circuitry 34 may mask the critical cache lines except for the lowest critical state (block 176) and determine whether at least one valid unmasked entry is found (decision block 178). If yes (decision block 178, "yes" branch), the criticality control circuitry 34 may select the LRU-most unmasked entry (block 162). If no (decision block 178, "no" branch) or if the selection is no (decision block 174, "no" branch), the criticality control circuitry 34 may continue similar iterations until an entry is found (block 180) or until all critical rows are not masked, thereby masking fewer of the highest criticality levels. Once an entry is found, the criticality control circuitry may select the LRU-most unmasked entry (block 162).

[0065] Embodiments that implement dynamic replacement data updates to preferentially retain critical cache lines that are closer to the MRU than other cache lines can successfully retain cache lines in LLC 14. However, once a critical cache line is no longer available, the same characteristics can increase the difficulty of replacing the critical cache line with a recently accessed cache line that is not critical.

[0066] As described above, during the selection of the victim cache line for replacement, LLC 14 may be configured to preferentially retain the cache line identified as critical by the corresponding criticality value rather than the cache line not identified as critical. LLC 14 may be configured to select the victim cache line according to the replacement data separated from the criticality value maintained by the cache (and also taking into account the criticality value). However, when the criticality control circuit 34 detects one or more indications that at least some of the cache lines identified as critical are no longer critical, the criticality control circuit 34 may be configured to terminate the priority retention of the cache line based on the one or more indications. In another way of looking at it, the criticality control circuit 34 may accelerate the eviction of the cache line identified as critical based on the one or more indications (compared to the reservation that will be applied before the one or more indications are detected). For example, in one embodiment, the criticality control circuit 34 may be configured to ignore the criticality value when selecting the victim cache line to terminate the priority retention of the critical cache line or accelerate the eviction of the critical cache line.

[0067] Fig.10 is a flow chart illustrating the operation of one embodiment of the criticality control circuit 34 for accelerating the eviction of critical cache lines that are no longer in use. However, the blocks are shown in a particular order for ease of understanding, and other orders may be used. The blocks may be executed in parallel in the criticality control circuit 34 in combinatorial logic. The blocks, combinations of blocks, and / or the flow chart as a whole may be pipelined over multiple clock cycles. The criticality control circuit 34 may be configured to implement Fig.10 The operation shown.

[0068] If the critical cache line is no longer being accessed (e.g., one or more accessing threads have completed execution), the critical cache line may eventually migrate toward an LRU position in the replacement data. Therefore, the criticality control circuit 34 may monitor the hit rate for the critical cache lines in the N most LRU positions (block 190). N may be selected in any desired manner. For example, N may be approximately one-quarter the number of ways in the group. Additionally, if snooped back copies of cache lines from LLC 14 are increasing (i.e., snoops are causing cache lines to be forwarded to another processor 10A to 10N), the thread that is accessing the critical cache line may have migrated to a different cluster of processors 10A to 10N coupled to a different LLC 14 in the system ( Figure 116). Thus, the criticality control circuitry 34 may monitor the snoop rate that causes cache lines to be forwarded to other agents in the system (not back to the memory controller 16) (block 192). In various embodiments, the criticality control circuitry 34 may monitor snoop forwarding for only critical cache blocks or all cache blocks. Another factor that may be monitored is inferred coprocessor requests (requests from the coprocessor 12) (block 194).

[0069] If the cache hit rate detected via monitoring represented by block 190 is less than a threshold value (decision block 196, "yes" branch), the criticality control circuitry 34 may ignore criticality values ​​in victim selection and LRU insertion and promotion (block 198). Thus, cache lines may be treated the same regardless of critical / non-critical status. Similarly, if the snoop forwarding rate appears above a threshold value (decision block 200, "yes" branch), the criticality control circuitry 34 may ignore criticality values ​​in victim selection and LRU insertion and promotion (block 198). If the inferred coprocessor requests are increasing (decision block 202, "yes" branch), the criticality control circuitry 34 may ignore criticality values ​​in victim selection and LRU insertion and promotion (block 198).

[0070] Another factor that may be used is that if the capacity available in MCache 40 for processors 10A to 10N / LLC 14 drops below a threshold (e.g., as indicated by the Capacity_CPU indication from monitoring circuit 42) (decision block 204, "yes" branch), then the criticality control circuit 34 treats all criticality levels as the lowest criticality (block 206). If none of the above is true (decision blocks 196, 200, 202, and 204, "no" branch), then the criticality control circuit 34 may maintain the use of criticality values ​​in victim selection and LRU insertion and promotion (block 208).

[0071] Therefore, in this embodiment, the one or more indications may include a cache hit rate below a threshold level for a cache line in a plurality of least recently used locations in the replacement data and having a criticality value indicating a critical state. The one or more indications may include a ratio of snoop hits occurring in the cache and causing the forwarding of the corresponding cache line in response to the snoop hit being above the threshold level. In a system including a coprocessor, the coprocessor is coupled to the cache and is configured to execute a coprocessor instruction issued by the one or more processors to the coprocessor, the one or more indications may include a memory request issued by the coprocessor to the cache. The criticality control circuit 34 may be configured to infer a coprocessor memory request based on a prefetch request indicated as a coprocessor prefetch request generated by the one or more processors. As also described above, the MCache 40 may provide an indication of the capacity of the second cache that can be allocated to data from the LLC 14, and the control circuit is configured to cover the multiple criticality levels with the lowest criticality level in the multiple criticality levels based on an indication that the capacity is less than the threshold.

[0072] In one embodiment, a method may include: assigning criticality values ​​to cache lines in a cache, wherein a given criticality value corresponds to a given cache line; during selection of a victim cache line for replacement, preferentially retaining cache lines identified as critical by corresponding criticality values ​​rather than cache lines not identified as critical, wherein the selection is also based on replacement data maintained by the cache separate from the criticality value; detecting one or more indications that at least some of the cache lines identified as critical are no longer critical; and ignoring criticality values ​​updated for victim selection and replacement data based on the one or more indications. For example, in one embodiment, the method also includes: monitoring a cache hit rate for cache lines in multiple least recently used locations in the replacement data and having a criticality value indicating a critical state, and one of the one or more indications is based on the cache hit rate being below a threshold level. In one embodiment, the method also includes: monitoring a rate at which snoop hits occur in the cache and causing a corresponding cache line to be forwarded in response to the snoop hits, and one of the one or more indications is based on the snoop hit rate being above a threshold level. In one embodiment, the one or more indications include a memory request issued by a coprocessor to a cache, wherein the coprocessor is coupled to the cache and is configured to execute coprocessor instructions issued by one or more processors to the coprocessor. The method may also include: inferring the coprocessor memory request based on a prefetch request generated by the one or more processors and indicated as a coprocessor prefetch request. In one embodiment, the method may also include: providing an indication of capacity in the second cache that is allocable to data from the cache from the second cache, wherein the criticality value indicates non-critical and a plurality of criticality levels; and overwriting the plurality of criticality levels with a lowest criticality level of the plurality of criticality levels based on an indication that the capacity is less than a threshold.

[0073] Fig.11 1 is a flow chart illustrating the operation of one embodiment of the criticality control circuit 34 and MCache 40 for inserting evicted cache lines from LLC 14 into MCache 40 replacement data. However, for ease of understanding, the blocks are shown in a particular order, and other orders may be used. The blocks may be executed in parallel in the criticality control circuit 34 and / or MCache 40 in combinational logic. The blocks, combinations of blocks, and / or the flow chart as a whole may be pipelined over multiple clock cycles. The criticality control circuit 34 / MCache 40 may be configured to implement Fig.11 The operation shown.

[0074] If the evicted cache line is a critical cache line (decision box 210, "yes" branch), the criticality control circuit 34 may generate an insert hint for MCache 40 to insert the cache line at the MRU position (box 212). Alternatively, MCache 40 may detect the criticality of the cache line and insert the cache line at the MRU position. If the cache line is non-critical (decision box 210, "no" branch), the criticality control circuit 34 may generate an index to the MCache insert LUT 36 based on the capacity_CPU telemetry data (box 214). For example, the capacity_CPU telemetry data may indicate the average number of ways of MCache 40 that can be used for cache lines from processors 10A to 10N / LLC 14. The index may be generated based on the average number of ways in various ranges. For example, up to one-eighth of the number of ways, one-eighth to one-quarter of the number of ways, one-quarter to one-half of the number of ways, and more than one-half of the number of ways may be an index for a two-bit insert hint. Criticality control circuitry 34 may generate insertion hints from the indexed entries in MCache insertion LUT 36 (block 216).

[0075] When the insertion control circuit and LUT 38 receive the evicted cache block, the insertion hint can be used as an index to LUT 38, and the insertion position can be read from the table (box 218). The insertion control circuit 38 can modify the insertion position based on whether the MCache way is powered off for power conservation. That is, each powered off way occupies an LRU position in the replacement data because it cannot be used. If the insertion position will be in one of the N LRU positions, where N is the number of powered off ways, the insertion position can be increased to N (box 220). MCache 40 can allocate entries for cache lines and update the entries with cache lines (box 222), and if the cache line (if any) evicted from MCache 40 is modified relative to the copy in memory 18, the cache line is written to memory 18. MCache 40 can update the replacement data to indicate the allocated entry at the insertion position (box 224).

[0076] In one embodiment, MCache 40 may also support dynamic insertion locations for non-critical cache lines from processors 10A to 10N / LLC 14. For example, MCache 40 may determine the MRU-most non-critical cache line (excluding the cache line for which the insertion point is detected), referred to as position H in this paragraph. If there is no non-critical cache line, MCache 40 may insert the cache line at the adjusted insertion position described in the previous paragraph. However, if there is a valid non-critical cache line in MCache 40 and the cache line being inserted is already at a position that is more MRU than position H, MCache 40 may insert the cache line at a position that is closer to the MRU than position H. Otherwise, MCache 40 may insert the cache line at position H. MCache 40 may update replacement data to indicate the allocated entry at the insertion position (box 224).

[0077] Fig.12 308A to 308B) (more briefly, "peripherals"), memory controller 16, and communication structure 312. Components 304, 306, 308A to 308B, and 16 may all be coupled to communication structure 312. Memory controller 16 may be coupled to memory 18 during use. In some embodiments, there may be more than one memory controller coupled to corresponding memory. In such embodiments, memory address space may be mapped across memory controllers in any desired manner. In the illustrated embodiment, processor cluster 304 may include multiple processors (P) 10A to 10N. Processors 10A to 10N may form a central processing unit (CPU) of SOC 300. Processor cluster 304 may also include one or more coprocessors (e.g., Fig.12 The processor cluster 304 may also include the LLC 14. The processor cluster 306 may be similar to the processor cluster 304. Thus, the SOC 300 may be Figure 1 A specific implementation of the system shown.

[0078] The memory controller 16 may generally include circuits for receiving memory operations from other components of the SOC 300 and for accessing the memory 18 to complete the memory operations. The memory controller 16 may be configured to access any type of memory 18. For example, the memory 18 may be a static random access memory (SRAM), a dynamic RAM (DRAM) such as a synchronous DRAM (SDRAM) including a double data rate (DDR, DDR2, DDR3, DDR4, etc.) DRAM. Low power / mobile versions of DDR DRAM (e.g., LPDDR, mDDR, etc.) may be supported. The memory controller 16 may include a queue for memory operations to sort (and potentially reorder) these operations and present them to the memory 18. The memory controller 16 may also include a data buffer for storing write data waiting to be written to the memory and read data waiting to be returned to the source of the memory operation. In some embodiments, the memory controller 16 may include a memory cache (MCache) 40 for storing recently accessed memory data. For example, in a SOC implementation, MCache 40 can reduce power consumption in the SOC by avoiding re-accessing data from memory 18 when it is expected to be accessed again soon. In some cases, MCache 40 may also be referred to as a system cache, as opposed to a private cache that only serves certain components (such as a cache in LLC 14 or processors 10A to 10N). In addition, in some embodiments, the system cache need not be located within memory controller 16.

[0079] The peripherals 308A-308B may be any set of additional hardware functions included in the SOC 300. For example, the peripherals 308A-308B may include video peripherals, such as one or more graphics processing units (GPUs), image signal processors configured to process image capture data from a camera or other image sensor, video encoders / decoders, scalers, rotators, mixers, display controllers, etc. The peripherals may include audio peripherals, such as microphones, speakers, interfaces to microphones and speakers, audio processors, digital signal processors, mixers, etc. The peripherals may include interface controllers for various interfaces external to the SOC 300, including interfaces such as a universal serial bus (USB), a peripheral component interconnect (PCI) (including PCI Express (PCIe), a serial port, and a parallel port, etc.). The interconnections to the peripherals are provided by Fig.12 300 is illustrated by a dashed arrow extending outside of SOC 300. Peripherals may include networking peripherals such as a media access controller (MAC). Any set of hardware may be included.

[0080] The communication fabric 312 may be any communication interconnect and protocol for communicating between components of the SOC 300. The communication fabric 312 may be bus-based, including a shared bus configuration, a crossbar configuration, and a hierarchical bus with bridges. The communication fabric 312 may also be packet-based and may be a hierarchical, crossbar, point-to-point, or other interconnect with bridges.

[0081] Note that the number of components of SOC 300 (and Fig.12 Subcomponents of those components shown, such as the number of processors 10A to 10N in each processor cluster 304 and 306, may vary from implementation to implementation. Additionally, the number of processors 10A to 10N in one processor cluster 304 may be different than the number of processors 10A to 10N in another processor cluster 306. Fig.12 Quantities shown may be greater or less than that of each component / subcomponent.

[0082] Based on the foregoing, in one embodiment, a system may include: one or more processors configured to issue a memory request to access a memory system; and a cache coupled to the one or more processors and configured to cache data from the memory system for access by the one or more processors. The cache may include a control circuit configured to assign a criticality value to a cache line based on multiple factors at the time when the cache line is filled into the cache. During the filling of a given cache line, the control circuit may be configured to represent the given cache line at a selected position in the replacement data for the cache based on the criticality value assigned to the given cache line. The control circuit may be configured to select a victim cache line to be evicted from the cache based on the replacement data. The control circuit may be configured to selectively prevent a cache line having a criticality value indicating a critical state from being selected as a victim cache line based on probability.

[0083] In one embodiment, the system further includes a second cache coupled to the cache and configured to cache data from the memory system for the cache and for one or more other agents in the system that access the cache. The second cache is configured to store a victim cache line and retain an indication of a criticality value assigned to the victim cache line by the control circuit. In one embodiment, the system further includes a memory controller configured to control one or more memory devices that form at least a portion of the system memory, and the memory controller includes the second cache. In one embodiment, the second cache may be configured to provide a criticality value with the victim cache line in filling the cache based on another memory request that occurs after the victim cache line is evicted from the cache. In one embodiment, the second cache is configured to hold second replacement data; and the second cache may be configured to evict the cache line from the second cache based on the second replacement data. The initial position of the victim cache line in the second replacement data may be based on the criticality value. In one embodiment, the system includes a monitoring circuit coupled to the second cache and configured to provide an indication of the capacity in the second cache that can be allocated to data from the cache. The cache may be configured to generate an insertion hint for transmission with the victim cache line based on the indication of capacity. For example, the cache may include a table coupled to the control circuit that maps a range of indications of capacity to values ​​for the insertion hint. In one embodiment, the second cache includes a second table. The second cache may be configured to select an entry in the second table based on the insertion hint. The second table may be configured to output an insertion point indication from the selected entry.

[0084] In one embodiment, the criticality status may include critical and non-critical. In one embodiment, the criticality status may also indicate that one or more criticality levels are assigned to a critical cache line.

[0085] In one embodiment, the control circuitry may be configured to update the replacement data based on a request that hits a second given cache line in the cache. The replacement data may be updated based on a criticality value assigned to the second given cache line and criticality values ​​of other cache lines represented in the replacement data to move the location where the second entry of the second given cache line is stored to be closer to a recently accessed location. For example, where the criticality value assigned to the second given cache line is less than the criticality value of one or more other cache lines represented in the replacement data, the control circuitry may be configured to update the replacement data to represent the second given cache line at a second location below the location occupied by the one or more other cache lines.

[0086] In one embodiment, the control circuit is configured to monitor a cache hit rate for cache lines in a plurality of low positions in the replacement data and having a criticality value indicating a critical state. The control circuit may be configured to ignore criticality values ​​updated for victim selection and replacement data based on the cache hit rate being below a threshold level. In one embodiment, the control circuit may be configured to monitor a snoop hit rate for snoops that cause cache lines to be forwarded from the cache. The control circuit may be configured to ignore criticality values ​​updated for victim selection and replacement data based on the snoop hit rate exceeding a threshold level.

[0087] Computer Systems

[0088] Next turn Fig.13 , a block diagram of one embodiment of a system 700 is shown. In the illustrated embodiment, the system 700 includes at least one instance of a system on a chip (SOC) 706 coupled to one or more peripheral devices 704 and an external memory 702. A power supply (PMU) 708 is provided that supplies a supply voltage to the SOC 706 and one or more supply voltages to the memory 702 and / or the peripheral devices 704. In some embodiments, more than one instance of the SOC may be included (and more than one memory 702 may also be included). In one embodiment, the memory 702 may include Figure 1 and Fig.12 The memory 18 is shown. In one embodiment, the SOC 706 may be Fig.12 An example of a SOC 300 is shown.

[0089] Depending on the type of system 700, peripheral device 704 may include any desired circuit system. For example, in one embodiment, system 704 may be a mobile device (e.g., personal digital assistant (PDA), smart phone, etc.), and peripheral device 704 may include equipment for various types of wireless communications, such as Wi-Fi, Bluetooth, cellular, global positioning system, etc. Peripheral device 704 may also include additional storage devices, which include RAM storage devices, solid-state storage devices, or disk storage devices. Peripheral device 704 may include user interface devices, such as display screens, which include touch display screens or multi-touch display screens, keyboards or other input devices, microphones, speakers, etc. In other embodiments, system 700 may be any type of computing system (e.g., desktop personal computers, laptops, workstations, network set-top boxes, etc.).

[0090] The external memory 702 may include any type of memory. For example, the external memory 702 may be an SRAM, a dynamic RAM (DRAM) such as a synchronous DRAM (SDRAM), a double data rate (DDR, DDR2, DDR3, etc.) SDRAM, a RAMBUS DRAM, a low power version of a DDR DRAM (e.g., LPDDR, mDDR, etc.), etc. The external memory 702 may include one or more memory modules to which a memory device may be mounted, such as a single inline memory module (SIMM), a dual inline memory module (DIMM), etc. Alternatively, the external memory 702 may include one or more memory devices mounted on the SOC 706 in a chip-on-chip or package-on-package implementation.

[0091] As shown, system 700 is shown as having applications in a wide range of fields. For example, system 700 can be used as a part of a chip, circuit system, component, etc. of a desktop computer 710, a laptop computer 720, a tablet computer 730, a cellular or mobile phone 740, or a TV 750 (or a set-top box coupled to a TV). A smart watch and a health monitoring device 760 are also shown. In some embodiments, the smart watch may include various general computing related functions. For example, the smart watch may provide access to email, mobile phone services, user calendars, etc. In various embodiments, the health monitoring device may be a dedicated medical device or otherwise include dedicated health-related functions. For example, the health monitoring device may monitor the vital signs of the user, track the proximity of the user to other users for the purpose of epidemiological social distance, contact tracking, provide communications to emergency services in the event of a health crisis, etc. In various embodiments, the above-mentioned smart watch may include or may not include some or any health monitoring related functions. Other wearable devices are also envisioned, such as devices worn around the neck, devices that can be implanted in the human body, glasses designed to provide enhanced and / or virtual reality experience, and the like.

[0092] The system 700 may also be used as part of a cloud-based service 770. For example, the previously mentioned devices and / or other devices may access computing resources in the cloud (i.e., remotely located hardware and / or software resources). Still further, the system 700 may be used in one or more devices in a home other than those previously mentioned. For example, a home appliance may monitor and detect noteworthy situations. For example, various devices in a home (e.g., a refrigerator, a cooling system, etc.) may monitor the status of the device and provide an alert to the homeowner (or, for example, a maintenance agency) if a particular event is detected. Alternatively, a thermostat may monitor the temperature in a home and may automatically adjust the heating / cooling system based on a history of responses by the homeowner to various situations. Fig.13The system 700 is also illustrated in the application to various modes of transportation. For example, the system 700 can be used for control and / or entertainment systems of airplanes, trains, buses, taxis, private cars, watercraft ranging from private boats to cruise ships, scooters (for rental or private use), etc. In various cases, the system 700 can be used to provide automated guidance (e.g., self-driving vehicles), general system control, etc. Any of these and many other embodiments are possible and contemplated. Note that Fig.13 The devices and applications shown are exemplary only and are not intended to be limiting. Other devices are possible and contemplated.

[0093] Computer readable storage medium

[0094] Now turn to Fig.14 , a block diagram of an embodiment of a computer-readable storage medium 800 is shown. Generally speaking, a computer-accessible storage medium may include any storage medium that can be accessed by a computer during use to provide instructions and / or data to a computer. For example, a computer-accessible storage medium may include a storage medium such as a magnetic or optical medium, for example, a disk (fixed or removable), a tape, a CD-ROM, a DVD-ROM, a CD-R, a CD-RW, a DVD-R, a DVD-RW or a blue-ray. The storage medium may also include a volatile or non-volatile memory medium, such as a RAM (for example, a synchronous dynamic RAM (SDRAM), a Rambus DRAM (RDRAM), a static RAM (SRAM) etc.), a ROM or a flash memory. The storage medium may be physically included in a computer to which the storage medium provides instructions / data. Alternatively, the storage medium may be connected to a computer. For example, the storage medium may be connected to a computer via a network or a wireless link such as a network attached storage device. The storage medium may be connected via a peripheral interface such as a universal serial bus (USB). Typically, the computer-accessible storage medium 800 can store data in a non-transitory manner, where non-transitory in this context can refer to not transmitting instructions / data via signals. For example, a non-transitory storage device can be volatile (and may lose the stored instructions / data in response to a power outage) or non-volatile.

[0095] Fig.14The computer-accessible storage medium 800 in the computer-accessible storage medium 800 may store a database 802 representing the SOC 300. In general, the database 802 may be a database that can be read by a program and used directly or indirectly to manufacture the hardware including the SOC 300. For example, the database may be a behavioral level description of the hardware functions in a high-level design language (HDL) such as Verilog or VHDL or a register transfer level (RTL) description. The description may be read by a synthesis tool, which may synthesize the description to generate a netlist including a list of gate circuits from a synthesis library. The netlist includes a set of gate circuits that also represent the functions of the hardware including the SOC 300. The netlist may then be placed and routed to generate a data set for describing the geometry to be applied to the mask. The mask may then be used in various semiconductor manufacturing steps to generate one or more semiconductor circuits corresponding to the SOC 300. Alternatively, the database 802 on the computer-accessible storage medium 800 may be a netlist (with or without a synthesis library) or a data set, as desired.

[0096] Although computer-accessible storage medium 800 stores a representation of SOC 300, other embodiments may carry a representation of any portion of SOC 300 as desired, including Fig.12 In addition, database 802 may represent any subset of the components shown. Figure 1 The processors 10A-10N, the coprocessor 12, or both are shown, and may also represent the LLC 14 and / or the memory controller 16. The database 802 may represent any of the above.

[0097] ***

[0098] This disclosure includes references to "an embodiment" or groups of "embodiments" (e.g., "some embodiments" or "various embodiments"). Embodiments are different specific implementations or examples of the disclosed concepts. References to "an embodiment," "one embodiment," "a particular embodiment," etc. do not necessarily refer to the same embodiment. A large number of possible embodiments are contemplated, including those specifically disclosed, as well as modifications or substitutions that fall within the spirit or scope of this disclosure.

[0099] The present disclosure may discuss potential advantages that may be generated by the disclosed embodiments. Not all of these implementations will necessarily exhibit any or all of the potential advantages. Whether a particular implementation achieves an advantage depends on many factors, some of which are outside the scope of the present disclosure. In fact, there are many reasons why a particular implementation that falls within the scope of the claims may not exhibit some or all of any of the disclosed advantages. For example, a particular implementation may include other circuits outside the scope of the present disclosure, which, in combination with one of the disclosed embodiments, negate or weaken one or more of the disclosed advantages. In addition, suboptimal design execution of a particular implementation (e.g., a specific implementation technique or tool) may also negate or weaken the disclosed advantages. Even if a specific implementation of the technology is assumed, the realization of the advantages may still depend on other factors, such as the environmental conditions in which the specific implementation is deployed. For example, the input provided to a particular implementation may prevent one or more problems solved in the present disclosure from occurring in a particular occasion, and as a result, the benefits of its solution may not be realized. In view of the existence of possible factors external to the present disclosure, any potential advantages described herein should not be understood as claim limitations that must be met in order to prove infringement. Instead, the identification of such potential advantages is intended to show one or more types of improvements available to designers who benefit from the present disclosure. Permanently describing such advantages (eg, stating that a particular advantage "may occur") is not intended to convey a doubt as to whether such advantage can actually be achieved, but rather to recognize that achievement of such advantage typically depends on technical realities of additional factors.

[0100] Unless otherwise indicated, the embodiments are non-limiting. That is, the disclosed embodiments are not intended to limit the scope of the claims drafted based on the present disclosure, even if only a single example is described for a particular feature. The embodiments disclosed by the present invention are intended to be illustrative and non-restrictive, without any contrary statement in the present disclosure. Therefore, the present application is intended to allow claims covering the disclosed embodiments, as well as such alternatives, modifications and equivalent forms, which will be apparent to those skilled in the art who are aware of the effective effects of the present disclosure.

[0101] For example, features in this application may be combined in any suitable manner. Accordingly, new claims may be made during the prosecution of this patent application (or a patent application claiming priority thereto) for any such combination of features. In particular, with reference to the appended claims, features of dependent claims may be combined, where appropriate, with features of other dependent claims, including claims that are dependent on other independent claims. Similarly, features from corresponding independent claims may be combined, where appropriate.

[0102] Thus, while the appended dependent claims may be drafted such that each dependent claim is dependent on a single other claim, additional dependencies are also contemplated. Any combination of dependent features consistent with the present disclosure is contemplated and may be claimed in this or another patent application. In short, the combinations are not limited to those specifically listed in the appended claims.

[0103] It is also contemplated that claims drafted in one format or legal type (eg, apparatus) are intended to support corresponding claims in another format or legal type (eg, method), where appropriate.

[0104] ***

[0105] Because this disclosure is a legal document, various terms and phrases may be subject to regulatory and judicial interpretation. Notice is hereby given that the definitions provided in the following paragraphs and throughout this disclosure will be used to determine how to interpret claims drafted based on this disclosure.

[0106] Unless the context clearly dictates otherwise, reference to an item in the singular (i.e., a noun or noun phrase preceded by "a," "an," or "the") is intended to mean "one or more." Thus, reference to "an item" in a claim does not exclude additional instances of that item without the accompanying context. A "plurality" item refers to a collection of two or more items.

[0107] The word "may" is used herein in a permissive sense (ie, having the potential to, being able to), rather than the mandatory sense (ie, must).

[0108] The terms "including" and "comprising" and forms thereof are open ended and mean "including, but not limited to."

[0109] When the term "or" is used in this disclosure with respect to a list of options, it will generally be understood to be used in an inclusive sense unless the context provides otherwise. Thus, the expression "x or y" is equivalent to "x or y, or both," thus covering 1) x but not y, 2) y but not x, and 3) both x and y. On the other hand, phrases such as "either x or y, but not both" make it clear that "or" is used in an exclusive sense.

[0110] The expressions "w, x, y, or z, or any combination thereof" or ". . . at least one of w, x, y, and z" are intended to cover all possibilities involving individual elements up to the total number of elements in the set. For example, given the set [w, x, y, z], these phrases cover any single element in the set (e.g., w but not x, y, or z), any two elements (e.g., w and x, but not y or z), any three elements (e.g., w, x, and y, but not z), and all four elements. The phrase ". . . at least one of w, x, y, and z" thus refers to at least one element in the set [w, x, y, z], thereby covering all possible combinations in the list of elements. The phrase should not be interpreted as requiring the presence of at least one instance of w, at least one instance of x, at least one instance of y, and at least one instance of z.

[0111] In this disclosure, various "labels" may precede a noun or noun phrase. Unless the context provides otherwise, different labels used for a feature (e.g., "first circuit," "second circuit," "particular circuit," "given circuit," etc.) refer to different instances of the feature. In addition, unless otherwise specified, the labels "first," "second," and "third" do not imply any type of ordering (e.g., spatial, temporal, logical, etc.) when applied to features.

[0112] The phrase "based on" or is used to describe one or more factors that influence a determination. This term does not exclude that there may be additional factors that may influence the determination. That is, the determination may be based only on the specified factors or on the specified factors and other unspecified factors. Consider the phrase "A is determined based on B." This phrase specifies that B is a factor used to determine A or that B influences the determination of A. This phrase does not exclude that the determination of A may also be based on some other factor such as C. This phrase is also intended to cover embodiments in which A is determined based only on B. As used herein, the phrase "based on" is synonymous with the phrase "based at least in part on."

[0113] The phrases "in response to" and "in response to" describe one or more factors that trigger an effect. The phrase does not exclude the possibility that additional factors may influence or otherwise trigger the effect, either in conjunction with the specified factors or independently of the specified factors. That is, the effect may be responsive to these factors alone, or may be responsive to the specified factors as well as other unspecified factors. Consider the phrase "in response to B, A is performed." The phrase specifies that B is a factor that triggers the execution of A or triggers a specific result of A. The phrase does not exclude that the execution of A may also be responsive to some other factor, such as C. The phrase also does not exclude that the execution of A may be performed jointly in response to B and C. This phrase is also intended to cover embodiments in which A is performed only in response to B. As used herein, the phrase "in response to" is synonymous with the phrase "at least partially in response to." Similarly, the phrase "in response to" is synonymous with the phrase "at least partially in response to."

[0114] ***

[0115] Within the present disclosure, different entities (which may be variously referred to as "units," "circuits," other components, etc.) may be described or claimed as being "configured to" perform one or more tasks or operations. This expression—[an entity] configured to [perform one or more tasks]—is used herein to refer to a structure (i.e., a physical thing). More specifically, this expression is used to indicate that this structure is arranged to perform one or more tasks during operation. A structure may be said to be "configured to" perform a task even if the structure is not currently being operated. Thus, an entity described or stated as "configured to" perform a task refers to a physical thing used to implement the task, such as a device, a circuit, a system with a processor unit, and a memory storing executable program instructions, etc. This phrase is not used herein to refer to an intangible thing.

[0116] In some cases, various units / circuits / components may be described herein as performing a set of tasks or operations. It should be understood that these entities are "configured to" perform those tasks / operations, even if not specifically stated.

[0117] The term "configured to" is not intended to mean "configurable to". For example, an unprogrammed FPGA would not be considered "configured to" perform a particular function. However, the unprogrammed FPGA may be "configurable to" perform that function. After being appropriately programmed, the FPGA may then be considered "configured to" perform the particular function.

[0118] For purposes of a U.S. patent application based on the present disclosure, stating in a claim that a structure is “configured to” perform one or more tasks is expressly intended not to invoke 35 U.S.C. §112(f) for that claim element. If the applicant wishes to invoke section 112(f) during prosecution of a U.S. patent application based on the present disclosure, it would use a “means for [performing the function]” structure to phrase the claim element.

[0119] Different "circuits" may be described in the present disclosure. These circuits or "circuitry" constitute hardware that includes various types of circuit elements, such as combinational logic, clock storage devices (e.g., flip-flops, registers, latches, etc.), finite state machines, memories (e.g., random access memory, embedded dynamic random access memory), programmable logic arrays, etc. The circuits may be custom designed or taken from a standard library. In various specific implementations, the circuitry may include digital components, analog components, or a combination of both, as appropriate. Certain types of circuits may be generally referred to as "units" (e.g., decoding units, arithmetic logic units (ALUs), functional units, memory management units (MMUs), etc.). Such units are also referred to as circuits or circuitry.

[0120] Thus, the disclosed circuits / units / components and other elements shown in the drawings and described herein include hardware elements, such as those described in the preceding paragraphs. In many cases, the internal arrangement of hardware elements in a particular circuit can be specified by describing the functionality of that circuit. For example, a particular "decode unit" may be described as performing the function of "processing an opcode for an instruction and routing the instruction to one or more of a plurality of functional units," meaning that the decode unit is "configured to" perform that function. For a person skilled in the computer arts, this functional specification is sufficient to suggest a set of possible structures for the circuit.

[0121] In various embodiments, as discussed in the preceding paragraphs, circuits, units, and other elements defined by the functions or operations they are configured to implement. The arrangement of such circuits / units / components relative to each other and the way they interact form a microarchitecture definition of hardware, which is ultimately manufactured in an integrated circuit or programmed into an FPGA to form a physical implementation of the microarchitecture definition. Therefore, the microarchitecture definition is considered by those skilled in the art to be a structure from which many physical implementations can be derived, all of which fall into a broader structure described by the microarchitecture definition. That is, a technician with a microarchitecture definition provided in accordance with the present disclosure can implement the structure by coding a description of a circuit / unit / component in a hardware description language (HDL) such as Verilog or VHDL without excessive experimentation and using the application of ordinary technicians. HDL descriptions are often expressed in a way that can be displayed as functional. However, for those skilled in the art, the HDL description is a way to convert the structure of a circuit, unit, or component into the next level of specific implementation details. Such HDL descriptions may take the form of behavioral code (which is generally non-synthesizable), register transfer language (RTL) code (which is generally synthesizable compared to behavioral code), or structural code (e.g., a netlist specifying logic gates and their connectivity). The HDL description may be sequentially synthesized for a library of cells designed for a given integrated circuit manufacturing technology, and may be modified for timing, power, and other reasons to obtain a final design database that is transmitted to the factory to generate masks and ultimately produce integrated circuits. Some hardware circuits or portions thereof may also be custom designed in a schematic editor and captured into an integrated circuit design along with the synthesized circuitry. The integrated circuit may include transistors and other circuit elements (e.g., passive elements such as capacitors, resistors, inductors, etc.), as well as interconnects between transistors and circuit elements. Some embodiments may implement multiple integrated circuits coupled together to implement the hardware circuit, and / or discrete elements may be used in some embodiments. Alternatively, the HDL design may be synthesized into a programmable logic array such as a field programmable gate array (FPGA), and may be implemented in an FPGA. This decoupling between the design of a set of circuits and the subsequent low-level implementation of those circuits often leads to a situation where a circuit or logic designer never specifies a specific set of structures for the low-level implementation beyond a description of what the circuits are configured to do, because that process is performed at different stages of the circuit implementation process.

[0122] The fact that many different low-level combinations of circuit elements can be used to achieve the same specification of a circuit results in a large number of equivalent structures for that circuit. As noted, these low-level circuit implementations can vary depending on variations in manufacturing technology, the foundry selected to manufacture the integrated circuit, the cell libraries provided for a particular project, etc. In many cases, the selection made by different design tools or methodologies to produce these different implementations can be arbitrary.

[0123] Furthermore, for a given embodiment, a single implementation of a particular functional specification of a circuit typically includes a large number of devices (e.g., millions of transistors). Thus, the shear volume of this information makes it impractical to provide a complete description of the low-level structure used to implement a single embodiment, let alone the large number of equivalent possible implementations. For this reason, the present disclosure describes the structure of the circuit using functional shorthand commonly used in the industry.

[0124] Various embodiments are contemplated, as illustrated in the following numbered examples:

[0125] 1. A system comprising:

[0126] one or more processors configured to issue memory requests to access the memory system; and

[0127] a cache coupled to the one or more processors and configured to cache data from the memory system for access by the one or more processors, wherein:

[0128] The cache includes control circuitry configured to assign criticality values ​​to cache lines, wherein a given criticality value corresponds to a given cache line;

[0129] During selection of a victim cache line for replacement, the cache is configured to preferentially retain cache lines identified as critical by corresponding criticality values ​​over cache lines not identified as critical, wherein the cache is further configured to select the victim cache line based on replacement data maintained by the cache separate from the criticality value;

[0130] The control circuitry is configured to detect one or more indications that at least some of the cache lines identified as critical are no longer critical;

[0131] and

[0132] The control circuitry is configured to terminate the preferential retention of the cache line based on the one or more indications.

[0133] 2. A system according to embodiment 1, wherein the control circuit is configured to monitor a cache hit rate for cache lines in multiple least recently used locations in the replacement data and having a criticality value indicating a critical state, and wherein one of the one or more indications is based on the cache hit rate being below a threshold level.

[0134] 3. A system according to embodiment 1 or 2, wherein the control circuit is configured to monitor a rate at which snoop hits occur in the cache and cause forwarding of corresponding cache lines in response to the snoop hits, and wherein one of the one or more indications is based on the snoop hit rate being above a threshold level.

[0135] 4. The system according to any one of embodiments 1 to 3 further includes a coprocessor, which is coupled to the cache and is configured to execute coprocessor instructions issued by the one or more processors to the coprocessor, and wherein the one or more instructions include a memory request issued by the coprocessor to the cache.

[0136] 5. The system of embodiment 4, wherein the control circuit is configured to infer coprocessor memory requests based on prefetch requests generated by the one or more processors that are indicated as coprocessor prefetch requests.

[0137] 6. The system of any preceding embodiment, further comprising a second cache coupled to the cache, wherein the second cache is configured to provide an indication of capacity in the second cache that is allocable to data from the cache, and wherein the criticality value indicates non-critical and a plurality of criticality levels, and wherein the control circuitry is configured to use the plurality of criticality levels based on the indication that the capacity is less than a threshold.

[0138] The plurality of criticality levels are covered by the lowest criticality level among the plurality of criticality levels.

[0139] 7. The system of any preceding embodiment, wherein the control circuit is configured to terminate the priority retention at least in part by ignoring the criticality value updated for victim selection and replacement data.

[0140] 8. A system comprising:

[0141] One or more processors configured to issue access

[0142] a memory request to the memory system; and

[0143] A cache memory coupled to the one or more processors and configured to cache data from the memory system for use by the one or more processors

[0144] access, wherein the cache includes control circuitry, and wherein:

[0145] The control circuitry is configured to assign criticality values ​​to cache lines, wherein a given criticality value corresponds to a given cache line;

[0146] During selection of a victim cache line for replacement, the cache is configured to preferentially retain cache lines identified as critical by corresponding criticality values ​​over cache lines not identified as critical, wherein the cache is further configured to select the victim cache line based on replacement data maintained by the cache separate from the criticality value;

[0147] The control circuitry is configured to detect one or more indications that at least some of the cache lines identified as critical are no longer critical; and

[0148] The control circuitry is configured to accelerate eviction of the cache line identified as critical based on the one or more indications.

[0149] 9. A system according to embodiment 8, wherein the control circuit is configured to monitor a cache hit rate for cache lines in multiple least recently used locations in the replacement data and having a criticality value indicating a critical state, and wherein one of the one or more indications is based on the cache hit rate being below a threshold level.

[0150] 10. A system according to embodiment 8 or 9, wherein the control circuit is configured to monitor a rate at which snoop hits occur in the cache and cause forwarding of corresponding cache lines in response to the snoop hits, and wherein one of the one or more indications is based on the snoop hit rate being above a threshold level.

[0151] 11. A system according to any one of embodiments 8 to 10, wherein the one or more instructions include a memory request issued by a coprocessor to the cache, wherein the coprocessor is coupled to the cache and is configured to execute coprocessor instructions issued by the one or more processors to the coprocessor.

[0152] 12. The system of embodiment 11, wherein the control circuit is configured to infer the coprocessor memory request based on a prefetch request generated by the one or more processors that is indicated as a coprocessor prefetch request.

[0153] 13. A system according to any one of embodiments 8 to 12, further comprising a second cache, wherein the second cache is configured to provide an indication of capacity in the second cache that can be allocated to data from the cache, and wherein the criticality value indicates non-critical and multiple criticality levels, and wherein the control circuit is configured to overwrite the multiple criticality levels with a lowest criticality level of the multiple criticality levels based on the indication that the capacity is less than a threshold.

[0154] 14. The system of any one of embodiments 8 to 13, wherein the control circuit is configured to accelerate eviction at least in part by ignoring the criticality value for victim selection and replacement data updates.

[0155] 15. A method comprising:

[0156] Assigning criticality values ​​to cache lines in a cache, where a given criticality

[0157] The property value corresponds to a given cache line;

[0158] During selection of a victim cache line for replacement, cache lines identified as critical by corresponding criticality values ​​are preferentially retained over cache lines not identified as critical, wherein the selection is also based on a cache memory retained by the cache memory corresponding to the cache memory.

[0159] Replacement data separated by criticality values;

[0160] Detecting at least some of the cache lines identified as critical

[0161] One or more indications that the row is no longer critical; and

[0162] The criticality value for victim selection and replacement data updates is ignored based on the one or more indications.

[0163] 16. The method according to embodiment 15 also includes monitoring a cache hit rate for cache lines in multiple least recently used locations in the replacement data and having a criticality value indicating a critical state, and wherein one of the one or more indications is based on the cache hit rate being below a threshold level.

[0164] 17. The method of embodiment 15 or 16 further includes monitoring a rate at which snoop hits occur in the cache and causing forwarding of a corresponding cache line in response to the snoop hits, and wherein one of the one or more indications is based on the snoop hit rate being above a threshold level.

[0165] 18. A method according to any one of embodiments 15 to 17, wherein the one or more instructions include a memory request issued by a coprocessor to the cache, wherein the coprocessor is coupled to the cache and is configured to execute coprocessor instructions issued by one or more processors to the coprocessor.

[0166] 19. The method of embodiment 18, further comprising inferring coprocessor memory requests based on prefetch requests generated by one or more processors that are indicated as coprocessor prefetch requests.

[0167] 20. The method of any one of embodiments 15 to 19, further comprising:

[0168] providing, from a second cache, an indication of capacity in the second cache that is allocable to data from the cache, wherein the criticality value indicates a non-critical

[0169] and multiple levels of criticality; and

[0170] The plurality of criticality levels are overwritten with a lowest criticality level of the plurality of criticality levels based on the indication that capacity is less than a threshold.

[0171] Numerous variations and modifications will become apparent to those skilled in the art once the above disclosure is fully appreciated. It is intended that the following claims be interpreted to encompass all such variations and modifications.

Claims

1. A device, include: one or more processors configured to issue memory requests to access a memory system; and a cache circuit configured to cache data from the memory system for access by the one or more processors, wherein: The cache circuit includes a control circuit configured to: assigning a criticality value to a cache line, wherein a given criticality value indicates that the corresponding cache line is non-critical or has a criticality level of one of a plurality of criticality levels; and adjusting a criticality value of a cache line in response to a criticality event; To select a victim cache line for replacement, the cache circuit is configured to: masking cache lines having one or more higher criticality levels; and The victim cache line is selected from the unmasked cache lines based on access recency data maintained by the cache circuit separately from the criticality value.

2. The device according to claim 1, further comprising: include: Loading queue circuits; Wherein the criticality event corresponds to a load accessing the cache line being the oldest outstanding load in the load queue circuit.

3. The device according to claim 1, further comprising: include: Memory management circuitry, wherein the criticality event corresponds to an access to the cache line requested by the memory management circuitry.

4. The device according to claim 1, further comprising: include: a miss queue circuit for fetching requests that miss in the instruction cache; Wherein the criticality event corresponds to a fetch request accessing the cache line being the oldest outstanding fetch request in the miss queue circuit.

5. The device of claim 1, wherein the adjustment is one of: Assigning the lowest critical threshold values ​​to replace non-critical values; and Increases the criticality of an already critical criticality value.

6. The apparatus of claim 1, wherein masking cache lines having one or more higher criticality levels comprises masking lesser criticality levels until a candidate victim cache line is found.

7. The apparatus of claim 1 , wherein the cache circuit is further configured to: The apparatus includes control circuitry configured to monitor a capacity of a higher level cache that can be allocated to data from the cache circuitry; and The cache circuit is configured to assign a cache line to a lowest criticality level based on the monitored capacity being below a threshold.

8. The device according to claim 1, in, To allocate a cache entry for a non-critical cache line, the cache circuit is configured to determine an artificial access recency data value for the non-critical cache line based on access recency data values ​​of cached critical cache lines.

9. The apparatus of claim 8, wherein the artificial access recency value is further back than the access recency value of the least recently used cached critical cache line.

10. The device according to claim 1, in: The apparatus includes control circuitry configured to monitor a capacity of a higher level cache that can be allocated to data from the cache circuitry; and To allocate a cache entry for a non-critical cache line, the cache circuit is configured to determine an artificial access recency data value for the non-critical cache line based on the monitored capacity.

11. The apparatus of claim 1, wherein the apparatus is a computing device, the computing device further comprising: include: Display control circuit; and Network interface circuit.

12. The device of claim 1, wherein the device is an integrated circuit.

13. A method, include: A memory request is issued by a computing device to access a memory system; caching, by the computing device, data from the memory system; assigning, by the computing device, a criticality value to a cache line, wherein a given criticality value indicates that the corresponding cache line is non-critical or has a criticality level of one of a plurality of criticality levels; adjusting, by the computing device, a criticality value of a cache line in response to a criticality event; Selecting, by the computing device, a victim cache line for replacement includes: masking cache lines having one or more higher criticality levels; as well as The victim cache line is selected from the unmasked cache lines based on access recency data maintained separately from the criticality value.

14. The method of claim 13, wherein the criticality event corresponds to a load accessing the cache line being the oldest outstanding load in a load queue.

15. The method of claim 13, wherein the criticality event corresponds to an access to the cache line requested by memory management circuitry.

16. The method of claim 13, wherein the criticality event corresponds to a fetch request to access the cache line being the oldest outstanding fetch request in a miss queue circuit.

17. The method of claim 13, wherein the masking comprises masking lesser criticality levels until a candidate victim cache line is found.

18. The method according to claim 13, further comprising: include: Allocating a cache entry for a non-critical cache line includes determining an artificial access recency data value for the non-critical cache line based on an access recency data value of a cached critical cache line.

19. The method of claim 18, wherein the artificial access recency value is further back than the access recency value of the least recently used cached critical cache line.

20. A non-transitory computer readable medium having stored thereon instructions in a hardware description programming language, which when processed by a computing system, program the computing system to generate a computer model, wherein the model represents a hardware circuit, the hardware circuit include: one or more processors configured to issue memory requests to access a memory system; and a cache circuit configured to cache data from the memory system for access by the one or more processors, wherein: The cache circuit includes a control circuit configured to: assigning a criticality value to a cache line, wherein a given criticality value indicates that the corresponding cache line is non-critical or has a criticality level of one of a plurality of criticality levels; and adjusting a criticality value of a cache line in response to a criticality event; To select a victim cache line for replacement, the cache circuit is configured to: masking cache lines having one or more higher criticality levels; and The victim cache line is selected from the unmasked cache lines based on access recency data maintained by the cache circuit separately from the criticality value.

21. A device, include: one or more processors configured to issue memory requests to access a memory system; a first cache configured to cache data from the memory system for access by the one or more processors; a second cache configured to cache data from the memory system for access by the first cache; and A control circuit, the control circuit being configured to: assigning a criticality value to a cache line of the first cache based on one or more operating conditions when the cache line is populated into the cache; for a first cache line evicted from the first cache to the second cache, setting an initial retention priority value in the second cache based on a criticality value assigned to the first cache line; as well as A victim cache line for eviction from the second cache is determined based on the initial retention priority value.

22. The apparatus of claim 21, wherein the control circuit is further configured to set the initial retention priority value based on available capacity in the second cache when the first cache line is evicted to the second cache.

23. The apparatus of claim 22, wherein the control circuit is configured to map a range of available capacity to a hint value of the initial retention priority.

24. The apparatus of claim 21, wherein the control circuitry is further configured to retain an indication of a criticality value of the first cache line in the second cache in response to the eviction.

25. The apparatus of claim 21, wherein the second cache is a memory cache associated with a memory controller.

26. The apparatus of claim 21, wherein the criticality value comprises a plurality of critical values ​​and at least one non-critical value.

27. The apparatus of claim 21, wherein the control circuitry is configured to determine, in response to an access hitting the first cache line, an updated replacement value for the first cache line of the first cache based on the criticality value assigned to the first cache line.

28. The apparatus of claim 21, wherein the control circuitry is configured to prevent cache lines indicated as critical in the first cache from being selected for eviction for at least one victim selection process.

29. The apparatus of claim 28, wherein the control circuit is configured to ignore criticality values ​​for one or more cache update processes based on a cache hit rate criterion.

30. The apparatus of claim 28, wherein the control circuit is configured to ignore criticality values ​​for one or more cache update processes based on a snoop hit rate criterion.

31. The apparatus of claim 21, wherein the one or more operating conditions include at least one of the following factors: whether the load accessing the given cache line is the oldest outstanding load in a load queue circuit of the device; whether access to a given cache line is requested by memory management circuitry of the device; as well as Whether the access to a given cache line is for a fetch request that is the oldest outstanding fetch request in a miss queue circuit of the device.

32. The apparatus of claim 21, wherein the control circuitry is further configured to assign a retention priority value to a cache line of the first cache.

33. The apparatus of claim 21, wherein the apparatus is a computing device, the computing device further comprising: include: Network interface circuit; and Display controller.

34. The apparatus of claim 21, wherein the control circuit is further configured to: The initial retention priority value is updated based on one or more accesses to the second cache, wherein the determination of the victim cache line is based on the updated retention priority value.

35. A method, include: issuing, by one or more processors of a computing system, a request to access a memory system of the computing system; caching, by a first cache of the computing system, data from the memory system for access by the one or more processors; caching data from the memory system by a second cache of the computing system for access by the first cache; assigning, by the computing system, a criticality value to a cache line of the first cache when the cache line is populated into the cache based on one or more operating conditions; setting, by the computing system, an initial retention priority value in the second cache for a first cache line evicted from the first cache to the second cache based on a criticality value assigned to the first cache line; as well as A victim cache line for eviction from the second cache is determined, by the computing system, based on the initial retention priority value.

36. The method of claim 35, wherein the setting is further based on available capacity in the second cache when the first cache line is evicted to the second cache.

37. The method according to claim 36, further comprising: include: Ranges of available capacity are mapped, by the computing system, to hint values ​​for initial reservation priorities.

38. The method according to claim 35, further comprising: include: An indication of the criticality value of the first cache line is retained in the second cache by the computing system in response to the eviction.

39. The method according to claim 35, further comprising: include: An updated replacement value for the first cache line of the first cache is determined by the computing system in response to an access hitting the first cache line based on the criticality value assigned to the first cache line.

40. A non-transitory computer readable medium having stored thereon instructions in a hardware description programming language, which when processed by a computing system, program the computing system to generate a computer model, wherein the model represents a hardware circuit, the hardware circuit include: one or more processors configured to issue memory requests to access a memory system; a first cache configured to cache data from the memory system for access by the one or more processors; a second cache configured to cache data from the memory system for access by the first cache; and A control circuit, the control circuit being configured to: assigning a criticality value to a cache line of the first cache based on one or more operating conditions when the cache line is populated into the cache; for a first cache line evicted from the first cache to the second cache, setting an initial retention priority value in the second cache based on a criticality value assigned to the first cache line; as well as A victim cache line for eviction from the second cache is determined based on the initial retention priority value.