Adaptive Cache Partitioning

Through adaptive cache partitioning technology, dynamically adjusting the partitioning of cache memory, solving the problem of balancing cache capacity and prefetching performance, improving prefetching performance and cache efficiency, and adapting to different workload conditions.

CN114077553BActive Publication Date: 2025-07-29MICRON TECHNOLOGY INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110838429.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-08-19
Filing Date
2021-07-23
Publication Date
2025-07-29
Estimated Expiration
2041-07-23

AI Technical Summary

Technical Problem

The prior art is difficult to effectively balance the capacity and prefetch performance of cache memory, especially when facing different workloads, resulting in poor prefetch performance and reduced cache efficiency.

Method used

Through adaptive cache partitioning technology, the partition of the cache memory is dynamically adjusted, and part of it is allocated to store metadata related to the address space, and the other part is used to store data, and the amount of cache memory allocated by the metadata is dynamically adjusted according to the prefetch energy degree.

Benefits of technology

Improves prefetch performance and cache efficiency, adapts to different workload conditions, reduces invalid prefetching, and improves overall memory access speed and hit rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114077553B_ABST
    Figure CN114077553B_ABST
Patent Text Reader

Abstract

This application relates to adaptive cache partitioning. The described apparatus and method partition a cache memory at least in part based on a measure indicative of prefetch performance. The amount of cache memory allocated for metadata related to prefetch operations on cache storage can be adjusted based on operating conditions. Accordingly, the cache memory can be partitioned into a first portion allocated for metadata related to an address space (prefetch metadata) and a second portion allocated for data associated with the address space (cache data). The amount of cache memory allocated to the first portion can be increased under workloads suitable for prefetching and decreased otherwise. The first portion can include one or more cache units, cache lines, cache ways, cache sets, or other resources of the cache memory.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to methods, devices, and systems for adaptive cache partitioning. Background Art

[0002] To operate effectively, some computing systems include a hierarchical memory system that may include multiple memory levels. Here, efficient operation may require cost efficiency and speed efficiency. Since faster memories are typically more expensive than relatively slower memories, designers attempt to balance their relative costs and benefits. One approach is to use a smaller amount of faster memory and a larger amount of slower memory. In a hierarchical memory system, the faster memory is deployed at a higher level than the slower memory so that the faster memory can be preferentially accessed first. An example of a relatively fast memory is a cache memory. An example of a relatively slow memory is a backing memory, which may include primary memory, main memory, a backing storage device, etc.

[0003] A cache memory can accelerate data operations by storing and retrieving data from a backing memory using, for example, high-performance memory cells. The high-performance memory cells enable the cache memory to respond to memory requests faster than the backing memory. Thus, a cache memory can achieve a faster response from the memory system based on the desired data present in the cache. One way to increase the likelihood that the desired data is present in the cache is to prefetch data before it is requested. To this end, a prefetch system attempts to predict which data the processor will request and then loads the predicted data into the cache. Although a prefetch system can make it more likely that a cache memory will accelerate memory access operations, data prefetching may introduce operational complexities that engineers and other computer designers strive to overcome. Summary of the Invention

[0004] Devices and methods partition a cache memory at least in part based on a metric indicative of prefetch performance. The amount of cache memory allocated for metadata related to prefetch operations on cache storage can be adjusted based on operating conditions. Thus, the cache memory can be partitioned into a first portion allocated for metadata related to an address space (prefetch metadata) and a second portion allocated for data related to the address space (cache data). The amount of cache memory allocated to the first portion can be increased under workloads suitable for prefetching and decreased otherwise. The first portion can include one or more cache units, cache lines, cache ways, cache sets, or other resources of the cache memory.

[0005] Another aspect of the present application relates to a method, which includes: allocating a first portion of a cache memory for metadata related to an address space; writing data associated with an address of the address space to a second portion of the cache memory different from the first portion of the cache memory; and modifying the size of the first portion of the cache memory allocated for the metadata related to the address space at least in part based on a metric related to data prefetched into the second portion of the cache memory.

[0006] Another aspect of the present application relates to an apparatus, which includes: a memory array configured as a cache memory; and logic coupled to the memory array and configured to: allocate a first portion of the cache memory to store metadata related to an address space; determine a metric related to cache data loaded into a second portion of the cache memory, the second portion being different from the first portion; and modify the amount of the cache memory allocated to the first portion at least in part based on the determined metric.

[0007] Another aspect of the present application relates to a system, which includes: a cache memory including a plurality of cache units; an interface configured to be coupled to an interconnect of a computing device; and logic coupled to the interface and the cache memory and configured to: partition the cache memory into a first portion and a second portion, allocate the first portion for metadata related to an address space associated with a memory, and the second portion is allocated for storing cache data corresponding to addresses of the address space; write the metadata related to the address space to the first portion of the cache memory; write data to the second portion of the cache memory at least in part based on the metadata written to the first portion of the cache memory; and modify the amount of the cache memory allocated to the first portion at least in part based on a metric related to the data written to the second portion of the cache memory. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] This document describes in detail one or more aspects of adaptive cache partitioning with reference to the following drawings. Throughout the drawings, like numerals are used to refer to like features and components:

[0009] Figures 1-1 to 1-3 An example environment is shown in which techniques for implementing adaptive cache partitioning may be employed.

[0010] Figure 2An example of a device that can implement some aspects of adaptive cache partitioning is shown.

[0011] Figure 3 Another example of a device that can implement some aspects of adaptive cache partitioning is shown.

[0012] Figures 4-1 to 4-3 An example operational implementation of adaptive cache partitioning is shown.

[0013] Figures 5-1 to 5-3 Another example operational implementation of adaptive cache partitioning is shown.

[0014] Figure 6 An example flowchart depicting the operation of adaptive cache partitioning is shown.

[0015] Figure 7 An example flowchart depicting the operation of adaptive cache partitioning is shown.

[0016] Figure 8 An example flowchart depicting the operation of adaptive cache partitioning at least in part based on a metric related to prefetch performance is shown.

[0017] Figure 9 An example flowchart depicting the operation of adaptive cache partitioning at least in part based on a metric related to cache and / or prefetch performance is shown.

[0018] Figure 10 An example of a system for implementing adaptive cache partitioning is shown. Detailed Description

[0019] Overview

[0020] Advances in semiconductor process technology and microarchitecture have led to a significant reduction in processor cycle time and an increase in processor density. At the same time, advances in memory technology have led to an increase in memory density but a relatively small reduction in memory access time. As a result, the memory latency measured in processor clock cycles is continuously increasing. However, cache memory can help bridge the processor-memory latency gap. Compared to the backing memory, the cache memory that can store the data of the backing memory can serve requests faster. In some aspects, the cache memory can be deployed "above" or "in front of" the backing memory in the memory hierarchy so that the cache memory is preferably accessed before accessing the slower backing memory.

[0021] Especially due to cost considerations, the capacity of the cache memory may be lower than the capacity of the backing memory or the main memory. The cache memory can thus load a selected subset of the address space of the backing memory. Data can be selectively admitted and / or evicted from the cache memory according to suitable criteria (such as cache admission policies, eviction policies, replacement policies, etc.).

[0022] During operation, data can be loaded into the cache in response to a "cache miss". A cache miss refers to a request related to an address that has not yet been loaded into the cache and / or is not included in the working set of the cache. Serving a cache miss may involve fetching data from a slower backing memory, which may significantly degrade performance. In contrast, serving a request that results in a "cache hit" may involve accessing the cache memory with relatively high performance, without incurring the latency of accessing the relatively low-performance backing memory.

[0023] In some cases, cache performance can be improved by prefetching. Prefetching generally involves loading an address into the cache memory before the requested address. A prefetcher can predict the addresses of upcoming requests and preload the addresses into the cache memory in the background, so that when a request related to the predicted address is subsequently received, the cache memory can serve the request instead of triggering a cache miss. In other words, the relatively high-performance cache memory can be used to serve requests related to the prefetched addresses, without incurring the latency of the relatively low-performance backing memory.

[0024] The benefits of prefetching can be quantified in terms of hit rate, the number of "useful" prefetches, the ratio of useful prefetches to "invalid" prefetches, etc. As used herein, a "useful" prefetch refers to a prefetch that results in a subsequent cache hit, which is referred to as a "prefetch hit". In other words, a useful prefetch is achieved by prefetching data associated with an address that is subsequently requested and / or otherwise accessed from the cache memory. In contrast, an "invalid" prefetch or "prefetch miss" refers to a prefetch of data that will not be requested subsequently and thus will not result in a cache or prefetch hit. Invalid prefetches may have an adverse effect on performance. Invalid prefetches may consume limited cache memory resources as well as data that is unlikely to be requested (e.g., poisoning the cache), resulting in an increased cache miss rate, a lower hit rate, increased jitter, higher bandwidth consumption, etc.

[0025] A prefetcher can attempt to avoid these issues by attempting to detect patterns of memory access and then prefetching data based on the detected patterns. The prefetcher can use metadata to detect, predict, derive, and / or exploit memory access patterns to determine accurate prefetch predictions (e.g., predicting the addresses of upcoming requests). The metadata used by the prefetcher can be referred to as "prefetcher metadata", "prefetch metadata", "request metadata", "access metadata", "memory metadata", "memory access metadata", etc. Such metadata can contain any suitable information related to the address space, including but not limited to: sequences of previously requested addresses or address offsets, address history, address history tables, index tables, access frequencies of corresponding addresses, access counts (e.g., accesses within a corresponding window), access times, last access times, and so on.

[0026] The prefetcher can implement prefetch operations suitable for the workload of prefetching. As used herein, a "suitable" workload or a workload "suitable for prefetching" refers to a "predictable" workload that generates memory accesses according to patterns detectable (and / or exploitable) by the prefetcher. A suitable workload can thus refer to a workload associated with metadata from which the prefetcher can derive predictable access patterns. Examples of suitable workloads include workloads in which the memory request offsets are consistent offsets or strides. These types of workloads can be generated by programs that repeatedly and / or in a regular pattern access structured data. By way of non-limiting example, a program can repeatedly access a data structure of size D, thereby generating a predictable workload in which the memory access offsets are a relatively constant offset Δ or stride, where Δ≈D. The stride and other types of access patterns can be derived from metadata related to the previous memory accesses of the workload. The prefetcher can utilize the memory access patterns derived from such metadata to prefetch data that may be requested in the future. In the above stride example, in response to a cache miss at address a, the prefetcher can load the data at addresses a+Δ, a+2Δ, …… to a+dΔ into the cache memory (where d is a configurable prefetch degree). Given the predictable memory access patterns derived from the metadata associated with the workload, the data prefetched from a+Δ to a+dΔ will likely result in subsequent prefetch hits, thereby preventing cache misses and improving performance.

[0027] Certain types of workloads may not be suitable for prefetching. As used herein, an "unsuitable" workload or a workload "not suitable for prefetching" refers to a workload that accesses memory in a manner that the prefetcher cannot predict, model, and / or otherwise utilize to produce accurate prefetch predictions. An unsuitable workload may refer to a workload associated with metadata from which the prefetcher cannot derive address predictions, patterns, models, etc. Examples of unsuitable workloads include workloads generated by programs that access memory in non-repetitive and / or non-regular patterns, programs that access memory at seemingly random addresses and / or address offsets, programs that access memory according to patterns that are too complex or variable to be detected by the prefetcher (and / or cannot be captured in prefetcher metadata), etc. Attempting to prefetch data for an unsuitable workload may result in poor prefetch performance. Since prefetch decisions for unsuitable workloads are not guided by a discernible access pattern, little (if any) of the prefetched data may subsequently be requested before being evicted from the cache. As disclosed herein, inaccurate prefetch predictions may result in invalid prefetches that consume the relatively limited capacity of the cache memory and data that is subsequently unlikely to be accessed, while excluding other data that may be accessed more frequently. Thus, attempting to prefetch for an unsuitable workload may degrade cache performance (e.g., result in lower hit rates, increased miss rates, jitter, increased bandwidth consumption, etc.). To avoid these and other problems, prefetching may not be implemented for unsuitable workloads (and / or within address regions associated with unsuitable workloads).

[0028] A cache may serve multiple different workloads, each having corresponding workload characteristics (e.g., corresponding memory access characteristics, patterns, etc.). As further detailed herein, the workload characteristics within corresponding regions of the address space may depend on a variety of factors that vary over time. Thus, programs running in different regions of the address space may generate workloads having different characteristics (e.g., different memory access patterns). For example, a first program running in a first region of the address space may access memory in a first stride pattern (Δ1); a second program running in a second region may access memory at a different stride pattern (Δ2) per second; a third program running in a third region of the address space may access memory according to a more complex pattern (such as a correlation pattern); additional programs running in a fourth region of the address space may access memory unpredictably; etc. Although the first stride pattern may be able to produce accurate prefetching within the first region, if used in other regions, the first stride pattern will likely produce poor results (and vice versa).

[0029] In some embodiments, prefetch performance can be improved by maintaining metadata related to corresponding regions of an address space. The metadata used by the prefetcher can include multiple entries, where each entry contains information related to memory accesses made using the corresponding region of the address space. The prefetcher can utilize the metadata related to the corresponding region to inform prefetch operations within the corresponding region. More specifically, the prefetcher can utilize the metadata related to the corresponding region of the address space to determine characteristics of the workload within the corresponding region, determine whether the workload is suitable for prefetching (e.g., distinguish workloads and / or regions suitable for prefetching from those not suitable for prefetching), determine access patterns within the corresponding region, perform prefetch operations within the corresponding region based on the determined access patterns, and so on. In some embodiments, the prefetcher metadata covers multiple fixed-size address regions. Alternatively, the prefetcher metadata can be configured to cover address regions of adaptive size, where the workload characteristics (such as access patterns) are consistent. In these embodiments, the size of the address range covered by a corresponding entry of the prefetcher metadata can vary within the corresponding region of the address space, depending in particular on the workload characteristics and / or prefetch performance within the corresponding region.

[0030] Metadata related to memory access is often tracked at and / or within performance-sensitive functions of a hierarchical memory system (such as a memory I / O path, etc.). Additionally, prefetch operations that utilize such metadata can be performance-sensitive (e.g., to ensure that the prefetched data is available before such data is requested). Therefore, it can be advantageous to maintain metadata related to memory access (prefetcher metadata) within high-performance memory resources. In some embodiments, the prefetch metadata can be maintained within a high-performance cache memory. For example, a fixed portion of the high-performance memory resources of the cache can be allocated for storing the prefetcher metadata (and / or allocating the fixed portion to the prefetcher and / or prefetch logic of the cache). The size and / or configuration of the fixed portion can be determined during the design, production, and / or manufacturing of the cache and / or the deployment of components of the cache (such as a processor, a system-on-chip SoC, etc.). In some embodiments, the fixed portion of the cache memory allocated for the prefetch metadata can be set in hardware, a register transfer level (RTL) implementation, etc.

[0031] A fixed allocation of cache memory can improve prefetch performance, especially by reducing the latency of metadata updates, address prediction, prefetch operations, and so on. Since the size of the cache memory is limited, allocating a fixed portion of the cache memory for metadata storage may have an adverse impact on other aspects of cache performance. For example, the allocation of the fixed portion may reduce the amount of data that can be loaded into the cache, which may result in a decrease in cache performance (e.g., an increase in the miss rate, a decrease in the hit rate, an increase in the replacement rate, etc.). In some cases, the benefits of improved prefetch performance may outweigh these drawbacks. For example, when serving a suitable workload with access patterns that can be accurately predicted and / or exploited, a prefetcher can utilize the metadata maintained within the fixed allocation of a high-performance cache memory to perform accurate, low-latency prefetch operations that result in better overall cache performance, despite the reduced cache capacity.

[0032] However, in other cases, the benefits of improved prefetch performance may not outweigh the drawbacks of reduced cache capacity. For example, when serving an unsuitable workload, a fixed portion of the cache memory resources allocated to prefetcher metadata may be effectively wasted. More specifically, when serving a workload with access patterns that may not be accurately predicted and / or exploited by the prefetcher, the fixed portion of the cache memory allocated for storing prefetcher metadata may not result in useful prefetching and may thus not improve cache performance, far less than the performance compensation that occurs due to the reduced cache capacity. When serving an unsuitable workload, the fixed allocation of cache memory would be better utilized to increase the available capacity of the cache rather than store prefetcher metadata.

[0033] To address these and other issues, the amount of cache memory allocated to prefetch metadata (fixed prefetch metadata capacity) can be determined in advance. The fixed prefetch metadata capacity can be configured to provide acceptable performance under a range of different operating conditions. In some embodiments, the fixed prefetch metadata capacity can be determined through testing, experience, simulation, machine learning, and so on. Although a fixed amount of prefetch metadata capacity may result in acceptable performance under some conditions, the performance may be affected under other conditions. In addition, the cache may not be able to adapt to changes in workload conditions.

[0034] Consider, for example, a cache service with a fixed prefetch metadata capacity or otherwise operating on a workload that is primarily unsuitable (e.g., a workload that generates memory accesses that a prefetcher cannot model, predict, and / or utilize to determine accurate prefetch predictions). The fixed prefetch metadata capacity may thus fail to produce a performance improvement. In such cases, cache performance can be improved by reducing the fixed prefetch metadata capacity or removing the fixed allocation entirely.

[0035] Consider other cases where a cache with a fixed prefetch metadata capacity services a workload that is primarily suitable, such as a large number of workloads with different respective access patterns (tracked in respective prefetcher metadata), workloads with more complex access patterns, workloads with access patterns involving a greater amount of prefetcher metadata, etc. The fixed prefetch metadata capacity may be insufficient to accurately capture the access patterns of the workload, resulting in reduced prefetch accuracy and degraded cache performance. Under these conditions, cache performance can be improved by increasing the metadata capacity available to the prefetcher (and / or further reducing the available cache capacity).

[0036] Workload characteristics, such as access patterns, can vary by address region. Additionally, the characteristics of the respective workload and / or corresponding address region may change over time. The workload characteristics within a respective region of the address space may depend on many factors, including but not limited to: the program utilizing the respective region, the state of the program, the processing tasks being performed by the program, the execution phase of the program, the characteristics of the data structures accessed by the program, the manner in which the data structures are accessed, etc. A prefetcher can utilize metadata related to the workload characteristics within a respective address region to determine accurate prefetch predictions for the respective address region. Thus, the amount of prefetcher metadata required to produce accurate prefetch predictions may depend on many factors that may change over time, including but not limited to: the number of workloads (and / or corresponding address regions), the amount of metadata required to track the access patterns within a respective address region, the prefetch techniques implemented by the prefetcher within a respective address region, the complexity of the access patterns, etc. The prefetch metadata capacity required to produce accurate prefetch predictions under a first operating condition (and / or during a first time interval) may differ from the prefetch metadata capacity required to produce accurate prefetch predictions under a second operating condition (and / or during a second time interval).

[0037] To address these and other drawbacks, this document describes adaptive cache partitioning techniques that enable dynamically adjusting the amount of cache memory allocated for storing prefetcher metadata. Thus, the cache memory capacity allocated to prefetch operations can be fine-tuned to improve cache performance.

[0038] In one embodiment, a prefetcher implements stride prefetching techniques for multiple workloads, where each workload corresponds to a respective region of an address space. The prefetcher may use metadata related to accesses within the respective region to detect stride patterns for the respective workloads. Detecting stride access patterns for Y workloads may involve maintaining metadata related to accesses within Y different address regions. However, a fixed prefetcher metadata capacity may not be able to maintain metadata that can capture Y patterns, which may reduce the accuracy of prefetch predictions and thus result in reduced cache performance. For example, a fixed prefetcher metadata capacity may only be able to track a subset of the Y patterns, leaving X address regions uncovered. This may cause the prefetcher to perform inaccurate prefetching within the X address regions or prevent prefetching within the X address regions altogether. In contrast, the disclosed adaptive cache partitioning techniques can improve cache performance by, in particular, increasing the amount of cache memory allocated to the prefetcher such that the prefetcher can store metadata related to the stride patterns of each of the Y workloads and / or regions. The disclosed adaptive cache partitioning can be able to modify the prefetcher metadata capacity in response to changing workload conditions. For example, one or more of the Y workloads may transition from being suitable to being unsuitable over time, resulting in reduced prefetch performance. In response to the reduced prefetch performance, the amount of cache memory allocated for prefetcher metadata can be reduced, and the reduction in the amount can cause the available cache capacity to increase correspondingly, thereby improving overall cache performance.

[0039] In another instance, a prefetcher may implement correlation prefetching techniques that learn access patterns that may be repetitive but not as consistent as simple stride or delta-address patterns (correlation patterns). The correlation patterns may include delta sequences that contain multiple elements and may thus be derived from a larger amount of metadata than simple stride patterns. For example, depending on the degree of correlation prefetch operation, correlation prefetching for a delta sequence containing two elements (delta1, delta2) may include prefetching addresses a + delta1, a + delta1 + delta2, a + 2delta1 + delta2, a + 2delta1 + 2delta2, and so on. Since correlation prefetching techniques attempt to capture more complex patterns, these techniques may involve a larger amount of metadata. In the case of tracking correlation patterns for multiple workloads and / or regions, a cache with a fixed prefetcher metadata capacity may be insufficient, resulting in reduced performance. In contrast, the adaptive cache partitioning techniques disclosed herein can increase the amount of cache memory allocated to the prefetcher, resulting in improved prefetch accuracy and better overall performance, even though the available cache capacity is reduced correspondingly. The disclosed adaptive cache partitioning techniques can adjust the cache memory allocation in response to changing workload conditions, such as workloads with simpler single-stride access patterns, workloads with fewer workloads, etc.

[0040] The disclosed adaptive cache partitioning techniques can also improve the performance of machine learning and / or machine learning (ML) prefetching implementations (such as classification-based prefetchers, artificial neural network (NN) prefetchers, deep neural network (DNN) prefetchers, recurrent NN (RNN) prefetchers, long short-term memory (LSTM) prefetchers, etc.). For example, an LSTM prefetchers can be trained to model the "local context" of memory accesses within an address space, where each "local context" corresponds to a corresponding address range of the address space. These types of ML prefetching techniques can attempt to exploit the local context because, as disclosed herein, data structures accessed by programs running within a corresponding local context tend to be stored in contiguous data structures or blocks that are accessed repeatedly and / or in a regular pattern. An ML prefetchers can be trained to develop and / or refine an ML model within the corresponding local context, and the ML prefetchers can use the ML model to perform prefetch operations. However, due to differences in the workloads generated by programs operating in various regions of the address space, the local context can vary significantly across the address space. An ML model trained to learn the local context within one region of the address space (and / or generated by one program) may not be able to accurately model the local context within other regions of the address space (and / or generated by another program). Thus, the ML model may rely on metadata that covers the corresponding local context. A fixed allocation of cache memory may not be sufficient to maintain the ML model for the workload served by the cache, resulting in poor prefetch performance. However, the disclosed adaptive cache partitioning techniques can be able to adjust the amount of prefetch metadata capacity allocated to the prefetchers based on the number and / or complexity of the ML models tracked thereby.

[0041] The techniques described for adaptive cache partitioning can be used with caches, prefetchers, and / or other components of a hierarchical memory system. In some embodiments, logic coupled to a cache memory is configured to balance the performance improvements achieved by allocating cache memory capacity for prefetch metadata against the corresponding reduction in available cache capacity. The logic can be configured to allocate a first portion of the cache memory for metadata related to an address space (e.g., prefetch metadata), allocate cache data to a second portion of the cache memory different from the first portion, and / or modify the size of the first portion of the cache memory allocated for the metadata at least in part based on a metric related to data prefetched into the second portion of the cache memory. The metric can be configured to quantify any suitable aspect of cache and / or prefetch performance, including but not limited to: prefetch hit rate, prefetch miss rate, number of useful prefetches, number of invalid prefetches, ratio of useful prefetches to invalid prefetches, cache hit rate, cache miss rate, request latency, average request latency, etc. When the metric exceeds a first threshold, the amount of cache memory allocated for prefetcher metadata can be increased, and when the metric is below a second threshold, the amount can be decreased.

[0042] Metadata maintained within the first portion of the cache memory can be updated in response to requests related to an address space (such as read requests, write requests, transfer requests, cache hits, cache misses, prefetch hits, prefetch misses, etc.). The prefetcher (and / or prefetch logic) of the cache can be configured to select data to be prefetched into the second portion of the cache memory at least in part based on the metadata related to the address space maintained within the first portion of the cache memory. The metadata can include any suitable information related to an address and / or a range of an address space, including but not limited to: address sequences, address histories, index tables, Δ sequences, stride patterns, correlation patterns, feature vectors, ML features, ML feature vectors, ML models, ML modeling data, etc.

[0043] In response to monitoring one or more metrics related to data prefetched into the second portion of the cache memory, the size of the first portion of the cache memory allocated for metadata may be modified. The metrics may be configured to quantify prefetch performance and may include, but are not limited to, prefetch hit rate, number of useful prefetches, number of invalid prefetches, or ratio of useful prefetches to invalid prefetches, etc. When one or more of the metrics exceed a first threshold, the size of the first portion may be increased, or when the metrics are below a second threshold, the size of the first portion may be decreased. In some embodiments, when prefetch performance remains above the first threshold, the amount of cache memory allocated for storing metadata related to an address space (prefetcher metadata) may be incrementally and / or periodically increased. The amount of cache memory allocated for the metadata may be increased until a maximum value or upper limit is reached. Conversely, when prefetch performance remains below the second threshold, the amount of cache memory allocated for the metadata may be incrementally and / or periodically decreased. The amount of cache memory allocated for the metadata may be decreased until a minimum value or lower limit is reached. In some aspects, at the lower limit, no cache resources are allocated for metadata storage, and substantially all of the cache memory is available as cache capacity.

[0044] In these ways, adaptive cache partitioning provides flexible means and techniques for effectively coping with different prefetch environments. The cache memory may include a first portion allocated for metadata related to an address space and a second portion allocated for caching data of the address space. In an example embodiment, the relative sizes of the first portion and the second portion are adjusted at least in part based on the current processing workload. If the current processing workload is suitable for prefetching, the size of the first portion may be set appropriately. For example, the logic may increase the size of the first portion and decrease the size of the second portion. The logic may convert some memory storage from being used for caching data to being used for storing metadata to increase the prefetching ability of the workload suitable for prefetching. On the other hand, if the current processing workload is not suitable for prefetching, the size of the first portion may be appropriately reduced to provide more resources for caching data. In this case, the logic may decrease the size of the first portion and increase the size of the second portion. Thus, the logic may convert some cache memory storage from being used for maintaining metadata to being used for storing cached data to reduce the resources consumed by the prefetcher for the workload not suitable for prefetching. Thus, the described cache partitioning may be adapted to effectively provide more prefetching capabilities or more cache storage according to the current processing workload.

[0045] Example Operating Environment

[0046] Figure 1-1 An example device 100 that illustrates some aspects in which adaptive cache partitioning may be implemented. The device 100 may be implemented as, for example, at least one electronic device. Example electronic device implementations include Internet of Things (IoT) device 100-1, tablet computer device 100-2, smartphone 100-3, notebook computer 100-4, desktop computer 100-5, server computer 100-6, server cluster 100-7, and so on. Other device examples include: wearable devices such as smartwatches or smart glasses; entertainment devices such as set-top boxes or smart TVs; motherboards or server blades; consumer appliances; vehicles; industrial equipment; network-attached storage (NAS) devices, and so on. Each type of electronic device includes one or more components to provide some computing functionality or features.

[0047] In an example implementation, the device 100 includes at least one host 102, at least one processor 103, at least one memory controller 104, an interconnect 105, a memory 108, and at least one cache 110. The memory 108 may represent main memory, system memory, backup memory, backup storage device, combinations thereof, and so on. The memory 108 may be implemented using any suitable memory and / or storage facility, including but not limited to: memory arrays, semiconductor memories, read-only memory (ROM), random access memory (RAM), static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), thyristor random access memory (TRAM), ferroelectric RAM (FeRAM), magnetic RAM (MRAM), resistive RAM (RRAM), magnetoresistive RAM (MRAM), spin torque transfer RAM (STT RAM), phase change memory (PCM), three-dimensional (3D) stacked DRAM, double data rate (DDR) memory, high bandwidth memory (HBM), hybrid memory cube (HMC), solid state memory, flash memory, NAND flash, NOR flash, 3D XPoint TM memory, and so on. Other examples of the memory 108 are described herein. In some aspects, the host 102 may further include and / or be coupled to a non-transitory storage device, which may be implemented using a device or module that includes any suitable non-transitory, persistent, solid state, and / or non-volatile memory.

[0048] As shown, host 102 or host device 102 may include a processor 103, a memory controller 104, and / or other components (e.g., cache 110-1). The processor 103 may be coupled to cache 110-1, and cache 110-1 may be coupled to memory controller 104. The processor 103 may also be directly or indirectly coupled to memory controller 104. Host 102 may be coupled to cache 110-2 via interconnect 105. Cache 110-2 may be coupled to memory 108.

[0049] The components of the depicted device 100 represent an example computing architecture having a memory hierarchy (or hierarchical memory system). For example, cache 110-1 may be logically coupled between processor 103 and cache 110-2. Further, cache 110-2 may be logically coupled between processor 103 (and / or cache 110-1) and memory 108. In Figure 1-1 an instance, cache 110-1 is located at a higher level in the memory hierarchy compared to cache 110-2. Similarly, cache 110-2 is located at a higher level in the memory hierarchy compared to memory 108. The indicated interconnect 105 and other interconnects coupling the various components may enable data to be transferred between or among the various components. Examples of interconnects include buses, switching fabrics, one or more wires carrying voltage or current signals, and the like.

[0050] Although Figure 1-1 a particular implementation of device 100 is depicted and described herein, device 100 may be implemented in alternative ways. For example, host 102 may include additional caches, including multi-level cache memories (e.g., multiple cache levels). In some implementations, processor 103 may include one or more internal memories and / or cache levels, such as instruction registers, data registers, L1 cache, L2 cache, L3 cache, and the like. Further, at least one other cache and memory pair may be coupled "below" the shown cache 110-2 and / or memory 108. Cache 110-2 and memory 108 may be implemented in various ways. In some implementations, both cache 110-2 and memory 108 are disposed on or physically supported by a motherboard, where memory 108 includes "main memory". In other implementations, cache 110-2 includes DRAM and / or is implemented by the DRAM, and memory 108 includes a non-transitory memory device or module and / or is implemented by the non-transitory memory device or module. Nevertheless, the components may be implemented in alternative ways, including implementing them in a distributed or shared memory system. Further, a given device 100 may include more, fewer, or different components.

[0051] The cache 110-2 can be configured to improve memory performance by storing data from the relatively lower-performance memory 108 within the relatively higher-performance cache memory 120. The cache memory 120 can be provided and / or embodied by cache hardware, which can include but is not limited to: semiconductor integrated circuit systems, memory cells, memory arrays, memory banks, memory chips, etc. In some aspects, the cache memory 120 includes a memory array. The memory array can be configured as the cache memory 120 that includes a plurality of cache units (such as cache lines, etc.). The memory array can be a collection of memory cells (e.g., a grid), where each memory cell is configured to store at least one bit of digital data. The cache memory 120 (and / or its memory array) can be formed on a semiconductor substrate (such as silicon, germanium, silicon-germanium alloy, gallium arsenide, gallium nitride, etc.). In some cases, the substrate is a semiconductor wafer. In other cases, the substrate can be a silicon-on-insulator (SOI) substrate, such as silicon-on-glass (SOG) or silicon-on-sapphire (SOS) or an epitaxial semiconductor material layer on another substrate. The conductivity of the substrate or sub-regions of the substrate can be controlled by doping using various chemical species including but not limited to phosphorus, boron, or arsenic. Doping can be performed by ion implantation or by any other doping mechanism during the initial formation or growth of the substrate. However, the present disclosure is not limited to this; the cache memory 120 can include any suitable memory and / or memory mechanism, including but not limited to: memories, memory arrays, semiconductor memories, volatile memories, RAM, SRAM, DRAM, SDRAM, non-volatile memories, solid-state memories, flash memories, etc.

[0052] Data can be loaded into the cache memory 120 in response to a cache miss, such that subsequent requests for the data can be served more quickly. Additional performance improvements can be achieved by prefetching data into the cache memory 120, which can include predicting addresses that may be requested in the future and prefetching the predicted addresses into the cache memory 120. When a request related to the prefetched address is subsequently received at the cache 110-2, the relatively higher-performance cache memory 120 can serve the request without triggering a cache miss (and without accessing the relatively lower-performance memory 108).

[0053] Addresses for prefetching can be selected based in particular on metadata 122 related to the address space of the memory 108. In some embodiments, access patterns within corresponding regions of the address space can be derived from the metadata 122, and the access patterns can be used to prefetch data into the cache memory 120. In some embodiments, at least some of the metadata 122 is maintained within the cache memory 120. A portion or partition of the cache memory 120 can be allocated for storing the metadata 122. The amount of the cache memory 120 allocated for the metadata 122 can be adjusted, fine-tuned, modified, changed, and / or otherwise managed based at least in part on one or more metrics. The metrics may be related to one or more aspects of prefetch performance (which may include one or more prefetch performance metrics), such as the number of useful prefetches, the number of invalid prefetches, the ratio of useful prefetches to invalid prefetches, the prefetch hit rate, the prefetch miss rate, etc. Alternatively or additionally, the metrics may be related to one or more aspects of cache performance (which may include one or more cache performance metrics), such as the cache hit rate, the cache miss rate, etc. When one or more of the metrics exceed a first threshold, the amount of the cache memory 120 allocated for the metadata 122 can be increased, and when the metrics are below a second threshold, the amount can be decreased, and so on.

[0054] In Figure 1-1 Examples, various aspects of the adaptive cache partitioning are implemented by the cache 110-2. However, the present disclosure is not limited to this. The disclosed techniques for adaptive cache partitioning can be implemented in any cache 110 (e.g., cache 110-1) and / or in cache hierarchies (including across multiple caches 110 and / or cache hierarchies). In some examples, the cache 110-2 can be configured to allocate cache memory for metadata 122 related to the address space. Alternatively or additionally, one or more internal caches of the processor 103 can be configured to implement adaptive cache partitioning as disclosed herein (e.g., the L3 cache of the processor 103 can allocate cache memory to store metadata 122 related to the address space).

[0055] Figure 1-2Shows additional examples of devices that can implement adaptive cache partitioning. Device 100 may include a cache 110 that is configured to cache data associated with an address space. Cache 110 may be configured to cache data related to any suitable address space, including but not limited to: a memory address space, a storage address space, a host address space, an input / output (I / O) address space, a main memory address space, a physical address space, a virtual address space, a virtual memory address space, especially an address space managed by a processor 103, a memory controller 104, a memory management unit (MMU), etc. In Figure 1-2 this example, cache 110 is configured to cache data related to the address space of memory 108. Thus, memory 108 may represent the backing memory of cache 110 within the memory hierarchy.

[0056] Cache 110 may load the address and / or corresponding data of relatively slower memory 108 into relatively faster cache memory 120. Data may be loaded in response to a cache miss (e.g., in response to a request related to an address and / or data not available within cache 110). Serving a cache miss may involve transferring data from relatively slower memory 108 to relatively faster cache memory 120. Thus, cache misses may result in increased request latency and poor performance. Before a request related to an address is received, cache 110 may address these and other issues by prefetching the address into cache memory 120. Accurate prefetching can result in prefetch hits that can be served using cache memory 120 without incurring the latency associated with cache misses. However, inaccurate prefetching may consume cache memory resources as well as data that will subsequently not be accessed, which may adversely affect performance (e.g., increase the miss rate, decrease the cache hit rate, increase bandwidth consumption, etc.).

[0057] Metadata 122 related to address spaces can be used, in particular, to notify prefetch operations. In some embodiments, an address access pattern is derived from the metadata 122 and the address access pattern is accurately utilized to predict the address of an upcoming request. The metadata 122 can contain any information related to the address space and / or data associated with the backing store of the cache 110 (e.g., the memory 108), including but not limited to: a sequence of previously requested addresses or address offsets, an address history, an address history table, an index table, the access frequency of the corresponding address, an access count (e.g., accesses within a corresponding window), an access time, a last access time, ML features or parameters, ANN features or parameters (e.g., weight and / or bias parameters), DNN features or parameters, LSTM features or parameters, and so on. In some aspects, the metadata 122 contains multiple entries, each entry containing information related to a corresponding region of the address space. The metadata 122 related to a corresponding region of the address space can be used, in particular, to determine an address access pattern within the corresponding region, and the address access pattern can be used to notify prefetch operations within the corresponding region.

[0058] The metadata 122 can be performance-sensitive. The metadata 122 related to the address space can be retrieved, updated, and / or otherwise accessed during performance-sensitive operations such as operations to service requests, cache operations, prefetch operations, and so on. Therefore, it can be advantageous to maintain the metadata 122 in a high-performance memory resource. In Figure 1-2 an instance, at least some of the metadata 122 related to the address space is maintained within the cache memory 120. A portion of the cache memory 120 can be allocated for storing the metadata 122. As Figure 1-2 shown, a first portion 124 (or first partition) of the cache memory 120 is reserved for the metadata 122. Data (cache data) corresponding to addresses associated with the memory 108 can be cached within a second portion 126 (or second partition) of the cache memory 120, and the second portion can be different from and / or separate from the first portion 124. The first portion 124 can contain any suitable resources of the cache memory 120, including but not limited to zero or more of: cache units, cache blocks, cache lines, hardware cache lines, sets, ways, rows, columns, banks, and so on. The first portion 124 (or first partition) can be referred to as a metadata portion, a metadata partition, a prefetch portion, a prefetch partition, and so on. The second portion 126 (or second partition) can be referred to as a cache portion, a cache partition, and so on.

[0059] The size of the first portion 124 can be adjusted based at least in part on one or more metrics. The metrics may be related to any suitable aspect of the cache 110 and / or the memory hierarchy, including but not limited to: request latency, average request latency, throughput, cache performance, cache hit rate, cache miss rate, prefetch performance, prefetch hit rate, prefetch miss rate, number of useful prefetches, number of invalid prefetches, ratio of useful prefetches to invalid prefetches, etc. The size of the first portion 124 can increase in response to metrics indicating that the cache and / or prefetch performance meets one or more first thresholds, and can decrease in response to metrics that fail to meet one or more second thresholds. The metrics can be configured to quantify the degree to which the workload on the cache 110 is suitable for prefetching. Accordingly, the amount of cache memory 120 allocated to store metadata 122 related to prefetch operations can correspond to the degree to which the workload is suitable for prefetching (as quantified by one or more metrics). Under workload conditions suitable for prefetching, the size of the first portion 124 may increase (and the size of the second portion 126 may decrease), which can enable more accurate prefetching and further improve overall cache performance, even if the available cache capacity is reduced. Under workload conditions not suitable for prefetching, the size of the first portion 124 may decrease (and the size of the second portion 126 may increase), which can increase the available capacity of the cache 110. Under suitable workloads, the increased availability of cache capacity can enable performance improvements (e.g., reducing cache miss rates, reducing replacement rates, etc.).

[0060] Figure 1-3 Additional examples of devices that can implement adaptive cache partitioning are shown. In Figure 1-3 an example, the cache 110 can be an internal cache and / or cache layer of the processor 103 (and / or its processor cores), such as an L1 cache, an L2 cache, an L3 cache, etc. In some aspects, the memory hierarchy can further include a cache 110 (such as Figure 1-1 the cache 110-2 shown) disposed between the processor 103 and the memory 108.

[0061] The cache 110 can be configured to cache data associated with addresses of an address space. In Figure 1-3 an example, the cache 110 is configured to cache data related to a virtual address space managed by an MMU (such as a memory controller 104, an operating system, etc.). The address space can be larger than the physical address space of the memory resources of the host 102 (e.g., the address space can be larger than the physical address space of the memory 108). The address space can be a 32-bit address space, a 64-bit address space, a 128-bit address space, etc.

[0062] Cache 110 may allocate a first portion 124 of cache memory 120 to store metadata 122 related to an address space. The metadata 122 may include information related to accesses to corresponding addresses and / or address regions of the address space, and this information may be used in particular to prefetch data into a second portion 126 of the cache memory 120 (prefetch cache data 128 related to the corresponding addresses of the address space). As disclosed herein, cache 110 may adjust the amount of cache memory 120 allocated to store metadata 122 at least in part based on one or more metrics. Cache 110 may increase the amount of cache memory 120 allocated to the first portion 124 under workload conditions suitable for prefetching, and may decrease the amount allocated to the first portion 124 under workload conditions not suitable for prefetching.

[0063] Example Schemes and Devices for Adaptive Cache Partitioning

[0064] Figure 2 An example device 200 for implementing adaptive cache partitioning is shown. The device shown includes a cache 110 configured to accelerate memory storage operations related to a memory 108 (backing memory). As disclosed herein, the memory 108 may be any suitable memory and / or storage facility.

[0065] Cache 110 may include and / or be coupled to an interface 215 that may be configured to receive requests 202 related to an address space associated with the memory 108 from at least one requester 201. The requester 201 may be a host 102, a processor 103, a processor core, a client, a computing device, a communication device (e.g., a smartphone), a personal digital assistant (PDA), a tablet computer, an Internet of Things (IoT) device, a camera, a memory card reader, a digital display, a personal computer, a server computer, a data management system, a database management system (DBMS), an embedded system, a system-on-chip (SoC) device, etc. The requester 201 may include a system motherboard and / or a backplane and may include processing resources (e.g., one or more processors, microprocessors, control circuitry, etc.). The interface 215 may be configured to couple the cache 110 to one or more interconnects, such as interconnect 105, host 102, etc.

[0066] Cache 110 may be configured to service requests 202 related to the memory 108 by using high-performance memory resources such as cache memory 120. The cache memory 120 may include cache memory resources. As used herein, "cache memory resources" refers to any suitable data and / or memory storage resources. InFigure 2 In an example, the cache memory resources of cache memory 120 include a plurality of cache units 220, and each cache unit 220 is capable of storing a corresponding amount of data. The cache unit 220 may comprise and / or correspond to any suitable type and / or arrangement of memory resources, including but not limited to: units, memory cells, blocks, memory blocks, cache blocks, cache memory blocks, pages, memory pages, cache pages, cache memory pages, cache lines, hardware cache lines, sets, ways, memory arrays, rows, columns, banks, memory banks, etc. In Figure 2 the example, cache memory 120 includes X cache units 220 (cache units 220-1 to 220-X).

[0067] In some embodiments, cache 110 may be logically positioned between requester 201 and memory 108 (e.g., may be interposed between requester 201 and memory 108). In Figure 2 the example, requester 201, cache 110, and memory 108 may be communicatively coupled to interconnect 105. Cache 110 may include and / or be coupled to logic (cache logic 210) configured to receive requests 202 related to the address space associated with memory 108, in particular by monitoring, filtering, snooping, fetching, intercepting, identifying, and / or otherwise retrieving requests 202 related to the address space on interconnect 105. Cache logic 210 may be further configured to map address 204 to cache unit 220. A request 202 related to address 204 that is mapped to cache unit 220 (including valid data associated with address 204) causes a cache hit, and the cache hit may be serviced by using relatively high-performance cache memory 120. A request 202 related to address 204 that is not mapped to valid data stored in cache memory 120 causes a cache miss. Servicing a request 202 related to address 204 that causes a cache miss may involve performing a transfer operation 203 to obtain data associated with address 204 from relatively slow memory 108, which may increase the latency of request 202.

[0068] In some embodiments, cache logic 210 includes and / or is coupled to prefetch logic 230. The prefetch logic 230 may be configured to predict the address 204 of an upcoming request 202. The prefetch logic 230 may be further configured to perform a transfer operation 203 (or cause the cache logic 210 to perform the transfer operation 203) to prefetch data corresponding to the predicted address 204 into the cache memory 120. The prefetch logic 230 may cause the transfer operation 203 to be performed before the cache 110 receives a request 202 related to the predicted address 204. Thus, a subsequent request 202 related to the predicted address 204 may result in a prefetch hit that can be serviced using the relatively high-performance cache memory 120, without incurring the latency associated with servicing a cache miss (or accessing the relatively low-performance memory 108). The latency of a request 202 related to the prefetched address 204 may not include the latency involved in loading the data for the predicted address 204 into the cache 110.

[0069] The prefetch logic 230 may determine the address prediction at least in part based on metadata 122 related to the address space (e.g., prefetcher metadata). As disclosed herein, the metadata 122 may include any suitable address access characteristics, including but not limited to: a sequence of previously requested addresses or address offsets, address history, address history table, index table, access frequency of the corresponding address, access count (e.g., accesses within a corresponding window), access time, last access time, and the like. The prefetch logic 230 may be configured to maintain and / or update the metadata 122 in response to an event related to the corresponding address 204, which may include but is not limited to: data access requests, read requests, write requests, copy requests, clone requests, trim requests, erase requests, delete requests, cache misses, cache hits, and the like. The prefetch logic 230 may utilize the metadata 122 to determine an address access pattern, and may use the determined address access pattern to predict the address 204 of an upcoming request 202. In some embodiments, the metadata 122 may include multiple entries, each entry including information related to a corresponding region of the address space. The prefetch logic 230 may utilize the metadata 122 to determine an access pattern within a corresponding region of the address space, and use the determined access pattern to predict the address 204 of an upcoming request 202 within the corresponding region.

[0070] In some embodiments, at least a portion of the metadata 122 is maintained within the cache memory 120. The cache logic 210 may allocate a first portion 124 of the cache memory 120 to store the metadata 122 and / or for use by the prefetch logic 230. The cache logic 210 may maintain data (cache data 128) related to the addresses of the address space within a second portion of the cache memory 120, which may be separate and / or different from the first portion 124 of the cache memory 120. In Figure 1-3 an example, the cache logic 210 allocates M cache units 220 to store the metadata 122. The first portion 124 may include cache units 220-1 to 220-M, and the second portion 126 for storing the cache data 128 may include C cache units, where C = X - M (cache units 220-M+1 to 220-X). Although Figure 1-3 an example of adaptive cache partitioning is shown, the present disclosure is not limited to this and may be adapted to partition the cache memory 120 according to any suitable partitioning scheme. In another example, the cache logic 210 may allocate cache units 220-1 to 220-C to the second portion 126, and allocate cache units 220-C+1 to 220-X to the first portion 124. In other examples, the cache logic 210 may allocate other groupings of the cache units 220, such as sets, ways, rows, columns, banks, etc.

[0071] The cache logic 210 may be configured to implement a first mapping scheme (metadata scheme) to map the metadata 122 and / or its entries to the cache units 220 within the first portion 124. The cache logic 210 may be further configured to implement a second mapping scheme (cache or address mapping scheme 316) to map the address 204 to the cache units 220 allocated to the second portion 126. The cache logic 210 may be configured to modify the first mapping scheme and / or the second mapping scheme in response to modifying the size and / or configuration of the cache memory 120 allocated to one or more of the first portion 124 and the second portion 126.

[0072] The cache logic 210 may adjust the amount of cache memory 120 allocated for storing the metadata 122 based at least in part on one or more metrics 212. The metrics 212 may be configured to quantify the degree to which the workload on the cache 110 is suitable for prefetching. More specifically, the metrics 212 may be configured to quantify various aspects of prefetch performance (e.g., may include one or more prefetch performance metrics 212 and / or metrics 212 related to prefetch performance), such as prefetch hit rate, quantity or useful prefetches, ratio of useful prefetches to invalid prefetches, etc. Alternatively or additionally, the metrics 212 may be configured to quantify other performance characteristics, including cache performance (e.g., may include one or more cache performance metrics 212 and / or metrics 212 related to cache performance), such as cache hit rate, cache miss rate, etc.

[0073] The cache logic 210 may use the metrics 212 to determine the degree to which the workload on the cache 110 is suitable for prefetching and accordingly perform dynamic partitioning of the cache memory 120. More specifically, the cache logic 210 may adjust the amount of cache memory 120 allocated to the first portion 124 and / or the second portion 126 based at least in part on one or more of the metrics 212. The cache logic 210 may periodically monitor the metrics 212 and may determine whether to modify the size of the first portion 124 in response to the monitoring. When one or more of the metrics 212 are above a first threshold (when prefetch performance exceeds the first threshold), the cache logic 210 may increase the amount of cache memory 120 allocated to the first portion 124, and when the one or more metrics 212 are below a second threshold (when prefetch performance is below the second threshold), the cache logic may decrease the amount. The cache logic 210 may monitor the one or more metrics 212 during background operations, during idle time periods (when not actively servicing requests 202, not actively performing prefetch operations, etc.), according to a determined schedule, etc. Increasing the size of the first portion 124 may include decreasing the size of the second portion 126, while decreasing the size of the first portion 124 may include increasing the size of the second portion 126. More specifically, increasing the size of the first portion 124 may include reallocating one or more cache units 220 of the second portion 126 to the first portion 124, while decreasing the size of the first portion 124 may include reallocating one or more cache units 220 of the first portion 124 to the second portion 126. Modifying the size of the first portion 124 may include modifying the number of cache units 220 included in the first portion 124 (e.g., modifying M), which may cause a modification of the number of cache units 220 included in the second portion 126 (e.g., modifying C, where C = X - M).

[0074] The amount of cache memory 120 reconfigured to the size allocated to metadata 122 may include manipulating metadata 122 and / or cache data 128. Reducing the amount of cache memory 120 allocated to metadata 122 may include evicting portions of metadata 122. Metadata 122 may be evicted according to a policy (metadata eviction policy). The metadata eviction policy may specify that the oldest and / or least recently used entries of metadata 122 will be evicted when the size of the first portion 124 decreases. Similarly, reducing the size of the second portion 126 may include evicting cache data 128 from one or more cache units 220. Cache data 128 may be evicted according to a policy (replacement or eviction policy), which may include but is not limited to: First In First Out (FIFO), Last In First Out (LIFO), Least Recently Used (LRU), Time-Aware LRU (TLRU), Most Recently Used (MRU), Least Frequently Used (LFU), random replacement, etc.

[0075] The cache logic 210, prefetch logic 230, and / or their components and functions may include but are not limited to: circuitry, logic circuitry, control circuitry, interface circuitry, input / output (I / O) circuitry, fuse logic, analog circuitry, digital circuitry, logic gates, registers, switches, multiplexers, arithmetic logic unit (ALU), state machine, microprocessor, processor-in-memory (PIM) circuitry, etc. The cache logic 210 may be configured as a controller of the cache 110 (or cache controller). The prefetch logic 230 may be configured as a prefetcher of the cache 110 (or cache prefetcher).

[0076] Figure 3 Another example 300 of an apparatus for implementing adaptive cache partitioning is shown. As disclosed herein, the cache 110 may be configured to cache data related to an address space associated with the memory 108. In Figure 3In an example, the device includes a cache 110 coupled between a requester 201 and a memory 108. In some embodiments, the cache 110 is interposed between the requester 201 and the memory 108. The cache 110 may include and / or be coupled to a first interface 215A and / or a second interface 215B. The first interface 215A may be configured to receive a request 202 related to an address 204 of an address space via a first interconnect 105A, and may thus be referred to as a front-end interconnect. As disclosed herein, the request 202 may correspond to one or more requesters 201. The second interface 215B may be configured to couple the cache 110 (and / or cache logic 210) to a backing memory, such as the memory 108, and may thus be referred to as a back-end interface. Cache data 128 may be loaded into the cache 110 in a transfer operation 203, which is implemented by and / or via the second interface 215B. Alternatively, as Figure 2 shown, the requester 201, the cache 110, and the memory 108 may be coupled to the same interconnect. The cache memory 120 may include a plurality of cache units 220 (e.g., X cache units 220-1 to 220-X). The cache units 220 may include memory cells, memory rows, memory columns, memory pages, cache lines, hardware cache lines, cache memory units 320, cache tags 326, etc. In some embodiments, the cache units 220 are organized into a plurality of sets, each set including a plurality of ways, each way including and / or corresponding to a respective cache unit 220.

[0077] In Figure 3 an example, each cache unit 220 includes and / or is associated with a respective cache memory unit (CMU) 320 and / or a cache tag 326. The cache tag 326 may be configured to identify data stored within the CMU 320. The CMU 320 may be capable of storing cache data 128 associated with one or more addresses 204 (one or more addressable data units) of an address space. In some aspects, each CMU 320 (and each cache unit 220) is capable of storing data for U addresses 204 (or U data units). In embodiments where the address space references corresponding bytes (is byte-addressable), each CMU 320 (and the corresponding cache unit 220) may have a capacity of U bytes.

[0078] The cache unit 220 may further include and / or be associated with cache metadata 322. The cache metadata 322 of the cache unit 220 may include information related to the cache data 128 stored in the CMU 320 of the cache unit 220 (e.g., cache metadata 322-1 to 322-X related to the cache data 128 stored in the CMU 320-1 to 320-X of the cache units 220-1 to 220-X, respectively). The cache metadata 322 may include any suitable information related to the content of the cache unit 220, including but not limited to: validity information indicating whether the cache data 128 stored in the CMU 320 of the cache unit 220 is valid, a "dirty" flag indicating whether the cache data 128 has been modified since being loaded from the memory 108 (which should be written to the memory 108 before eviction), an access count, a last access time, an access frequency, a prefetch flag indicating whether the cache data 128 was loaded during a prefetch operation, and so on. In some embodiments, the cache metadata 322 of the cache unit 220 may be maintained within the CMU 320 of the cache unit 220. Alternatively, the cache metadata 322 may be maintained within a separate cache memory resource.

[0079] The cache logic 210 may implement, include, and / or be coupled to a partitioning logic 310, which may be configured to partition the cache memory 120 into a first part 124 and a second part 126 (e.g., divide the cache memory 120 into a first partition and a second partition), in particular. The cache logic 210 may be configured to map, assign an address 204, and / or otherwise associate the address with a cache unit 220 assigned to the second part 126. The cache logic 210 may associate the address 204 with the cache unit 220 according to an address-cache mapping scheme (address mapping scheme 316 or address mapping logic). In Figure 3 an example, the address mapping scheme 316 may logically divide the address 204 into a tag region (address tag 206) and an offset region 205. The offset region 205 may be defined within the least significant bit (LSB) address region. The number of bits included in the offset region 205 may correspond to the capacity of the cache unit 220 (e.g., O = log2U, where O is the number of bits included in the offset region 205 and U is the capacity of the CMU 320 of the cache unit 220 (in terms of addressable data units)). The address tag 206 may be defined within the remaining most significant bit (MSB) address region. The number of bits included in the address tag 206 (T A ) may be T A = A L - U, where A Lis the number of bits contained in address 204 (e.g., 64 bits). Although the example address 204, offset region 205, and address tag 206 are shown and described herein with reference to the big - endian format, the present disclosure is not limited to this and can be adapted to work with address 204 in any suitable format, encoding, or endianness.

[0080] Cache logic 210 can be configured to map address tag 206 to cache unit 220 using any suitable address mapping scheme 316 (or mapping logic), the address mapping scheme including but not limited to: modulo scheme (I U = T A % C, where I U is the index of cache unit 220, and address tag (T A ) 206 is mapped to said index within a set or group of C cache units 220), mapping function (e.g., according to I U = f U (T A , C), where f U is a function that maps address tag (T A ) 206 to index I U within a set of C available cache units 220), hash function (e.g., according to I U = f h (T A , C), where f h is a function that maps address tag (T A ) 206 to index I A through a hash value derived from address tag (T U ), direct mapping, fully associative mapping, set associative mapping, etc. Cache logic 210 can look up cache unit 220 for address 204 and / or determine whether cache memory 120 contains valid data corresponding to address 204 (e.g., determine whether address 204 is a cache hit or a cache miss). Cache logic 210 can look up cache unit 220 for a corresponding address 204 in particular by matching the address tag 206 of address 204 with the cache tag 326 of cache unit 220. An address 204 that matches cache tag 326 can be identified as a cache hit, while an address 204 that does not match cache tag 326 can be identified as a cache miss. In some embodiments, cache logic 210 implements a hierarchical or set - based address mapping scheme 316, in which address tag 206 is first mapped to one of a plurality of sets, and then compared with the cache tags 326 of multiple ways within the set, each way corresponding to a respective cache unit 220.

[0081] As disclosed herein, cache logic 210 may include and / or be coupled to prefetch logic 230, which may utilize metadata 122 related to an address space to predict the address of an upcoming request 202 and prefetch data corresponding to the predicted address into cache memory 120. At least a portion of the metadata 122 may be maintained within the cache memory 120. Cache logic 210 may include, implement, and / or be coupled to partitioning logic 310, which may be configured to partition the cache memory 120 into a first portion 124 and a second portion 126 (partitioning the cache memory 120). The first portion 124 may be allocated for the metadata 122 related to the address space. The partitioning logic 310 may use the remaining available capacity of the cache memory 120 (the second portion 126) as the available cache capacity. As disclosed herein, the cache logic 210 may use the second portion 126 of the cache memory 120 to maintain cache data 128.

[0082] The partitioning logic 310 may be configured to partition the cache memory 120 into a first partition including the first portion 124 of the cache memory resources of the cache memory 120 (e.g., a first number of cache units 220) and a second partition including the second portion 126 of the cache memory resources (e.g., a second number of cache units 220). The first portion 124 of the cache memory 120 may be allocated for storing the metadata 122 related to the address space, and the second portion 126 may be allocated as the available cache capacity of the cache 110 (e.g., allocated for storing cache data 128). The partitioning logic 310 may be configured to adjust the amount of cache memory resources allocated to the first portion 124 and / or the second portion 126 at least in part based on a metric 212 indicating prefetch performance and / or the degree to which the workload served by the cache 110 is suitable for prefetching. The partitioning logic 310 may be configured to allocate zero or more cache units 220 to the first portion 124 and one or more cache units 220 to the second portion 126.

[0083] In Figure 3In the illustrated example 300, cache logic 210 (and / or partitioning logic 310) allocates M of the X available cache units 220 of cache memory 120 to a first portion 124 such that C cache units 220 are allocated to a second portion 126, where C = X - M. The first portion 124 of cache memory 120 may include cache units 220-1 through 220-M, and the second portion 126 may include cache units 220-M+1 through 220-X. However, the present disclosure is not limited to this, and cache memory 120 may be partitioned and / or cache units 220 may be allocated in any suitable pattern or according to any suitable scheme or arrangement.

[0084] Cache logic 210 may implement, include, and / or be coupled to a metadata mapping scheme 314 (and / or metadata mapping logic) that may be configured to map, address, associate, reference cache units 220 allocated to the first portion 220 and / or otherwise provide access to the cache units. The metadata mapping scheme 314 may enable prefetch logic 230 (or an external prefetcher) to access metadata 122 maintained within the first portion 124 of cache memory 120. In some embodiments, the metadata mapping scheme 314 implemented by cache logic 210 (and / or partitioning logic 310) maps metadata addresses to cache units 220 (and / or offsets within the corresponding cache units 220) allocated to the first portion 124. The metadata mapping scheme 314 may define a metadata address space (M A ),M A ∈{0,…,(M·U)-I}, where U is the capacity of cache unit 220 (the capacity of CMU 320), and M is the number of cache units 220 allocated to the first portion 124. Alternatively or additionally, the metadata address space (M A ) may define a range of cache unit indices (M I ), each index corresponding to a respective one of the M cache units 220 allocated to the first portion 124, M A ∈{0,…,M-1}. Although examples of a metadata mapping scheme 314 (and / or metadata addressing and / or access scheme) are described herein, the present disclosure is not limited to this and may be adapted to provide access to cache memory 120 allocated to the first portion 124 by any suitable mechanism or technique.

[0085] Partitioning cache memory 120 into multiple portions (e.g., a first portion 124 and a second portion 126) can include mapping logic and / or a mapping scheme for configuring the portions to allocate, include, and / or incorporate designated cache memory resources of cache memory 120. As used herein, "allocating," "partitioning," or "assigning" a portion of cache memory 120 (or "allocating," "partitioning," or "assigning" cache memory resources to a portion or partition of cache memory 120) can include configuring the mapping logic and / or mapping scheme of the portion (or partition) to "include" or "reference" the cache memory resources. Configuring the mapping logic and / or mapping scheme to "include" or "reference" cache memory resources allocated to the portion or partition of cache memory 120 can include configuring the mapping logic and / or mapping scheme to reference, allocate, include, incorporate, add, map, address, associate, and / or otherwise access (or provide access to the cache memory resources) the cache memory resources. Allocating cache memory resources to a portion or partition of cache memory 120 (e.g., the first portion 124) can further include "deallocating," "removing," or "excluding" the memory resources from other partitions or portions of cache memory 120 (e.g., the second portion 126). As used herein, "deallocating," "removing," or "excluding" cache memory resources from a portion or partition of cache memory 120 can include configuring the mapping logic and / or mapping scheme of the portion (or partition) to "remove," "exclude," or "dereference" the cache memory resources. Configuring the mapping logic and / or mapping scheme to "remove," "exclude," or "dereference" cache memory resources can include configuring the mapping logic and / or the mapping scheme to remove, disable, ignore, deallocate, dereference, demap, bypass, and / or otherwise exclude the cache memory resources from the partition or the portion (e.g., prevent the cache memory resources from being accessed by and / or through the mapping logic and / or the mapping scheme).

[0086] In Figure 3In an example, the cache logic 210 (and / or partitioning logic) may be configured to allocate M cache units 220 to a first portion 124 of the cache memory 120 (e.g., cache units 220-1 to 220-M). As disclosed herein, allocating the M cache units 220 to the first portion 124 may include configuring the metadata mapping scheme 314 (and / or metadata mapping logic) to include and / or reference cache units 220-1 to 220-M. Allocating the M cache units to the first portion 124 may further include deallocating and / or excluding the M cache units 220 from a second portion 126. Deallocating or excluding cache units 220-1 to 220-M from the second portion 126 of the cache memory 120 may include configuring the address mapping scheme 316 to remove, exclude, and / or dereference cache units 220-1 to 220-M. The address mapping scheme 316 may be configured such that the address 204 (and / or address tag 206) does not map to the cache units 220 allocated to the first portion 124. In Figure 3 an example, the address mapping scheme 316 may be configured to index a subset of X available cache units 220 of the cache memory 120, the subset including C cache units 220, where C = X - M (e.g., cache units 220-M+1 to 220-X), and exclude the M cache units 220 allocated to the first portion 124 (e.g., cache units 220-1 to 220-M). In some embodiments, cache units 220 may be excluded from the address mapping scheme 316, particularly by disabling the cache tags 326 associated with the cache units 220. Thus, allocating cache units 220-1 to 220-M to the first portion 124 of the cache memory 120 may include disabling cache tags 326-1 to 326-M. In Figure 3 the cache tags 326 of the cache units 220 allocated to the first portion 124 of the cache memory 120 (and excluded from the second portion 126 and / or the address mapping scheme 316) are highlighted with cross-hatching. In the address mapping scheme 316, the address tag 206 may remain enabled and / or index the cache tags 326-M+1 to 326-X corresponding to cache units 220-M+1 to 220-X included in the second portion 126 of the cache memory 120.

[0087] In some aspects, cache logic 210 (and / or partitioning logic 310) partitions cache memory 120 according to a partitioning scheme 312. The partitioning scheme 312 can logically define how to divide cache memory resources between a first portion 124 and a second portion 126 of the cache memory 120. The partitioning scheme 312 can also logically define how to allocate cache resources between partitions. The partitioning scheme 312 can define rules, schemes, logic, and / or criteria that can be used to dynamically allocate and / or partition the cache memory 120 between the first portion 124 and the second portion 126. The partitioning scheme 312 can be further configured to specify the amount, quantity, and / or capacity of cache memory resources that will be allocated to the first portion 124 and / or the second portion 126, respectively. Adapting the partitioning scheme 312 can include modifying the amount, quantity, and / or capacity of cache memory resources allocated to the first portion 124 and / or the second portion 126. In Figure 3 an example, the partitioning scheme 312 allocates M cache units 220 to store metadata (e.g., allocates M cache units 220 to the first portion 124), and allocates X - M cache units 220 as available cache capacity (e.g., allocates the remaining X - M cache units 220 to the second portion 126).

[0088] In some embodiments, the partitioning scheme 312 configures the cache logic 210 (and / or partitioning logic 310) to allocate cache units 220 to the first portion 124 (the cache memory 120 can be partitioned according to cache units or a cache unit-based scheme). The partitioning scheme 312 can allocate cache units 220 in cache unit address or index order. In the sequential partitioning scheme 312, allocating M cache units 220 to the first portion 124 can include allocating cache units 220 - 1 to 220 - M to the first portion 124, such that cache units 220 - M + 1 to 220 - X are allocated to the second portion 126, as Figure 3As shown. Increasing the size of the first portion 124 can include sequentially allocating additional cache units 220 to the first portion 124. In a sequential scheme, increasing the amount of cache units 220 allocated to the first portion 124 from M cache units 220 to M + R cache units 220 (e.g., increasing the size of the first portion 124 by R cache units 220) can include allocating cache units 220 - M + 1 to 220 - M + R from the second portion 126 to the first portion 124. Thus, the first portion 124 can include cache units 220 - 1 to 220 - M + R, and the second portion 126 can include cache units 220 + M + R + 1 to 220 - X. Conversely, decreasing the size of the first portion 124 from M cache units 220 to M - R cache units 220 (e.g., decreasing the size of the first portion 124 by R cache units 220) can include allocating cache units 220 - M - R to 220 - M from the first portion 124 to the second portion 126. Thus, the first portion 124 can include cache units 220 - 1 to 220 - M - R, and the second portion 126 can include cache units 220 + M - R + 1 to 220 - X.

[0089] Although an example of the partitioning scheme 312 is described herein, the present disclosure is not limited thereto. In other embodiments, the partitioning scheme 312 can configure the cache logic 210 (and / or the partitioning logic 310) to allocate the cache units 220 in other modes, orders, and / or schemes. In one example, the partitioning scheme 312 can define an interleaved allocation mode, a modulo mode, a hash mode, and can allocate the cache units 220 according to the hardware structure of the cache memory 120 and / or the manner in which the cache units 220 of the cache memory 120 are organized, etc.

[0090] In some embodiments, cache memory 120 includes a plurality of sets, each set including a plurality of ways, and each way including and / or corresponding to a respective cache unit 220. Partitioning scheme 312 may allocate cache memory resources by way, set, etc. In a way-based scheme, cache logic 210 (and / or partitioning logic 310) may partition cache memory 120 by way. In one instance, cache logic 210 may allocate a first number of zero or more ways within one or more sets to first portion 124, and may allocate a second number of one or more ways within one or more sets to second portion 126. In another instance, first portion 124 includes a first number of zero or more ways within each set of cache memory 120, and second portion 126 includes a second number of one or more ways within each set. In cache memory 120 that includes N-way sets (e.g., cache units 220 where each set includes N ways), first portion 124 may include a first group of ways within each set, and second portion 126 may include a second group of ways within each set (e.g., may include ways not allocated to first portion 124). Allocating M cache units 220 to first portion 124 of cache memory 120 by way may include allocating W1 ways to the first portion within each set, where and S is the number of sets included in the cache memory. W2 ways within each set may be allocated to second portion 126, where W2 = N - W1 or and N is the number of ways within each set. First portion 124 may include ways 1 to W1 within each set, and second portion 126 may include ways W1 + 1 to N within each set. However, the present disclosure is not limited to this, and ways may be allocated between first portion 124 and second portion 126 in any suitable manner, scheme, and / or pattern.

[0091] In a way-based scheme, increasing the amount of cache memory 120 allocated to first portion 124 from M cache units 220 to M + R cache units may include allocating an additional W 1A ways to first portion 124 for each set (and deallocating W 1A ways from second portion 126 for each set), where Thus, first portion 124 may include ways 1 to W 1+ W 1A , and second portion 126 may include ways W1 + W 1A + 1 to N for each set. Conversely, decreasing the amount of cache memory 120 allocated to first portion 124 from M cache units 220 to M - R cache units may include W for each set2A The paths are allocated from the first part 124 to the second part 126, where Thus, the first part 124 may include paths 1 to W1-W within each set 2A , and the second part 126 may include paths W1-W 2A +1 to N within each set.

[0092] Alternatively or additionally, the cache logic (and / or the partitioning logic 310) may be configured to partition the cache memory 120 by sets. The allocated sets may include allocating each path (and / or the corresponding cache unit 220) of the set. In a set-based scheme, the first part 124 may include a first group of zero or more sets of the cache memory 120, and the second part 126 may include a second group of one or more sets of the sets (which may include each set of the cache memory 120 not allocated to the first part 124). Allocating M cache units 220 of the cache memory 120 to the first part 124 by sets may include allocating E1 sets to the first part 124, where and N is the number of paths included in each set, such that E2 sets are allocated to the second part 126, where E2 = S - E1, and S is the number of sets included in the cache memory 120. The first part 124 may include sets 1 to E1 of the cache memory 120, and the second part 126 may include sets E1+1 to S. However, the present disclosure is not limited to this, and the sets may be allocated between the first part 124 and the second part 126 in any suitable manner, scheme, and / or pattern.

[0093] In a set-based scheme, increasing the amount of the cache memory 120 allocated to the first part 124 from M cache units 220 to M+R cache units may include allocating an additional E 1A sets to the first part 124 (and deallocating E 1A sets from the second part 126), where Thus, the first part 124 may include sets 1 to E1+E 1A , and the second part 126 may include sets E1+E 1A +1 to S. Conversely, reducing the amount of the cache memory 120 allocated to the first part 124 from M cache units 220 to M-R cache units may include allocating E 2A sets from the first part 124 to the second part 126, where Thus, the first part 124 may include sets 1 to E1-E 2Aand the second portion 126 may include sets E1 - E 2A +1 to S.

[0094] As disclosed herein, cache logic 210 (and / or partitioning logic 310) may adjust the amount of cache memory 120 allocated to the first portion 124 (and / or the second portion 126) at least in part based on one or more metrics 212. The metrics 212 may be configured to quantify the degree to which the workload on cache 110 is suitable for prefetching. The metrics 212 may be configured to quantify various aspects of prefetch performance. Cache logic 210 (and / or prefetch logic 310) may determine and / or monitor any suitable aspects of prefetch performance, such as prefetch hit rate, prefetch miss rate, number of useful prefetches, number of invalid prefetches, ratio of useful prefetches to invalid prefetches, etc. The prefetch hit rate may be determined by tracking accesses to the prefetched cache data 128 within cache memory 120. As used herein, "prefetched" cache data 128 refers to cache data 128 that is loaded into cache memory 120 before being requested (e.g., by prefetch logic 230 and / or during a prefetch operation). In contrast, non-prefetched cache data 128 refers to cache data 128 that is loaded in response to a request 202, cache miss, etc. In some embodiments, cache logic 210 tracks prefetched cache data 128 by using cache metadata 322. Cache logic 210 may record a prefetch tag or other indicator in cache metadata 322 to distinguish prefetched cache data 128 from non-prefetched cache data 128. The prefetch hit rate may be determined based on access metrics (such as access count, access frequency, last access time, etc.) of the prefetched cache data 128 maintained in cache metadata 322. Alternatively or additionally, the prefetch miss rate may be determined by identifying prefetched cache data 128 that has no access or access below a threshold number or frequency.

[0095] In some aspects, the metrics 212 are further configured to quantify other aspects of cache performance, such as cache hit rate, cache miss rate, request latency, etc. Cache logic 210 may be configured to determine and / or monitor various aspects of cache performance. Cache logic 210 may be configured to determine the cache hit rate, in particular, by monitoring the number of requests 202 that cause cache hits, monitoring the number of requests 202 that cause cache misses, etc.

[0096] In some embodiments, one or more metrics 212 are configured to quantify aspects of cache and / or prefetch performance within corresponding regions of an address space. Cache logic 210 (and / or prefetch logic 230) can determine and / or monitor prefetch performance within the regions of the address space covered by corresponding entries of metadata 122. Accordingly, metrics 212 can quantify the degree to which the workload within a corresponding region of the address space is suitable for prefetching. Prefetch logic 230 can utilize metrics 212 to determine whether to perform prefetching within a corresponding address region, the degree of prefetching for a corresponding address region, the amount of metadata 122 maintained for a corresponding address region, and the like.

[0097] Cache logic 210 (and / or partitioning logic 310) can utilize one or more metrics 212 to perform dynamic partitioning of cache memory 120. More specifically, cache logic 210 can utilize metrics 212 to determine, fine-tune, adapt, and / or otherwise manage the amount of cache memory 120 allocated for storing metadata 122 related to the address space (the amount of cache memory 120 allocated to the first portion 124) and / or the amount of cache memory 120 allocated for storing cache data 128 (the amount of cache memory 120 allocated to the second portion 126). Cache logic 210 (and / or partitioning logic 310) can be configured to: a) increase the number of cache units 220 allocated to the first portion 124 when one or more of metrics 212 exceed a first threshold (thereby reducing the number of cache units 220 allocated for storing cache data 128 within the second portion 126); or b) reduce the number of cache units 220 allocated to the first portion 124 when one or more of metrics 212 are below a second threshold (thereby increasing the number of cache units 220 allocated for storing cache data 128 within the second portion 126). Since metrics 212 can be configured to quantify prefetch performance, the adjustments implemented by cache logic 210 can dynamically allocate cache memory resources between the first portion 124 (metadata 122) and the second portion 126 (cache data 128) based on the degree to which the workload served by cache 110 is suitable for prefetching. Cache logic 210 can increase the amount of cache memory 120 allocated for metadata 122 under workload conditions suitable for prefetching and can reduce (or eliminate) the allocation under workload conditions not suitable for prefetching, thereby increasing the amount of available cache capacity when serving an unsuitable workload.

[0098] Increasing the number of cache units 220 allocated to the first portion 124 may include assigning or allocating one or more cache units 220 of the second portion 126 to the first portion 124. As disclosed herein, allocating cache units 220 to the first portion 124 may include configuring the metadata mapping scheme 314 (and / or metadata mapping logic) to include the cache units 220, thereby providing access to the cache units 220 (or their CMUs 320) to the prefetch logic 230 and / or otherwise making the CMUs 320 of the cache units 220 available for storing metadata 122 related to the address space.

[0099] Allocating cache units 220 to the first portion 124 of the cache memory 120 may further include deallocating cache units 220 from the second portion 126. Deallocating cache units 220 from the second portion 126 of the cache memory 120 may include configuring the address mapping scheme 316 (and / or address mapping logic) to remove or exclude the cache units 220. The address mapping scheme 316 may be configured to dereference the cache units 220 such that the cache units 220 are excluded from the C cache units 220 included in the second portion 126. The address mapping scheme 316 may be modified to remove the cache units 220 from an index or other mechanism through which the address 204 and / or address tag 206 are associated with the cache units 220 (e.g., by disabling the cache tag 326 of the cache units 220). Deallocating cache units 220 from the second portion 126 may further include evicting cache data 128 from the cache units 220, setting the validity flag of the cache metadata 322 to "false", etc. In some embodiments, deallocating cache units 220 from the second portion 126 further includes identifying "dirty" cache data 128 within the cache units 220 (at least in part based on the cache metadata 322 associated with the cache data 128, such as a "dirty" indicator) and clearing and / or demoting the identified cache data 128 (if any) to a backing memory such as the memory 108 (e.g., writing the identified cache data 128 back to the memory 108).

[0100] The cache logic 210 (and / or the partitioning logic 310) can be configured to maintain the cache state when repartitioning the cache memory 120 to increase the size of the first portion 124 and / or decrease the size of the second portion 126. When reducing the amount of the cache memory 120 allocated to the second portion 126, the cache logic 210 can maintain the cache state by, in particular, compressing the cache data 128 stored within the second portion 126 to store the cache data in fewer cache units 220. Allocating R cache units 220 from the second portion 126 of the cache memory 120 to the first portion 124 can include compressing the cache data 128 stored within the second portion 126 of the cache memory 120 from a current size corresponding to C1 cache units 220 to a compressed size corresponding to C2 cache units 220, where C2 = C1 - R. Compressing the cache data 128 can include evicting a first subset of the cache data 128 currently stored within the second portion 126 of the cache memory 120, the first subset including an amount of cache data 128 equivalent to R cache units 220 (and / or the cache data 128 stored within the R cache units 220 currently allocated to the second portion 126). The first subset of the cache data 128 can be selected for eviction based on a suitable eviction or replacement policy and / or criteria (such as FIFO, LIFO, LRU, TLRU, MRU, LFU, random replacement, etc.). Evicting the cache data 128 from the cache unit 220 can make the cache unit 220 available for storing other cache data 128 (changing the cache unit 128 from "occupied" to "available" or empty). The cache logic 210 (and / or the partitioning logic 310) can be further configured to move the remaining cache data 128 stored within the cache unit 220 to be allocated to the first portion 124 (if any) to the available cache units 220 that will remain allocated to the second portion 126.

[0101] In some embodiments, increasing the size of a first portion 124 of a cache memory 120 that includes X cache units 220 from M cache units 220 to M+R cache units 220 can include: a) selecting cache units 220 of a second portion 126 for allocation to the first portion 124 (e.g., selecting R cache units 220 for reallocation); b) evicting cache data 128 from R of the C cache units 220 currently allocated to the second portion 126 (where C = M-X); c) moving the cache data 128 stored in the selected cache units 220 (if any) to available cache units 220 of the C-R cache units 220 to maintain the allocation to the second portion 126; and d) allocating the R selected cache units 220 from the second portion 126 to the first portion 124. The cache units 220 selected for eviction can be different from the cache units 220 selected for reallocation. As disclosed herein, cache logic 110 can select cache units 220 for eviction based on an eviction or replacement policy. In contrast, cache logic 110 (and / or partitioning logic 310) can select cache units 220 for reallocation from the second portion 126 to the first portion 124 (or vice versa) based on separate, independent criteria. As disclosed herein, the cache units 220 reallocated from the second portion 126 to the first portion 124 (or vice versa) can be selected according to a partitioning scheme 312. The partitioning scheme 312 can define the rules, schemes, logic, and / or other criteria used to partition (and / or dynamically allocate) cache units 220 between the first portion 124 and the second portion 126. The partitioning scheme 312 can partition the cache memory 120 in any suitable pattern or scheme, the pattern or scheme including but not limited to: a sequential scheme, a way-based scheme, a set-based scheme, etc.

[0102] Reducing the number of cache units 220 allocated to the first portion 124 can include increasing the number of cache units 220 allocated to the second portion 126 (e.g., increasing the available cache capacity). Reducing the size of the first portion 124 can include assigning or allocating one or more cache units 220 from the first portion 124 to the second portion 126. Allocating a cache unit 220 to the second portion 126 can include removing the cache unit 220 from the metadata mapping scheme 314 such that the cache unit 220 is no longer included in the group of M cache units 220 available for storing metadata 122 (e.g., modifying the metadata address scheme M A)Allocating the cache unit 220 to the second portion 126 may further include modifying the address mapping scheme 316 to reference the cache unit 220 (e.g., including the cache unit 220 in a group of C cache units 220 available for storing cache data 128). The address mapping scheme 316 may be modified to enable the address 204 and / or the address tag 206 to be mapped and / or assigned to the cache unit 220, particularly by enabling the cache tag 326 of the cache unit 220.

[0103] Reducing the number of cache units 220 allocated to the first portion 124 may reduce the amount of cache memory 120 available for storing metadata 122. Thus, reducing the size of the first portion 124 may include compressing the metadata 122 to store it in a smaller amount of cache memory 120. The metadata 122 may be compressed to be stored within a smaller memory range (e.g., from a first size M1 to a smaller second size M2). Compressing the metadata 122 may include removing a portion of the metadata 122, such as one or more entries of the metadata 122. The portion of the metadata 122 may be selected based on a removal criterion (such as a stage criterion (oldest first removal, newest first removal, etc.), a least recently used criterion, a least frequently used criterion, etc.).

[0104] Alternatively or additionally, individual portions of the metadata 122 may be selected for removal at least in part based on one or more metrics 212. The metadata 122 may include multiple entries, each entry including access information related to a corresponding region of the address space. The prefetch logic 230 may utilize the corresponding entries of the metadata 122 to perform prefetch operations within the address regions covered by the corresponding entries. The one or more metrics 212 may be configured to quantify the prefetch performance within the address regions covered by the corresponding entries of the metadata 122. Compressing the metadata 122 may include selecting entries of the metadata 122 for removal at least in part based on the prefetch performance, as quantified by the metrics 212, within the address regions covered by the entries. In some embodiments, entries of the metadata 122 with prefetch performance below a threshold may be removed (and / or the amount of memory capacity allocated to the entries may be reduced). Alternatively, entries of the metadata 122 that exhibit higher prefetch performance may be retained, while entries that exhibit lower prefetch performance may be removed (e.g., the R entries of the metadata 122 with the lowest performance may be selected for removal). Thus, compressing the metadata may include removing the metadata 122 from one or more cache units 220 and / or moving the metadata 122 (and / or entries of the metadata 122) from cache units 220 reallocated to the second portion 126 to the remaining cache units 220 allocated to the first portion 124.

[0105] Figure 4-1 shows another example 400 of a device for implementing adaptive cache partitioning. Device 400 includes a cache 110 configured to cache data related to an address space associated with a memory 108. Cache 110 may include and / or be coupled to one or more interconnects. In Figure 4-1 an example, cache 110 includes and / or is coupled to a first interface 215A configured to couple cache 110 to a first interconnect 105A and a second interface 215B configured to couple cache 110 to a second interconnect 105B. Cache 110 may be configured to service requests 202 from one or more requesters 201 related to an address 204 of the address space. Cache 110 may service requests 202 by using a cache memory 120, and the requests may include loading data associated with an address 204 of the address space in a transfer operation 203. The transfer operation may be implemented in response to a cache miss, a prefetch operation, etc. Cache 110 may include and / or be coupled to an interface 215 that may be configured to couple cache 110 (and / or cache logic 210) to one or more interconnects, such as interconnects 105A and / or 105B.

[0106] In Figure 4-1 an example, cache memory 120 includes a plurality of cache units 220, and each cache unit 220 includes and / or corresponds to a respective cache line. Cache units 220 may be arranged in a plurality of sets 430 (e.g., sets 430-1 to 430-S). Sets 430 may be N-way associative; each set 430 may include N ways 420, and each way 420 includes and / or corresponds to a respective cache unit 220 (respective cache line). As shown, each set 430 may include N ways 420-1 to 420-N, and each way 420 includes and / or corresponds to a respective cache unit 220. Cache memory 120 may include X cache units 220 (or X cache lines), where X = S·N.

[0107] An address mapping scheme 316 (or address mapping logic) implemented by cache logic 210 may be configured to partition an address 204 into an offset 205, a set region (set tag 406), and an address tag 206. As disclosed herein, offset 205 may correspond to the capacity of cache unit 220 (e.g., the capacity of CMU 320). Address mapping scheme 316 may utilize set tag 406 to associate address 204 with a corresponding set 430. In some aspects, address mapping scheme 316 includes a set mapping scheme by which set tag 406 is mapped to a set of available sets (SC )、S C One set from {420-1, ……, 420-S}, as follows S I = f S (T S , S C ), where f s is a set mapping function, T S is a set tag 406, and S I is an index or other identifier of the selected set 430. The address mapping scheme 316 can further include a way mapping scheme, by which the address tag 206 is mapped to one of the N ways 420 of the selected set 430 (e.g., by comparing the address tag 206 with the cache tag 326 of the way 420).

[0108] The cache logic 210 can include, implement, and / or be coupled to the partitioning logic 310, which can be configured to partition the cache memory 120 into a first part 124 and a second part 126. As disclosed herein, the first part 124 can be allocated for storing metadata 122 related to the address space, and the second part 126 can be allocated for storing cache data 128 (which can be allocated as the available cache capacity). The cache logic 210 can allocate the cache memory 120 between the first part 124 and the second part 126 according to the partitioning scheme 312. The partitioning scheme 312 can specify the amount of the cache memory 120 (the first part 124) to be allocated for the metadata 122. The partitioning scheme 312 can also specify the way of allocating the cache units 220 to the first part 124 and / or the second part 126. In Figure 4-1 an example, the cache logic 210 is configured to partition the cache memory 120 by way 420 (a way-based or way partitioning scheme 312-1 can be implemented). The way partitioning scheme 312-1 can specify that zero or more ways 420 within zero or more sets 430 of the cache memory 120 are allocated for the first part 124. In some embodiments, the way partitioning scheme 312-1 specifies that zero or more ways 420 within each set 430 of the cache memory 120 are allocated for the first part 124.

[0109] In some embodiments, cache memory 120 includes multiple banks (e.g., SRAM banks). The ways 420 of cache memory 120 can be organized within the respective banks. More specifically, the ways 420 of each set 430 can be split across multiple banks of cache memory 120. In some instances, each way 420 can be implemented by a respective one of the banks: the ways 420-1 of each set 430-1 to 430-S can be implemented by the first bank, the ways 420-2 of each set 430-1 to 430-S can be implemented by the second bank, and so on, where the ways 420-N of each set 430-1 to 430-S are implemented by the Nth bank of cache memory 120. The banks of cache memory 120 can include individual memory blocks. Thus, the first portion 124 allocated for metadata 122 can include zero or more banks (or blocks) of cache memory 120. The metadata mapping scheme 314 can address the banks allocated to the first portion 124 as a linear (or flat) memory bank block. Thus, the metadata mapping scheme 314 can enable the metadata 122 to be arranged and / or organized in any suitable manner (e.g., as specified by a prefetcher, prefetcher logic 230, etc.).

[0110] In Figure 4-1 an instance, the partitioning scheme 312-1 allocates two ways 420 within each set 430 to the prefetcher logic 230. Thus, the first portion 124 of cache memory 120 allocated for storing metadata 122 can include the ways 420-1 and 420-2 within each set 430-1 to 430-S (which can include M cache units 220, where M = 2·S). The second portion 126 of cache memory 120 available for storing cache data 128 can include N-2 ways within each set 430-1 to 430-S (which can include C cache units 220, where M = (N-2)·S). In Figures 4-1 to 4-3 the figure, the ways 420 (and / or cache units 220) allocated to the first portion 124 are shown hatched to distinguish them from the ways 420 allocated to the second portion 126. As shown, the first portion 124 can include the first portions 124-1 to 124-S within each set 430-1 to 430-S of cache memory 120, and the second portion 126 can include the second portions 126-1 to 126-S within the sets 430-1 to 430-S.

[0111] Allocating cache units 220 (or ways 420) to the first portion 124 can include configuring the address mapping scheme 316 to disable or ignore the cache units 220. In Figure 4-1In this case, two ways 420 of each set 430 are assigned to the first portion 124, and the address mapping scheme 316 is adapted to disable or ignore ways 420-1 and 420-2 of each set 430 (e.g., by disabling cache tags 326-1 and 326-2 of the corresponding cache units 220-1 and 220-2). In Figure 4-1 an example, the number of sets 430 available for storing cache data 128 (S C ) can remain substantially unchanged. As shown, the second portion 126 of the cache memory 120 can include S sets 430, each set including N-2 ways 420 (ways 420-3 to 420-N). Thus, the address mapping scheme 316 can distribute addresses 204 among the S sets 430 of the cache memory 120 (e.g., via set tags 406, etc.). The way mapping scheme implemented by the cache logic 210 (and / or the address mapping scheme 316) can be adapted to modify the associativity of the sets 430. In Figure 4-1 an example, the address mapping scheme 316 manages the sets 430 as [N-2]-way associative instead of N-way associative. More specifically, the address mapping scheme 316 maps N-2 address tags 206 to the corresponding sets 430 instead of N address tags 206.

[0112] As disclosed herein, assigning ways 420 to the first portion 124 can include evicting cache data 128 from the ways 420. As disclosed herein, cache data 128 can be selected for eviction from the corresponding sets 430 according to an eviction or replacement policy. Assigning R ways 420 of a set 430 to the first portion 124 can include compressing the cache data 128 stored within the set 430 from the capacity of N cache units 220 to the capacity of N-R cache units 220. The cache logic 210 can select cache data 128 to be retained within the corresponding set 430 and move the selected cache data 128 to the N-R ways 429 of the corresponding set 430 that will remain assigned to the second portion 126. In Figure 4-1 an example, assigning ways 420-1 and 420-2 of each set 430 to the first portion 124 can include, in particular, evicting cache data 128 from a first group of two ways 420 of the set 430 (such that a second group of N-2 ways 420 of the set 430 is retained), moving the cache data 128 stored within the second group of ways 420 to ways 420-3 to 430-N of the set 430 (if necessary), and assigning ways 420-1 and 420-2 to the first portion 124 to compress the cache data 128 within each set 430 from the capacity of N cache units 220 to the capacity of N-2 cache units 220.

[0113] Metadata 122 maintained within the first portion 124 of the cache memory 120 can be accessed according to a metadata mapping scheme 314. In Figure 4-1 an example, the metadata mapping scheme 314 can define a metadata address space (M A ), which includes ways 420-1 and 420-2 of each set 430-1 to 430-S. For example, the metadata address space (M A ) can include an address range {0, …, (R·U·S)-1}, where R is the number of ways 420 allocated to the first portion 124 within each of the S sets 430, and U is the capacity of each way 420 (in terms of addressable data units). Alternatively, the metadata address space (M A ) can define an address corresponding to an index and / or offset of the respective ways 420 (or cache units 220) of the first portion 124, as {0, …, (R·S)-1}.

[0114] The cache logic 210 (and / or prefetch logic 230) can be configured to determine and / or monitor one or more metrics 212. As disclosed herein, the metrics 212 can be configured to quantify cache and / or prefetch performance. The cache logic 210 (and / or partitioning logic 310) can adjust the way partitioning scheme 312-1 at least in part based on one or more of the metrics 212. The cache logic 210 can adjust the way partitioning scheme 312-1 to: increase the size of the first portion 124 (and decrease the size of the second portion 126) when one or more of the metrics 212 exceed a first threshold, or decrease the size of the first portion 124 (and increase the size of the second portion 126) when one or more of the metrics 212 are below a second threshold. Increasing the size of the first portion 124 can include increasing the number of ways 420 of the first portion 124 allocated to each set 430 of the cache memory 120. Decreasing the size of the first portion 124 can include decreasing the number of ways 420 of the first portion 124 allocated to each set 430 of the cache memory 120.

[0115] Figure 4-2Illustrates Example 401, in which the number of ways 420 allocated to store metadata 122 related to the address space is increased (e.g., from two ways 420 within each set 430 to three ways 420 within each set 430). The amount of cache memory 120 allocated to the first portion 124 can be increased in response to determining and / or monitoring a metric 212 (e.g., in response to prefetch performance quantified by a metric 212 exceeding a first threshold). As shown, the first portion 124 allocated to store metadata 122 can include ways 420-1 to 420-3 of each set 430-1 to 430-S. The capacity available to store metadata 122 can be increased to M = 3·S·U, where U is the capacity of cache unit 220 (or CMU 320) in terms of addressable data units. Allocating way 420-3 to the first portion 124 can include modifying the metadata mapping scheme 314 to reference way 420-3 within each set 430 (e.g., defining a metadata address scheme that includes addresses 0 to (3·S·U)-1, way indices 0 to (3·S)-1, etc.).

[0116] As Figure 4-2 shown, increasing the size of the first portion 124 can cause a decrease in the amount of cache memory 120 allocated to store cache data 128 (decreasing the size of the second portion 126). As disclosed herein, allocating way 420-3 of each set 430 to the first portion 124 can include compressing the cache data 128 stored within each set 430 into N-3 ways 420 (by selecting cache data 128 within each set 430 for eviction and moving the data to remain within each [N-3] associated set 430 of ways 420-3 to 420-N). Allocating way 420-3 can further include modifying the address mapping scheme 316 to associate address 204 with an [N-3] way associated set 430, rather than with an [N-2] or N way associated set 430. Allocating way 420-3 of each set 430 to the first portion 124 can include disabling the cache tag 326-3 of way 420-3 within each set 430.

[0117] Figure 4-3Shows another example 402 in which the number of ways 420 allocated to store metadata 122 related to the address space is reduced (e.g., reduced to one way 420 within each set 430). The amount of cache memory 120 allocated to the first portion 124 can be reduced in response to determining and / or monitoring one or more metrics 212 (e.g., in response to the prefetch performance quantified by the metric 212 being below a second threshold). As shown, the first portion 124 allocated to store the metadata 122 can include a single way 420-1 within each set 430-1 to 430-S. The capacity available for storing the metadata 122 can be reduced to M = S·U, where U is the capacity of a cache unit 220 (or CMU 320) in terms of addressable data units (or S cache units 220). As disclosed herein, reducing the size of the first portion 124 can include compressing the metadata 122 to store it within a reduced number of ways 420 (e.g., by portions of the metadata 122, one or more entries of the metadata 122, etc.). Reducing the size of the first portion 124 can further include modifying the metadata mapping scheme 314 to reference a reduced number of ways 420 allocated to the first portion 124. In Figure 4-3 , the metadata mapping scheme 314 can reference the way 420-1 within each set 430 and / or define a metadata address scheme, including addresses 0 to (S·U)-1, way indices 0 to (S)-1, etc.

[0118] As Figure 4-3 shown, reducing the amount of cache memory 120 allocated to the first portion 124 can cause an increase in the amount of cache memory 120 allocated to store cache data 128 within the second portion 126. In Figure 4-3 the example, ways 420-2 and 420-3 of each set are allocated to the second portion 126. Allocating ways 420-2 and 420-3 to the second portion 126 can include modifying the address mapping scheme 316 to include ways 420-2 and 420-3 of each set 430 (e.g., by enabling cache tags 326-2 and 326-3 for ways 420-2 and 420-3).

[0119] Figure 5-1Shows another example 500 of a device for implementing adaptive cache partitioning. Device 500 includes a cache 110 configured to cache data related to an address space associated with a memory 108. Cache 110 may include and / or be coupled to an interface 215 that may be configured to couple cache 110 (and / or cache logic 210) to an interconnect, such as interconnect 105 of host 102. Cache 110 may be configured to service requests 202 from a requester 201 related to an address 204 of the address space. Cache 110 may service request 202 by using a cache memory 120, and the request may include loading data associated with address 204 of the address space in a corresponding transfer operation 203. The transfer operation may be implemented in response to a cache miss, a prefetch operation, etc. Memory 108, cache 110, and requester 201 may be communicatively coupled via interconnect 105.

[0120] As shown, cache memory 120 may include a plurality of cache units 220 that may be organized into a plurality of N-way set associative sets 430 (e.g., S sets 430-1 to 430-S, each set including N ways 420-1 to 420-N). Cache logic 210 may implement, include, and / or be coupled to a partitioning logic 310 configured to partition cache memory 120 into a first portion 124 and a second portion 126. Cache logic 210 partitions and / or divides cache memory 120 according to a partitioning scheme 312-1 that may specify the amount of cache memory 120 to be allocated for storing metadata 122 (first portion 124), cache data 128 (second portion 126), etc. As disclosed herein, the amount of cache memory 120 allocated to the first portion 124 may be at least partially based on one or more metrics 212 that particularly quantify prefetch performance.

[0121] In Figure 5-1 an example, cache logic 210 (and / or partitioning logic 310) partitions cache memory 120 by sets and / or according to a set or set-based partitioning scheme 312-2. More specifically, cache logic 210 may allocate zero or more of the sets 430 of cache memory 120 to store metadata 122 (first portion 124) related to the address space, and allocate one or more of the sets 430 to store cache data 128 (second portion 126). In Figure 5-1In, a first portion 124 of the cache memory 120 allocated to the prefetch logic 230 includes two sets 430 (e.g., set 430-1 and 430-2), and a second portion 126 of the cache memory 120 allocated as available cache capacity includes S - 2 sets 430 (e.g., sets 430-3 to 430-S). In Figure 5-1 the sets 430 allocated to the first portion 124 are highlighted with a cross-hatched fill pattern.

[0122] As disclosed above, the address mapping scheme 316 (or address mapping logic) implemented by the cache logic 210 can be configured to map the address 204 to the cache unit 220, particularly by associating the address 204 with the corresponding set 430 and matching the address tag 206 of the address 204 to the cache tag 326 of the associated set 430. However, allocating one or more sets 430 to the first portion 124 can reduce the number of sets 430 included in the second portion 126 (reduce the number of sets 430 to which the address 204 can be mapped). Allocating R sets 430 to store metadata can reduce the number of available sets 430 to S - R (or Figure 5-1 S - 2 in the example). Allocating R sets to the first portion 124 can include modifying the address mapping scheme 316 (and / or its set mapping scheme) to distribute the address 204 among a group of S - R sets 430 (S C ), S C {420 - R,..., 420 - S} or {420 - 1,..., 420 - [S - R]}, as follows S I = f S (T S , S C ) where f s is the set mapping function, T S is the set tag 406, and S I is the index or other identifier of the selected set among the available sets 430 (S C ). In some embodiments, the address mapping scheme 316 modifies the way the address 204 is partitioned (and / or the size of the corresponding address region). The address mapping scheme 316 can adjust the number of bits included in the set tag 406-1 at least in part based on the number of sets 430 allocated to the first portion 124. The address mapping scheme 316 can, for example, reduce the number of bits included in the set tag 406-1 by log2R, where R is the number of sets 430 allocated to store the metadata 122 (reduce Figure 5-1 one bit in the example).

[0123] The metadata mapping scheme 314 can be configured to associate metadata 122 (and / or metadata addresses) with the cache memory 120 allocated to the first portion 124. The metadata mapping scheme 314 can define a series of metadata addresses from 0 to (R·N·U) - 1 or indices from 0 to R·N, where R is the number of sets 430 allocated for the metadata 122, N is the number of ways 420 included in each set 430, and U is the capacity of each way 420 (and / or the corresponding cache unit 220).

[0124] The cache logic 210 can be further configured to adjust the set partitioning scheme 312-2 at least in part based on one or more metrics 212 related to prefetch performance. When one or more of the metrics 212 exceed a first threshold, the cache logic 210 can increase the number of sets 430 allocated for the metadata 122, and when one or more of the metrics 212 are below a second threshold, the cache logic can reduce the number of sets 430 allocated for the metadata 122 (and increase the number of sets 430 available for storing cache data 128).

[0125] Figure 5-2 An example 501 is shown, in which the amount of cache memory 120 allocated for the metadata 122 is Figure 5-1 increased compared to the example 500 shown. The size of the first portion 124 can be increased based on prefetch performance within one or more regions of the address space. In Figure 5-2 , the number of sets 430 included in the first portion 124 can be increased to four (e.g., from sets 430-1 to 430-2 to sets 430-1 to 430-4). Allocating additional sets 430-3 and 430-4 to store metadata can include adjusting the address mapping scheme 316 to allocate addresses among S - 4 sets 430 (as opposed to S - 2 or S sets 430). The address mapping scheme 316 can be modified to reduce the number of bits included in the set tag 406-2 by two bits (or by a single bit compared to Figure 5-1 the set tag 406-1 of the example). Allocating sets 430 to the first portion 124 can include evicting cache data from the sets 430, disabling cache tags 326-1 to 326-N of each way 420 of the sets 430, and so on. Allocating additional sets 430 to the first portion 124 can further include adjusting the metadata mapping scheme 314 to include the additional sets 430. In Figure 5-2 the example, the metadata mapping scheme 314 can be adapted to define a series of metadata addresses from 0 to (4·N·U) - 1 or indices from 0 to 4·N.

[0126] Figure 5-3shows Instance 502, in which the amount of cache memory 120 allocated for metadata 122 is reduced compared to Figure 5-2 Instance 501 (and Figure 5-1 Instance 500). As disclosed herein, the size of the first portion 124 can be reduced based on the prefetch performance within one or more regions of the address space. In Figure 5-3 , the number of sets 430 included in the first portion 124 can be reduced to one (e.g., reduced to a single set 430-1). Thus, reducing the size of the first portion 124 can include allocating additional sets 430-4 to 430-2 to store cache data 128 (allocated to the second portion 126). Allocating one or more sets 430 to the second portion 126 can include compressing the metadata 122 and storing the compressed metadata in a reduced number of cache units 220. In Figure 5-3 the instance, the metadata 122 can be compressed to be stored in N cache units 220. Compressing the metadata 122 can include removing portions of the metadata 122, such as one or more metadata entries. The entries can be selected based on any suitable criteria, including but not limited to: phase criteria (oldest first removed, newest first removed, etc.), least recently used criteria, least frequently used criteria, prefetch performance criteria (e.g., prefetch performance within the address region covered by the corresponding entry of the metadata 122), etc. The metadata mapping scheme 314 can be modified to reduce the number of cache units 220 referenced thereby (reducing the metadata address range to N·U or reducing the metadata index range to N), etc. The address mapping scheme 316 can be modified to increase the number of available sets 430 to S-1. The address mapping scheme 316 can be configured to distribute the address 204 among a larger number of sets 430, particularly by increasing the number of bits in the set tag 406-2 included in the address 204. Allocating sets 430 to store cache data can further include enabling the cache tags 326 for each way 420 of the sets 430 (e.g., enabling the cache tags 326-1 to 326-N for each way 420-1 to 420-N of the sets 430 allocated to store cache data 128).

[0127] Example Method for Adaptive Cache Partitioning

[0128] This section will describe the example method with reference to the Figures 6 to 9 flowchart / flow diagram. These descriptions refer only by way of example to the Figures 1-1 to 5-3 components, entities, and other aspects depicted therein. Figure 6A flow chart 600 illustrates an example method for implementing adaptive cache partitioning for a device. The flow chart 600 includes blocks 602 through 606. In some embodiments, the host device 102 (and / or its components) may perform one or more operations of the flow chart 600 (and / or the operations of other flow charts described herein) to implement at least one method for adaptive cache partitioning. Alternatively or additionally, one or more of the operations may be performed by a memory, a memory controller, PIM logic, cache 110, cache memory 120, cache logic 210, prefetch logic 230, an embedded processor, etc.

[0129] At 602, a first portion of the cache memory 120 of the cache 110 is allocated to store metadata 122 related to an address space that is associated with the backing memory of the cache 110 (e.g., the address space is associated with the memory 108). For example, allocating the first portion may include partitioning the cache memory 120 into a first portion 124 and a second portion 126. The first portion 124 may be allocated to store the metadata 122, and the second portion 126 may be allocated to store cache data 128 (which may be allocated as available cache capacity). The metadata 122 maintained within the first portion of the cache memory 120 may include information related to accesses to corresponding addresses and / or regions of the address space. The prefetcher and / or prefetch logic 230 of the cache 110 may utilize the metadata 122 to predict the address 204 of an upcoming request 202 and prefetch data associated with the predicted address 204 into the second portion 126 of the cache memory 120. The metadata 122 may include any suitable information related to the address space, including but not limited to: a sequence of previously requested addresses 204 or address offsets, an address history, an address history table, an index table, the access frequency of the corresponding address 204, an access count (e.g., accesses within a corresponding window), an access time, a last access time, etc. In some aspects, the metadata 122 includes multiple entries, each entry including information related to a corresponding region of the address space. The metadata 122 related to a corresponding region of the address space may in particular be used to determine the address access pattern within the corresponding region, which may be used to inform prefetch operations within the corresponding region.

[0130] The cache memory 120 can be partitioned into a first portion 124 (e.g., a first partition) and a second portion 126 (e.g., a second partition) according to any suitable partitioning scheme 312 (such as a sequential scheme, a way-based partitioning scheme 312-1, a set-based partitioning scheme 312-2, etc.). The first portion 124 can include any suitable portion, quantity, and / or amount of the cache memory resources of the cache memory 120, including but not limited to zero or more: cache units 220, CMUs 320, cache blocks, cache lines, hardware cache lines, ways 420 (and / or corresponding cache units 220), sets 430, rows, columns, banks, etc. In some embodiments, the cache logic 210 allocates M cache units 220 to the first portion 124 and allocates X - M cache units 220 to the second portion 126 as available cache capacity (where X is the number of available cache units 220 included in the cache memory 120). Allocating M cache units 220 can include: allocating cache units 220-1 to 220-M to the first portion 124 (e.g., according to a sequential scheme); allocating cache units 220 within way W1 of each set 430 of the cache memory 120, where and S is the set 430 included in the cache memory 120 (e.g., according to the way-based partitioning scheme 312-1); allocating cache units 220 within sets 1 to E1, where and N is the number of ways 420 included in each set 430 of the cache memory 120 (e.g., according to the set-based partitioning scheme 312-2), etc.

[0131] Allocating M cache units 220 can include purging and / or demoting cache data 128 from the M cache units 220, and the purging and / or demoting can include writing dirty cache data 128 stored within the cache units 220 to the memory 108, etc. Allocating M cache units 220 can further include configuring an address mapping scheme 316 through which the address 204 is mapped to the corresponding cache units 220, sets 430, and / or ways 420 to disable, remove, and / or ignore the M cache units 220 such that the address 204 is not mapped to the M cache units 220 (and the M cache units 220 are not available for storing cache data 128). In some embodiments, at 602, the cache logic 210 disables the cache tags 326 of the M cache units 220 allocated to the first portion.

[0132] In one example, cache logic 210 partitions cache memory 120 by way 420 (e.g., by allocating ways 420 within corresponding sets 430 of cache memory 120). In way partitioning scheme 312-1, allocating M cache units 220 to the first portion 124 can include allocating W1 ways 420 within each of S sets 430-1 to 430-S of cache memory 120 to the first portion 124, where W2 ways 420 within each set 430 are allocated to the second portion 126, where W2 = S - W1 or ).

[0133] In another example, cache logic 210 can implement a set-based partitioning scheme 312-2, by which cache memory 120 is partitioned by sets 430. Allocating M cache units 220 to the first portion 124 according to set-based partitioning scheme 312-2 can include allocating E1 sets 430 to the first portion 124, where and N is the number of cache units 220 (ways 420) included in each set 430, such that E2 sets 430 are allocated to the second portion 126, where E2 = S - E1 or

[0134] Allocating M cache units 220 to store metadata can further include configuring metadata mapping scheme 314 to provide access to the memory storage capacity of the M cache units 220. Metadata mapping scheme 314 implemented by cache logic 210 can provide access to the memory storage capacity of the M cache units 220 included in the first portion 124 of cache memory 120. Metadata mapping scheme 314 can define a metadata address space (M A ), M A ∈ {0, …, (M·U)-1}, where U is the capacity of cache unit 220 (the capacity of CMU 320). Alternatively or additionally, the metadata address space (M A ) can define the range of cache unit indices (M I ), each index corresponding to a respective one of the M cache units 220 allocated to the first portion 124, M A ∈ {0, …, M-1}. Although examples of metadata mapping scheme 314 (and / or metadata addressing and / or access scheme) are described herein, the present disclosure is not limited thereto and can be adapted to provide access to cache memory 120 allocated to the first portion 124 by any suitable mechanism or technique.

[0135] At 604, data associated with the address space is written to the second portion 126 of the cache memory 120. For example, cache logic 210 may load data into the cache memory 120 in response to a request 202 related to an address 204 that triggered a cache miss (e.g., an address 204 that has not been loaded into the second portion 126 of the cache memory 120). Alternatively or additionally, at 604, cache logic 210 (and / or prefetch logic 230) may prefetch cache data 128 into the second portion 126 of the cache memory 120. Prefetcher logic 230 may utilize metadata 122 related to the address space to predict the address 204 of an upcoming request 202 and configure cache logic 210 to prefetch cache data 128 corresponding to the predicted address 204 before a request 202 related to the predicted address 204 is received. In a transfer operation 203, the prefetched cache data 128 may be transferred from a relatively slow memory 108 to a relatively fast cache memory 120. The transfer operation 203 for prefetching cache data 128 may be implemented as a background operation (e.g., during an idle period when the cache 110 is not servicing a request 202).

[0136] In some aspects, at 604, cache logic 210 (and / or prefetch logic 230) may be further configured to determine and / or monitor one or more metrics 212 related to the cache 110. The metrics 212 may be configured to quantify any suitable aspect of cache and / or prefetch performance, including but not limited to: request latency, average request latency, cache performance, cache hit rate, cache miss rate, prefetch performance, prefetch hit rate, prefetch miss rate, number of useful prefetches, number of invalid prefetches, ratio of useful prefetches to invalid prefetches, etc.

[0137] In some embodiments, cache logic 210 (and / or prefetch logic 230) may be further configured to record, update, and / or otherwise maintain metadata 122 related to an address space within the first portion of the cache memory 120 allocated at 604 (e.g., within the first portion of the cache memory 120). As disclosed herein, the metadata 122 may be accessed by and / or through a metadata mapping scheme 314.

[0138] At 606, the size of the first portion of the cache memory 120 allocated to metadata 122 associated with an address space is modified based at least in part on one or more metrics 212 related to cache data 128 prefetched into a second portion of the cache memory. When one or more of the metrics 212 exceed a first threshold, the amount of cache memory 120 allocated to the first portion 124 can be increased. At 606, the size of the first portion 124 can be incrementally and / or periodically increased while the prefetch performance remains above the first threshold and / or until a maximum value or upper limit is reached. Conversely, at 606, when one or more of the metrics 212 are below a second threshold, the amount of cache memory allocated to the first portion 124 can be decreased. At 606, the size of the first portion 124 can be incrementally and / or periodically decreased while the prefetch performance remains below the second threshold and / or until a lower limit is reached. In some aspects, at the lower limit, no cache resources are allocated for storing the metadata 122, and substantially all of the cache memory 120 is available as cache capacity.

[0139] At 606, the amount of cache memory 120 allocated for the metadata 122 can be increased when the workload on the cache 110 is suitable for prefetching and can be decreased when the workload is not suitable for prefetching (as indicated by one or more metrics 212). Thus, the cache 110 can be made to adapt to different workload conditions. For example, when servicing a workload suitable for prefetching, increasing the amount of cache memory 120 allocated to prefetch the metadata 122 can improve performance, even though the available cache capacity is reduced, while decreasing the amount of cache memory 120 allocated for the metadata 122 can make the available capacity of the cache 110 increase, thus improving performance under a workload not suitable for prefetching.

[0140] In some embodiments, modifying the size of the first portion 124 of the cache memory 120 allocated to metadata 122 may include completing outstanding requests 202 (e.g., flushing the pipeline of cache 110), clearing cache 110, resetting prefetch logic 230 (and / or the prefetcher), repartitioning cache memory 120 to modify the amount of cache memory 120 allocated to the first portion 124 and / or the second portion 126, and resuming operations using the resized cache memory 120 (e.g., using the first portion 124 and / or the second portion 126 of the resized cache memory 120). Alternatively, modifying the size of the first portion 124 of cache memory 120 may include maintaining cache and / or prefetcher state. When increasing the amount of cache memory 120 allocated to the first portion 124, cache logic 210 may maintain cache state by, among other things, compressing cache data 128 maintained within the second portion 126 (e.g., selecting cache data 128 of R cache units 220 for eviction), moving cache data 128 from cache units 220 designated to be allocated to the first portion 124 to cache units 220 that will remain allocated to the second portion 126, and so on. When decreasing the amount of cache memory 120 allocated to the first portion, cache logic 210 may maintain prefetcher state by, among other things, compressing metadata 122 maintained within the first portion 124 to store it within a smaller number of cache units 220 (e.g., by removing portions of metadata 122, such as entries associated with address regions exhibiting poor prefetch performance), moving the compressed metadata 122 to cache units 220 that will remain allocated to the first portion 124, and so on.

[0141] At 606, increasing the amount of cache memory 120 allocated to the first portion (e.g., first portion 124) can include allocating one or more cache units 220 from the second portion (e.g., second portion 126) to the first portion 124. As disclosed herein, allocating one or more cache units 220 to the first portion 124 can include flushing and / or demoting cache units 220, modifying the address mapping scheme 316 to disable, remove, and / or ignore cache units 220 (e.g., disabling cache tags 326 of cache units 220), modifying the metadata mapping scheme 314 to include and / or reference cache units 220, and so on. Thus, increasing the amount of cache memory 120 allocated to the first portion can include reducing the amount of cache memory 120 allocated to the second portion (and / or reducing the amount of cache memory 120 available for storing cache data 128). Reducing the size of the second portion 126 can include compressing the cache data 128 stored within the second portion 126 of the cache memory 120, and the compression can include selecting cache data 128 for removal and / or eviction from the cache 110. The cache data 128 can be selected according to any suitable replacement or eviction policy (such as FIFO, LIFO, LRU, TLRU, MRU, LFU, random replacement, etc.). Compressing the cache data 128 can include reducing the amount of cache memory 120 consumed by the cache data 128 by R cache units 220, where R is the number of cache units 220 allocated from the second portion 126 to the first portion 124 (or R·U, where U is the capacity of a cache unit 220, CMU 320, or way 420).

[0142] In some aspects, the cache logic 210 selects a first set of cache units 220 for reallocation to the first portion 124 and selects a second set of cache units 220 for eviction. Each of the first and second sets may include R cache units 220, where R is the number of cache units 220 to be reallocated to the first portion 124. The first and second sets may be selected independently and / or according to respective selection criteria. The first set of cache units 220 may be selected according to an address mapping scheme 316, a metadata mapping scheme 314, a partitioning scheme 312, etc. (the schemes may allocate cache units 220 for storing metadata 122 according to a predetermined pattern or scheme, such as a sequential scheme, a way-based partitioning scheme 312-1, a set-based partitioning scheme 312-2, etc.). As disclosed herein, the second set of cache units 220 may be selected according to an eviction or replacement policy. Reallocating R cache units 220 may include: a) clearing the second set of cache units 220; and b) moving cache data 128 from cache units 220 included in the first set (but not in the second set) to the second set of cache units 220. Thus, when reducing the available cache data 128 capacity of the cache 110, the cache logic 210 can retain data that is more frequently accessed within the cache memory 120.

[0143] In some embodiments, the cache logic 210 partitions the cache memory 110 according to a way or way-based partitioning scheme 312-1. Allocating R cache units 220 from the second portion 126 to the first portion 124 may include allocating one or more ways 420 within each set 430 of the cache to the first portion 124. Allocating R cache units 220 from the second portion 126 to the first portion 124 may include allocating an additional W 1A ways 420 from each of S sets 430-1 to 430-S from the second portion 126 to the first portion 124, where Alternatively or additionally, the cache logic 210 may partition the cache memory 110 according to a set or set-based partitioning scheme 312-2. Allocating R cache units 220 from the second portion 126 to the first portion 124 may include allocating an additional E 1A sets 430 of the cache memory 120 from the second portion 126 to the first portion 124, where and N is the number of cache units 220 (or ways 420) included in each set 430.

[0144] At 606, reducing the amount of cache memory 120 allocated to a first portion (e.g., first portion 124) can include allocating one or more cache units 220 from the first portion to a second portion (e.g., second portion 126). As disclosed herein, allocating one or more cache units 220 to the second portion 126 can include modifying an address mapping scheme 316 to enable, include, and / or otherwise reference the cache unit 220 (e.g., enable a cache tag 326 of the cache unit 220), modifying a metadata mapping scheme 314 to remove the cache unit 220, and so on.

[0145] Reducing the amount of cache memory 120 allocated to the first portion can further include compressing metadata 122. The metadata 122 can be compressed to be stored within fewer R cache units 220, where R is the number of cache units 220 to be allocated from the first portion 124 to the second portion 126. At 606, compressing the metadata 122 can include removing a portion of the metadata 122, such as one or more entries of the metadata 122. The portion of the metadata 122 can be selected based on a removal criterion, such as a stage criterion (oldest first removal, newest first removal, etc.), a least recently accessed criterion, a least frequently accessed criterion, etc.

[0146] Alternatively or additionally, respective portions of the metadata 122 can be selected for removal at least in part based on one or more metrics 212. The metadata 122 can include a plurality of entries, each entry including access information related to a corresponding region of an address space. The prefetch logic 230 can utilize the corresponding entries of the metadata 122 to perform prefetch operations within the address regions covered by the corresponding entries. The one or more metrics 212 can be configured to quantify prefetch performance within the address regions covered by the corresponding entries of the metadata 122. Compressing the metadata 122 can include selecting entries of the metadata 122 for removal at least in part based on prefetch performance as quantified by the metrics 212 within the address regions covered by the entries. In some embodiments, entries of the metadata 122 with prefetch performance below a threshold can be removed (and / or the amount of memory capacity allocated to the entries can be reduced). Alternatively, entries of the metadata 122 that exhibit higher prefetch performance can be retained, while entries that exhibit lower prefetch performance can be removed (e.g., the R entries of the metadata 122 with the lowest performance can be selected for removal). Thus, compressing the metadata can include removing the metadata 122 from one or more cache units 220 and / or moving the metadata 122 (and / or entries of the metadata 122) from cache units 220 that are reallocated to the second portion 126 to the remaining cache units 220 allocated to the first portion 124.

[0147] Figure 7Another example of a method for a device to implement adaptive cache partitioning is shown in flow chart 700. Flow chart 700 includes blocks 702 through 708. At 702, the logic of cache 110 (e.g., cache logic 210) implements partitioning scheme 312 to partition cache memory 120, in particular, into a first portion 124 and a second portion 126. The first portion 124 may include a first portion of cache memory 120, and the second portion 126 may include a second portion of cache memory 120, where the second portion is different from the first portion 124. The first portion 124 may be allocated to store metadata 122 related to an address space (such as the address space associated with the backing memory (memory 108) of cache 110). The second portion 126 may be allocated to store cache data 128 related to the address space (e.g., the available cache capacity of cache 110). Partitioning cache memory 120 may include implementing metadata mapping scheme 314 to access cache units 220 allocated to the first portion 124, and implementing address mapping scheme 316 to map addresses 204 of the address space to cache units 220 allocated to the second portion 126.

[0148] At 704, cache 110 services requests related to the address space, which may include maintaining metadata related to the address space within the first portion 124 (e.g., within metadata 122 maintained within the first portion 124) and loading data associated with the addresses of the address space into the second portion 126. Data may be loaded into cache memory 120 in response to a cache miss (such as a request 202 related to an address 204 not available within cache 110). Alternatively or additionally, at 704, data may be prefetched into cache memory 120. The prefetcher (and / or prefetch logic 230 of cache 110) may utilize metadata 122 maintained within the first portion 124 to predict the address 204 of an upcoming request 202, and data corresponding to the predicted address 204 may be prefetched into the second portion 126 before a request 202 related to the predicted address 204 is received at cache 110.

[0149] At 706, cache logic 210 (and / or prefetch logic 230) may determine whether to adjust the partitioning scheme 312 of cache memory 120. More specifically, at 706, cache logic 210 (and / or prefetch logic 230) may determine whether to modify the size of the first portion 124 allocated for metadata 122 (and / or modify the size of the second portion 126 allocated for storing cache data 128). As disclosed herein, the determination may be based at least in part on one or more metrics 212 that may be configured to quantify prefetch performance. Determining whether to adjust the partitioning scheme 312 may include determining and / or monitoring one or more metrics 212 related to data prefetched into the second portion 126 and comparing the metrics 212 to one or more thresholds. At 708, the partitioning scheme 312 may be adjusted in response to one or more of the metrics 212 being greater than a first threshold and / or less than a second threshold; otherwise, the process may continue at 704, where cache 110 may continue to service requests related to the address space.

[0150] At 708, cache logic 210 adjusts the partitioning scheme to, in particular, modify the amount of cache memory 120 allocated to the first portion 124 and / or the second portion 126. At 708, when one or more of the metrics 212 exceed one or more first thresholds (e.g., when prefetch performance exceeds one or more first thresholds), the size of the first portion 124 allocated for metadata 122 may be increased (and the size of the second portion 126 allocated for cache data 128 may be decreased). Conversely, when one or more of the metrics 212 are below one or more second thresholds (e.g., when prefetch performance is below one or more second thresholds), the size of the first portion 124 may be decreased (and the size of the second portion 126 may be increased).

[0151] Increasing the size of the first portion 124 can include allocating cache resources from the second portion 126 to the first portion 124 (e.g., one or more cache units 220, ways 420, sets 430, etc.). Increasing the size of the first portion 124 can include decreasing the size of the second portion 126. As disclosed herein, decreasing the size of the second portion 126 can include compressing the cache data 128 stored within the second portion 126 (e.g., by selecting cache data 128 for eviction, moving cache data 128 to remaining cache units 220 allocated to the second portion 126, etc.). Conversely, decreasing the size of the first portion 124 can include allocating cache resources from the first portion 124 to the second portion 126. As disclosed herein, decreasing the size of the first portion 124 can include compressing the metadata 122 stored within the first portion 124 of the cache memory 120 (e.g., by selecting portions of the metadata 122 for removal, moving portions of the metadata 122 to remaining cache units 220 allocated to the first portion 124, etc.). As disclosed herein, in response to adjusting the partitioning scheme 312 of the cache memory 120 at 708, the process can continue at 704, where the cache 110 can service requests related to the address space.

[0152] Figure 8 Another example flowchart 800 depicting operations for adaptively partitioning a cache based at least in part on a metric 212 related to prefetch performance is shown. Flowchart 800 includes blocks 802 through 816. At 802, cache logic 210 (and / or prefetch logic 230) partitions the cache memory 120 into a first portion 124 and a second portion 126. The first portion 124 includes a first partition of the cache memory 120 allocated for metadata 122 related to an address space, and the second portion 126 can include a second partition of the cache memory 120 allocated for cache data 128 (separate from the first portion 124).

[0153] At 804, the cache 110 services requests related to the address space, which can include loading data into the second portion 126 of the cache memory 120, retrieving data associated with a corresponding address 204 of the address space from the second portion 126 of the cache memory 120 in response to a request 202 related to the address 204, maintaining metadata 122 related to access to corresponding addresses and / or regions of the address space within the first portion 124 of the cache memory 120, using the metadata 122 maintained within the first portion 124 of the cache memory 120 to prefetch cache data 128 into the second portion 126 of the cache memory 120, etc.

[0154] At 806, cache logic 210 (and / or prefetch logic 230) determines whether to evaluate the partitioning scheme 312 of cache memory 120. In some embodiments, the partitioning scheme 312 can be evaluated in the background and / or by using the idle resources of cache 110. The determination at 806 can be based at least in part on whether cache 110 is idle (e.g., whether one or more requests 202 are being serviced), whether idle resources are available, etc. The determination at 806 can be based on one or more time-based criteria (e.g., the partitioning scheme can be evaluated periodically and / or at determined time intervals), a predetermined schedule, etc. Alternatively or additionally, the determination at 806 can be triggered by workload conditions and / or prefetch performance metrics (e.g., one or more metrics 212). Cache logic 210 can be configured to periodically and / or continuously determine and / or monitor metrics 212 related to prefetch performance, and can trigger an evaluation of the partitioning scheme 312 at 806 in response to the metrics 212 exceeding and / or falling below one or more thresholds.

[0155] If the determination at 806 is to evaluate the partitioning scheme 312, the process continues at 808; otherwise, the process continues to service requests related to the address space at 804.

[0156] At 808, cache logic 210 (and / or prefetch logic 230) can determine and / or monitor one or more aspects of prefetch performance, such as prefetch hit rate, prefetch miss rate, number of useful prefetches, number of invalid prefetches, ratio of useful prefetches to invalid prefetches, etc. At 808, cache logic 210 (and / or prefetch logic 230) can determine and / or monitor one or more metrics 212 related to prefetch performance, as disclosed herein. The prefetch hit rate can be based on the access metrics of the prefetched cache data 128 maintained in the cache metadata 122 associated with the prefetched cache data 128. The cache data 128 prefetched into cache memory 120 can be identified by using a prefetch indicator, such as a prefetch tag associated with the cache data 128, which can be maintained in the cache metadata 322 associated with the cache unit 220 storing the cache data 128.

[0157] At 810, the prefetch performance determined at 806 is compared with a first threshold. If the prefetch performance exceeds the first threshold, the process continues at 812; otherwise, the process continues at 814. In some embodiments, the determination at 810 is based on whether the prefetch performance determined at 810 exceeds the first threshold and whether the amount of cache memory 120 currently allocated to the first portion 124 is less than a maximum amount, threshold, or upper limit. If so, the process continues at 812; otherwise, the process continues at 814.

[0158] At 812, cache logic 210 modifies the partitioning scheme 312 to increase the amount of cache memory 120 allocated to store metadata 122 related to the address space (e.g., increase the size of the first portion 124 of cache memory 120). Increasing the amount of cache memory 120 allocated to the first portion 124 can include decreasing the amount of cache memory 120 allocated to the second portion 126 (e.g., decreasing the available capacity of cache 110). At 812, cache logic 210 can reallocate specified cache memory resources (such as one or more cache units 220, ways 420, sets 430, etc.) from the second portion 126 to the first portion 124. For example, cache memory 120 can be partitioned into a first portion 124 including a first group of cache units 220 and a second portion 126 including a second group of cache units, the second group being different from the first group. Increasing the amount of cache memory allocated to the first partition (first portion 124) can include allocating one or more cache units 220 from the second group to the first group, particularly by evicting and / or moving cache data 128 from one or more cache units 220, removing one or more cache units 220 from the address mapping scheme 316 (e.g., disabling the cache tags 326 of one or more cache units 220), adding one or more cache units 220 to the metadata mapping scheme 314, and so on.

[0159] As disclosed herein, cache logic 210 can be further configured to: compress cache data 128 stored within the second portion 126 to store it within a smaller amount of cache memory 120; configure the address mapping scheme 316 to remove, disable, and / or dereference specified cache memory resources; configure the metadata mapping scheme 314 to include, reference, and / or otherwise provide access to specified cache resources for storing metadata 122, and so on. In response to implementing the modified partitioning scheme 312 to increase the amount of cache memory 120 allocated for metadata 122 (and decrease the amount of available cache capacity), the process can continue at 804.

[0160] At 814, the prefetch performance determined and / or monitored at 808 is compared with a second threshold. If the prefetch performance is below the second threshold, the process continues at 816; otherwise, the process continues at 804. In some embodiments, the determination at 814 is based on whether the prefetch performance determined at 810 is below the second threshold and whether the amount of cache memory 120 currently allocated to the first portion 124 is above a minimum amount, threshold, or lower bound. If so, the process continues at 816; otherwise, the process continues at 814.

[0161] At 816, the cache logic 210 modifies the partitioning scheme 312 to reduce the amount of cache memory 120 allocated to store metadata 122 related to the address space (e.g., reducing the size of the first portion 124 of the cache memory 120). Reducing the amount of cache memory 120 allocated to the first portion 124 can include increasing the amount of cache memory 120 allocated to the second portion 126 (e.g., increasing the available capacity of the cache 110). At 816, the cache logic 210 can reallocate specified cache memory resources (such as one or more cache units 220, ways 420, sets 430, etc.) from the first portion 124 to the second portion 126. At 816, the cache logic 210 can be further configured to: compress the metadata 122 stored within the first portion 124 to store it within a smaller amount of cache memory 120; configure the address mapping scheme 316 to enable, reference, and / or otherwise utilize the specified cache memory resources for caching data 128; configure the metadata mapping scheme 314 to remove, exclude, and / or dereference the specified cache resources, etc., as disclosed herein. In response to implementing the modified partitioning scheme 312 to reduce the amount of cache memory 120 allocated for the metadata 122 (and increasing the amount of available cache capacity), the process can continue at 804.

[0162] Figure 9Shows an example flowchart 900 depicting operations for adaptively partitioning a cache based at least in part on metrics related to cache and / or prefetch performance. Flowchart 900 includes blocks 902 through 916. At 902, cache 110 partitions its cache memory 120 into a first portion 124 and a second portion 126. The first portion 124 may include a first portion of the cache memory 120 (e.g., zero or more cache units 220, cache lines, hardware cache lines, ways 420, sets 430, etc.). The second portion 126 may include a second portion of the cache memory 120 that is different from the first portion 124. The first portion 124 may be allocated for metadata 122 related to the address space, and the second portion 126 may be allocated for storing cache data 128 (which may be the available cache capacity).

[0163] At 904, cache 110 services requests related to the address space, which may in particular include receiving request 202, loading cache data 128 into the second portion 126 of the cache memory 120 in response to a cache miss, servicing request 202 by using the cache data 128 stored within the second portion 126 of the cache memory 120, and so on.

[0164] At 906, cache 110, cache logic 210, prefetch logic 230, and / or a prefetcher coupled to cache 110 maintains metadata 122 related to address access characteristics within the first portion 124 of the cache memory 120. As disclosed herein, metadata 122 may include any suitable information related to accesses to corresponding addresses and / or address regions of the address space. At 908, cache data 128 is prefetched into the second portion 126 of the cache memory 120 at least in part based on the metadata 122 maintained within the first portion 124 of the cache memory 120.

[0165] At 910, a cache, cache logic 210, prefetch logic 230, and / or a prefetcher coupled to cache 110 determine and / or monitor one or more metrics 212. As disclosed herein, metrics 212 can be configured to quantify cache and / or prefetch performance. At 912, the metrics 212 are evaluated to determine whether to adjust the partitioning scheme 312 of cache memory 120 (e.g., determine whether to adjust the amount of cache memory 120 allocated to the first portion 124 or the second portion 126). The determination at 912 can be based at least in part on the metrics 212 determined and / or monitored at 910. The determination at 912 can adjust the partitioning scheme 312 based on cache performance and / or prefetch performance. The partitioning scheme 312 can be adjusted at 914 in response to: a) the metrics 212 exceeding one or more thresholds; b) the prefetch performance exceeding one or more prefetch thresholds; c) the cache performance exceeding one or more cache thresholds, etc. The determination at 912 can be based on whether the prefetch performance (e.g., prefetch hit rate) is above an upper prefetch threshold or below a lower performance threshold, whether the cache performance (e.g., cache hit rate) is above an upper cache threshold or below a lower cache threshold, etc. In some embodiments, the determination at 912 can be based on both prefetch performance and cache performance (which can be configured to balance prefetch performance and cache performance). The determination at 912 can be based on: a) whether the prefetch performance exceeds a first prefetch threshold and the cache performance is below a first cache threshold; b) whether the prefetch performance is below a second prefetch threshold and the cache performance is above a second cache performance threshold, etc.

[0166] Alternatively or additionally, the determination at 912 can be based in particular on the amount of cache memory 120 currently allocated to metadata 122 in the first portion 124 (metadata capacity). The determination 912 can be based on whether the prefetch performance is above a first prefetch performance threshold and the metadata capacity is below a first capacity threshold (e.g., a first prefetch capacity threshold or metadata capacity threshold), whether the prefetch performance is below a second prefetch performance threshold and the metadata capacity is above a second capacity threshold (e.g., a second prefetch capacity threshold or metadata capacity threshold), etc. In some embodiments, the determination at 912 is based on cache performance. The determination can be based on whether the cache performance quantified by the metrics 212 (e.g., cache performance metrics 212) is below a cache performance threshold. At 914, the amount of cache memory 120 allocated to store metadata 122 related to the address space can be adjusted iteratively and / or periodically to improve cache performance (e.g., increase or decrease).

[0167] At 914, sizing of the first portion 124 and / or the second portion 126 is determined. The sizing can be at least partially based on the metric 212 determined and / or monitored at 910 (and / or the evaluation of the metric 212 at 912). At 914, when the prefetch performance quantified by the metric 212 is at or above an upper prefetch threshold (and the metadata capacity is below the determined maximum value), the size of the first portion 124 allocated for metadata 122 related to the address space can be increased. Conversely, when the prefetch performance quantified by the metric 212 is at or below a lower prefetch threshold, the size of the first portion 124 can be decreased. In another example, the amount of cache memory 120 allocated to the first portion 124: a) can be increased when the prefetch performance is above a first prefetch threshold and the cache performance is below a first cache threshold; or b) can be decreased when the prefetch performance is below a second prefetch threshold and the cache performance is above a second cache performance threshold, etc.

[0168] Alternatively or additionally, the sizing can be based in particular on the amount of cache memory 120 (metadata capacity) currently allocated for the first portion 124 of the metadata 122. When the prefetch performance is above a first prefetch performance threshold and the amount of cache memory 120 currently allocated to the first portion 124 is below a first capacity threshold, the amount of cache memory 120 allocated to the first portion 124 can be increased. Conversely, when the prefetch performance is below a second prefetch performance threshold and the amount of cache memory 120 currently allocated to the first portion 124 is above a second capacity threshold, the amount of cache memory 120 allocated to the first portion 124 can be decreased, etc. In some embodiments, the sizing at 914 can be based on a cache performance metric, such as a cache hit rate. At 914, the amount of cache memory 120 allocated to store the metadata 122 related to the address space can be adjusted iteratively and / or periodically to achieve an improved cache hit rate (e.g., increase or decrease). In some embodiments, the determination at 912 and the sizing at 914 can be implemented according to an optimization algorithm that can be configured to converge to an optimal (or locally optimal) partitioning scheme 312 that results in an optimal (or locally optimal) cache performance as quantified by the metric 212.

[0169] Example System for Adaptive Cache Partitioning

[0170] Figure 10 An example system 1000 for adaptive cache partitioning is shown. As disclosed herein, the system 1000 can include a cache device 1001, which can include a cache 110 and / or means for implementing the cache 110. Figure 10 The description of involves the above aspects, such as in a plurality of other figures (e.g.,Figures 1-1 to 5-3 ) the cache 110 depicted in. The system 1000 may further include an interface 1015 for coupling the cache device 1001 to the interconnect 1005, receiving requests 202 related to the address 204 of the address space associated with the memory 108 (e.g., from the requester 201), performing a transfer operation 203 to fetch cache data 128 from the memory 108, and so on. The interface 1015 may be configured to couple the cache device 1001 to any suitable interconnect, including but not limited to: an interconnect, a physical interconnect, a bus, the interconnect 105 of the host device 102, the front-end interconnect 105A, the back-end interconnect 105B, etc. The interface 1015 may include but not limited to: circuitry, logic circuitry, interface circuitry, interface logic, switching circuitry, switching logic, routing circuitry, routing logic, interconnect circuitry, interconnect logic, I / O circuitry, analog circuitry, digital circuitry, logic gates, registers, switches, multiplexers, ALUs, state machines, microprocessors, embedded processors, PIM circuitry, logic 220, interface 215, first interface 215A, second interface 215B, etc.

[0171] The cache device 1001 may include and / or be coupled to a cache memory 120, which may include but not limited to: a memory, a memory array, a semiconductor memory, a volatile memory, a RAM, an SRAM, a DRAM, an SDRAM, etc. In Figure 10 an example, the cache memory 120 includes a plurality of cache units 220 (e.g., cache units 220-1 to 220-X), and each cache unit 220 includes and / or corresponds to a respective CMU 320 and / or cache tag 326. In some aspects, the cache units 220 are arranged in a plurality of sets 430 (e.g., sets 430-1 to 430-S), each set 430 includes a plurality of ways 420 (e.g., ways 420-1 to 420-N), and each way 420 includes and / or corresponds to a respective cache unit 220.

[0172] System 1000 may include component 1010, which is configured to allocate a first portion 124 of cache memory 120 for metadata 122 related to an address space, cache data in a second portion 126 of cache memory 120 that is different from the first portion 124 of cache memory 120, and / or modify the size of the first portion 124 of cache memory 120 allocated for metadata 122 at least in part based on a metric 212 related to data prefetched into the second portion 126 of cache memory 120. Component 1010 may be configured to partition cache memory 120 into a first partition 1024 that includes the first portion 124 of cache memory 120 and a second partition 1026 that includes the second portion 126 of cache memory 120. The first partition 1024 may be allocated for storing metadata 122, and the second partition 1026 may be allocated for storing cache data 128. Component 1010 may include, but is not limited to: circuitry, logic circuitry, memory interface circuitry, memory interface logic, switching circuitry, switching logic, routing circuitry, routing logic, memory interconnect circuitry, memory interconnect logic, I / O circuitry, analog circuitry, digital circuitry, logic gates, registers, switches, multiplexers, ALUs, state machines, microprocessors, embedded processors, PIM circuitry, cache logic 210, partitioning logic 310, partitioning scheme 312, metadata mapping scheme 314 (and / or metadata logic 1014), address mapping scheme 316 (and / or address logic 1016), etc.

[0173] Component 1010 may be configured to partition cache memory 120 according to partitioning scheme 312. Partitioning scheme 312 may define the logic, rules, criteria, and / or other mechanisms for partitioning cache memory resources (e.g., cache units 220) of cache memory 120 between the first partition 1024 and the second partition 1026. Partitioning scheme 312 may also be further configured to specify the amount, quantity, capacity, and / or size of the first partition 1024 and / or the second partition 1026 (e.g., may specify the amount, quantity, capacity, and / or size of the first portion 124 and / or the second portion 126). Partitioning scheme 312 may define the logic, rules, criteria, and / or other mechanisms for dynamically reallocating and / or reassigning cache memory resources between the first partition 1024 and / or the second partition 1026, such as a cache unit-based scheme, a way-based partitioning scheme 312-1, a set-based partitioning scheme 312-2, etc. In Figure 10 an example, partitioning scheme 312 configures component 1010 to allocate M cache units 220 to the first partition 1024 (and allocate X - M cache units 220 to the second partition 1026).

[0174] In some examples, the partitioning scheme 312 defines a cache unit-based scheme. In a cache unit-based scheme, allocating M cache units 220 to the first partition 1024 can include allocating cache units 220-1 to 220-M to the first portion 124 and / or allocating 220-M+1 to 220-X to the second portion 126, as Figure 10 illustrated. In other examples, the partitioning scheme 312 defines a way-based scheme (e.g., way partitioning scheme 312-1). In a way-based scheme, allocating M cache units 220 to the first partition 1024 can include allocating W1 ways 420 within each set 430 of the cache memory 120 to the first partition 1024, where and S is the number of sets 430 included in the cache memory 120 such that W2 ways 420 within each set 430 are allocated to the second partition 1026, where W2 = N - W1 or Alternatively, the partitioning scheme 312 can define a set-based scheme (e.g., set partitioning scheme 312-2). In a set-based scheme, allocating M cache units 220 to the first partition 1024 can include allocating E1 sets 430 to the first partition 1024, where and N is the number of ways 420 included in each set 430 such that E2 sets are allocated to the second partition 1026, where E2 = S - E1 or

[0175] Component 1010 can implement, include, and / or be coupled to metadata logic 1014. The metadata logic 1014 can be configured to map, address, associate, reference, and / or otherwise access (and / or provide access to the cache units) the cache units 220 allocated to the first partition 1024. As disclosed herein, the metadata logic 1014 can implement and / or include a metadata mapping scheme 314. The metadata logic 1014 can include but is not limited to: circuitry, logic circuitry, memory interface circuitry, memory interface logic, switching circuitry, switching logic, routing circuitry, routing logic, memory interconnect circuitry, memory interconnect logic, I / O circuitry, analog circuitry, digital circuitry, logic gates, registers, switches, multiplexers, ALUs, state machines, microprocessors, embedded processors, PIM circuitry, cache logic 210, partitioning logic 310, partitioning scheme 312, metadata mapping scheme 314, etc.

[0176] Component 1010 may implement, include, and / or be coupled to address logic 1016. Address logic 1016 may be configured to map, address, associate, reference, and / or otherwise access (and / or provide access to the cache unit) cache unit 220 assigned to second partition 1026. Address logic 1016 may be configured to map address 204 of the address space and / or associate the address with cache data 128 stored within cache unit 220 assigned to second partition 1026. As disclosed herein, address logic 1016 may implement and / or include address mapping scheme 316. Address logic 1016 may include, but is not limited to: circuitry, logic circuitry, memory interface circuitry, memory interface logic, switching circuitry, switching logic, routing circuitry, routing logic, memory interconnect circuitry, memory interconnect logic, I / O circuitry, analog circuitry, digital circuitry, logic gates, registers, switches, multiplexers, ALUs, state machines, microprocessors, embedded processors, PIM circuitry, cache logic 210, partitioning logic 310, partitioning scheme 312, address mapping scheme 316, etc.

[0177] Component 1010 may be further configured to adjust partitioning scheme 312 at least in part based on one or more metrics 212. As disclosed herein, metrics 212 may be configured to quantify prefetch performance. Alternatively or additionally, metrics 212 may be configured to quantify other aspects, such as cache performance (e.g., cache hit rate, cache miss rate, etc.). Component 1010 may be configured to determine and / or monitor metrics 212. Component 1010 may modify the size of first partition 1024 (and / or first portion 124) of cache memory 120 assigned to metadata 122 at least in part based on one or more of metrics 212.

[0178] Component 1010 may implement, include, and / or be coupled to a prefetcher 1030 to update metadata 122 maintained within a first portion 124 of a cache memory 120 in response to a request 202 related to an address 204 of an address space and / or to select data to be prefetched into a second portion 126 of the cache memory 120 based at least in part on the metadata 122 maintained within the first portion 124 of the cache memory 120. The metadata 122 may include any suitable information related to an address of an address space, including but not limited to: access characteristics, access statistics, address sequences, address histories, index tables, Δ sequences, stride patterns, correlation patterns, feature vectors, ML features, ML feature vectors, ML models, ML modeling data, and the like. The prefetcher 1030 may include but not limited to: circuitry, logic circuitry, memory interface circuitry, cache circuitry, switching circuitry, switching logic, routing circuitry, routing logic, interconnect circuitry, interconnect logic, I / O circuitry, analog circuitry, digital circuitry, logic gates, registers, switches, multiplexers, ALUs, state machines, microprocessors, embedded processors, PIM circuitry, cache logic 210, prefetch logic 230, stride prefetcher, correlation prefetcher, ML prefetcher, LSTM prefetcher, and the like.

[0179] In some aspects, component 1010 is configured to determine and / or monitor a metric 212 related to data prefetched into a second portion 126 of the cache memory 120 and, in response to the monitoring, modify the size of the first portion 124 of the cache memory 120. Component 1010 may be configured to increase the size of the first portion 124 of the cache memory 120 allocated for the metadata 122 (and decrease the size of the second portion 126) in response to the metric 212 being above a first threshold, or decrease the size of the first portion 124 (and increase the size of the second portion 126) in response to the metric 212 being below a second threshold. Alternatively or additionally, component 1010 may be configured to increase the size of the first portion 124 in response to the current size of the first portion 124 being below a metadata capacity threshold and one or more of the following: a) a prefetch performance metric 212 above a prefetch performance threshold and / or b) a cache performance metric 212 below a cache performance threshold. Conversely, component 1010 may be configured to decrease the size of the first portion 124 of the cache memory 120 in response to the current size of the first portion 124 being above a prefetch capacity threshold and one or more of the following: a) a prefetch performance metric 212 below a prefetch performance threshold and / or b) a cache performance metric 212 above a cache performance threshold.

[0180] Component 1010 can be configured to allocate one or more cache units 220 to a first partition 1024. Allocating cache unit 220 to the first partition 1024 (and / or the first portion 124) can include configuring metadata logic 1014 to address, reference one or more cache units 220 and / or provide access to the one or more cache units to store metadata 122 and / or removing, disabling, ignoring, and / or otherwise excluding cache unit 220 from address logic 1016. Conversely, allocating cache unit 220 to a second partition 1026 and / or a second portion 126 can include configuring address logic 1016 to address, reference cache unit 220 and / or otherwise use the cache unit as available cache capacity (e.g., to store cache data 128) and / or removing, disabling, ignoring, and / or otherwise excluding cache unit 220 from metadata logic 1014. Allocating cache unit 220 to the first portion 124 can include evicting cache data 128 from cache unit 220 and disabling cache tag 326 of cache unit 220. Allocating cache unit 220 to the second portion 126 can include removing metadata 122 from cache unit 220 and enabling cache tag 326 of cache unit 220.

[0181] Component 1010 can be configured to increase the size of the first portion 124 (e.g., in response to metric 212 being higher than a first threshold). Increasing the size of the first portion 124 can include compressing cache data 128 stored within the second portion 126. Component 1010 can be configured to maintain at least a portion of cache data 128 maintained within cache 110 when increasing the size of the first portion 124 (and decreasing the size of the second portion 126). In response to increasing the size of the first portion 124, component 1010 can be configured to evict cache data 128 from selected cache units 220 that remain allocated to the second portion 126. Component 1010 can be further configured to move cache data 128 to the selected cache units 220. Cache data 128 can be moved from cache units 220 that are to be allocated from the second portion 126 to the first portion 124.

[0182] Conversely, component 1010 may be configured to reduce the size of the first portion 124 (e.g., in response to the metric 212 being below a second threshold). Reducing the size of the first portion 124 may include compressing the metadata 122 stored within the first portion 124. Component 1010 may be configured to maintain at least a portion of the metadata 122 when reducing the size of the first portion 124. Component 1010 may be configured to reduce the amount of cache memory 120 allocated for the metadata 122 from a first set of cache units 220 to a second set of cache units 220, the second set being less than the first set. Component 1010 may be further configured to compress the metadata 122 to store it within the second set of cache units 220. Component 1010 may move the metadata 122 stored within a cache unit 220 included in the first set of cache units 220 to a cache unit included in the second set of cache units 220.

[0183] Conclusion

[0184] Although embodiments of adaptive cache partitioning have been described in language specific to certain features and / or methods, the subject matter of the appended claims is not necessarily limited to the specific features or methods described. Rather, the specific features and methods are disclosed as example embodiments of adaptive cache partitioning.

Claims

1. A method for operating a memory, comprising: Allocating a first portion of a cache memory for metadata related to an address space; Writing data associated with an address of the address space to a second portion of the cache memory different from the first portion of the cache memory; And Modifying the size of the first portion of the cache memory allocated for the metadata related to the address space at least in part based on a metric related to data prefetched into the second portion of the cache memory.

2. The method according to claim 1, further comprising: Updating the metadata maintained in the first portion of the cache memory in response to a request related to an address of the address space; And Prefetching data into the second portion of the cache memory at least in part based on the metadata related to the address space maintained in the first portion of the cache memory.

3. The method according to claim 1, further comprising maintaining in the first portion of the cache memory one or more of the following: an address sequence, an address history, an index table, a Δ sequence, a step pattern, a correlation pattern, a feature vector, a machine learning ML feature, an ML feature vector, an ML model, or ML modeling data.

4. The method according to claim 1, further comprising: Monitoring the metric related to the data prefetched into the second portion of the cache memory; And Modifying the size of the first portion of the cache memory in response to the monitoring.

5. The method according to claim 4, further comprising monitoring one or more of a prefetch hit rate, a number of useful prefetches, a number of invalid prefetches, or a ratio of useful prefetches to invalid prefetches.

6. The method according to claim 1, further comprising one of the following: Increasing the size of the first portion of the cache memory allocated for the metadata related to the address space in response to the metric exceeding a first threshold; or Decreasing the size of the first portion of the cache memory allocated for the metadata related to the address space in response to the metric being below a second threshold.

7. The method according to claim 6, further comprising one of the following: Decreasing the size of the second portion of the cache memory in response to the metric exceeding the first threshold; or Increasing the size of the second portion of the cache memory in response to the metric being below the second threshold.

8. The method according to claim 1, further comprising increasing the size of the first portion of the cache memory allocated for the metadata related to the address space in response to the size of the first portion of the cache memory being below a metadata capacity threshold and one or more of the following: A prefetch performance metric higher than a prefetch performance threshold; or A cache performance metric lower than a cache performance threshold.

9. The method according to claim 1, further comprising reducing the size of the first portion of the cache memory allocated for the metadata related to the address space in response to the size of the first portion of the cache memory being higher than a prefetch capacity threshold and one or more of the following: a prefetch performance metric below a prefetch performance threshold; or a cache performance metric above a cache performance threshold.

10. The method according to claim 1, further comprising: in response to allocating a set of cache units of the cache memory for the metadata related to the address space, removing the set of cache units from an address mapping scheme, and adding the set of cache units to a metadata mapping scheme.

11. The method according to claim 1, further comprising: in response to allocating a set of cache units from the second portion to the first portion, evicting cache data from the cache units in the set of cache units, and disabling cache tags associated with the cache units in the set of cache units.

12. The method according to claim 11, further comprising: evicting cache data from selected cache units that will remain allocated to the second portion; and moving cache data stored in the cache units in the set of cache units allocated from the second portion to the first portion to the selected cache units.

13. The method according to claim 1, further comprising reducing the amount of cache memory allocated for the metadata related to the address space from a first set of cache units of the cache memory to a second set of cache units of the cache memory, the second set being smaller than the first set, the reduction comprising: compressing the metadata to store it within the second set of cache units; moving the metadata stored in the cache units included in the first set of cache units to the cache units included in the second set of cache units; and allocating one or more cache units included in the first set of cache units to the second portion of the cache memory.

14. The method according to claim 1, further comprising: allocating a certain number of ways of the cache memory for the metadata related to the address space; and modifying the number of ways of the cache memory allocated for the metadata related to the address space at least in part based on the metric.

15. The method according to claim 14, further comprising: dividing the ways of the set of cache memories into a first set allocated for the metadata related to the address space and a second set allocated for caching data associated with the address space; and moving cache data from the ways within the first set to the ways within the second set.

16. The method according to claim 1, further comprising: Allocate a certain number of sets of the cache memory for the metadata related to the address space; and Modify at least in part the number of sets of the cache memory allocated for the metadata related to the address space based on the metric.

17. The method according to claim 1, further comprising: Selecting a subset of the metadata related to the address space in response to reducing the capacity of the first portion of the cache memory from a first capacity to a second capacity less than the first capacity; and Storing the selected subset of the metadata related to the address space within the reduced capacity of the first portion of the cache memory.

18. An apparatus for operating a memory, comprising: A memory array configured as a cache memory; and Logic coupled to the memory array and configured to: Allocate a first portion of the cache memory to store metadata related to an address space; Determine a metric related to cache data loaded into a second portion of the cache memory, the second portion being different from the first portion; and Modify at least in part the amount of the cache memory allocated to the first portion based on the determined metric.

19. The apparatus according to claim 18, further comprising partitioning logic configured to allocate a first set of cache units among a plurality of cache units of the cache memory to the first portion and a different second set of cache units among the plurality of cache units to the second portion.

20. The apparatus according to claim 19, wherein the cache units of the cache memory include one or more of the following: memory cells, blocks, memory blocks, cache blocks, cache memory blocks, pages, memory pages, cache pages, cache memory pages, cache lines, hardware cache lines, ways, rows of the memory array, or columns of the memory array.

21. The apparatus according to claim 19, further comprising prefetch logic configured to: Store metadata related to the address space within the first set of cache units; Determine at least in part addresses of upcoming requests based on the metadata stored within the first set of cache units; and Write data corresponding to one or more of the determined addresses to the second set of cache units.

22. The apparatus according to claim 19, wherein: The cache memory includes a plurality of sets, each set including a plurality of ways, each way corresponding to a respective cache unit among the plurality of cache units of the cache memory; and The first portion includes one or more ways within each of the plurality of sets.

23. The apparatus according to claim 22, wherein to increase the amount of the cache memory allocated to the first portion, the logic is configured to: Evict cache data from a selected way of one of the plurality of sets; Move the cache data from the designated way of the set to the selected way of the set; Disable the cache tag of the designated way; And Include the designated way of the set in the first portion of the cache memory.

24. The apparatus according to claim 19, wherein: The cache memory includes a plurality of sets, each set including a plurality of ways; The first portion includes a first group of the plurality of sets; The second portion includes a second group of the plurality of sets, the second group being different from the first group; and To increase the amount of the cache memory allocated to the first portion, the logic is configured to: Evict cache data from sets of the second group; Disable the cache tags of each way included in the sets; and Include the sets in the first group.

25. A system for operating a memory, comprising: A cache memory, the cache memory including a plurality of cache units; An interface configured to be coupled to an interconnect of a computing device; And Logic coupled to the interface and the cache memory, the logic being configured to: Partition the cache memory into a first portion and a second portion, allocate the first portion for metadata related to an address space associated with the memory, and the second portion is allocated for storing cache data corresponding to addresses of the address space; Write metadata related to the address space to the first portion of the cache memory; Write data to the second portion of the cache memory at least in part based on the metadata written to the first portion of the cache memory; And Modify the amount of the cache memory allocated to the first portion at least in part based on a metric related to the data written to the second portion of the cache memory.

26. The system according to claim 25, wherein the logic is further configured to: Allocate a first number of ways within each of a plurality of sets of the cache memory to the first portion; and Store the metadata related to the address space within the first number of ways, the first number of ways being allocated within each of the plurality of sets of the cache memory.

27. The system according to claim 26, wherein the logic is further configured to: Allocate a second number of ways within each of the plurality of sets of the cache memory to the second portion; and Cache data associated with addresses of the address space within the second number of ways, the second number of ways being allocated within each of the plurality of sets of the cache memory.

28. The system according to claim 27, wherein to increase the amount of the cache memory allocated to the first portion, the logic is further configured to allocate a specified way within each of the plurality of sets of the cache memory from the second portion to the first portion, the allocation including being further configured to: evict cache data from a selected way of one of the plurality of sets; move cache data from the specified way to the selected way; and disable the cache tag of the specified way.

29. The system according to claim 25, wherein the logic is further configured to: allocate a first group of one or more sets of the plurality of sets of the cache memory to the first portion; allocate a second group of one or more sets of the plurality of sets of the cache memory to the second portion, the second group being different from the first group; store the metadata related to the address space within the first group of one or more sets; and cache data associated with the addresses of the address space within the second group of one or more sets.

Citation Information

Patent Citations

  • Memory controllers employing memory capacity and / or bandwidth compression with next read address prefetching, and related processor-based systems and methods

    CN106462495A

  • Cache space management method and device

    CN110688062A