Methods to provide cache coherence based on cache type

By expanding coherency directory states to support multiple agent classes and using hardware-based solutions, the inefficiencies of conventional cache coherency mechanisms for huge cache types are addressed, improving performance in multi-chip systems.

DE102018005453B4Active Publication Date: 2025-08-28INTEL CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
DE102018005453
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Priority Date
2017-08-07
Filing Date
2018-07-09
Publication Date
2025-08-28
Estimated Expiration
2038-07-09

AI Technical Summary

Technical Problem

Conventional cache coherency mechanisms are inefficient for huge cache types, leading to high snoop latency and negative impacts on processor performance, and existing software-based methods are resource-intensive and inefficient.

Method used

Implement cache coherency processes that support both small and huge cache types by expanding the state tracked in a coherency directory to include multiple agent classes, using hardware-based solutions that track coherency without intensive resource paging, and employing directory-based protocols to manage cache coherency across multi-chip configurations.

Benefits of technology

Improves cache coherency performance for huge cache types, reducing snoop latency and enhancing processor performance in multi-chip computing environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000023_0000
    Figure 00000023_0000
  • Figure 00000024_0000
    Figure 00000024_0000
  • Figure 00000025_0000
    Figure 00000025_0000
Patent Text Reader

Abstract

Device (104, 405, 505) comprising: at least one processor (120, 420); at least one cache memory (124a ... 124n, 184a, ..., 184n, 484a, 484b); at least one system memory (156, 510) storing a directory (152, 520) comprising cache status information to indicate whether the at least one cache memory (124a ... 124n, 184a, ..., 184n, 484a, 484b) is a huge cache (484b) or a small cache (484a). and logic (180) at least partially contained in hardware, wherein the logic (180) is designed to: receive a memory operation request associated with the at least one cache memory (124a ... 124n, 184a, ..., 184n, 484a, 484b), determine a cache status of the memory operation request, wherein the cache status indicates either a huge cache status (484b) or a small cache status (484a), execute the memory operation request via a small cache coherence process (484a) in response to a determination that the cache state is a small cache state (484a), and execute the memory operation request via a giant cache coherency process (484b) in response to determining that the cache state is a giant cache state (484b).
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments herein relate generally to information processing, and more particularly to providing cache coherence techniques that support various types of caches within a computer system. BACKGROUND

[0002] Computer systems typically include various components coupled together to cooperate and perform various processing functions under the control of a central processor, commonly referred to as a central processing unit (CPU). Most systems also include a collection of elements such as processors, processor cores (or "cores"), accelerators, memory devices, peripherals, associated processing units, and the like, as well as additional semiconductor devices that may act as system memory to provide storage for information used by the processing units. In many systems, multiple memories may be present, each of which may be associated with a given element, such as a core or accelerator, which may act as local memory or cache for the corresponding element.Local caches can be composed of different types of storage devices, differentiated based on size, speed, or other operating characteristics.

[0003] To maintain the coherence of data across the system, a cache coherence protocol may be implemented, such as a snoop-based protocol, a directory-based protocol, combinations thereof, and / or variations thereof. Each form of coherence protocol may operate differently for specific operating characteristics of local caches of system elements. For example, a particular coherence protocol may operate efficiently for a particular cache type, but may operate less efficiently or not at all for other cache types.

[0004] US 2014 / 0 281 239 A1 shows a method for determining an enclosure strategy comprising determining a ratio of a capacity of a large cache to a capacity of a core cache in a cache subsystem of a processor and selecting an enclosure strategy in response to the cache ratio exceeding an enclosure threshold.

[0005] US 2016 / 0321288 A1 shows a cloud-based storage server connected to one or more storage devices that store shared content that can be accessed by two or more users over a network.

[0006] US 2017 / 0 139 836 A1 shows a near-memory acceleration method for offloading data operations from a processing element.

[0007] The stated object is achieved according to the invention by the features of patent claim 1. Further embodiments of the invention are presented in the subclaims and independent claims. BRIEF DESCRIPTION OF THE DRAWINGS Fig. 1 illustrates an embodiment of a first operating environment. Fig. 2 illustrates an embodiment of a first logical flow. Fig. 3 illustrates an embodiment of a second logic flow. Fig. 4 illustrates an embodiment of a second operating environment. Fig. 5 illustrates an embodiment of a third operating environment. Fig. 6 illustrates an embodiment of a third logic flow. Fig. 7 illustrates an embodiment of a fourth logic flow. Fig. Figure 8 illustrates an example of a storage medium. Fig. 9 illustrates one embodiment of a computer architecture. DETAILED DESCRIPTION

[0008] Various embodiments may be generally directed to methods for providing at least one cache coherence process between multiple components within a processing system. In some embodiments, the cache coherence process may operate within a multi-component (e.g., multi-chip) computing environment to, among other things, track cache lines homed in an associated memory of a first component (e.g., a first chip) when they are cached on a second component (e.g., a second chip). In some embodiments, the first chip may include a processor chip. In various embodiments, the second chip may include a processor chip, an accelerator chip, or the like.

[0009] In some embodiments, the plurality of components may include a processor, such as a central processing unit (CPU), a processor chip (including, for example, a multi-chip processor chip), a processor core (or "core"), an accelerator, an input / output (I / O) device, an intelligent I / O device, a storage device, combinations thereof, and / or the like. In some embodiments, at least a portion of the plurality of components may include and / or be operatively coupled to memory (a "local memory," "cache memory," or "cache"). The caches of the plurality of components may include one or more cache types based on one or more cache characteristics. Non-limiting examples of cache characteristics may include size, speed, location, associated component, usage, and / or the like.The cache type of a cache may affect the operation of the at least one cache coherence process interacting with the cache. For example, the at least one cache coherence process may include a first cache coherence process that operates efficiently for a first cache type and inefficiently for a second cache coherence process.

[0010] In some embodiments, the cache type may be based on a cache size. For example, caches may be identified as a "giant cache type" or a "small cache type." In some embodiments, a small cache type may include a cache with a memory size of a typical cache used in a processor core, while a giant cache type may include caches with a size larger than a typical cache used in a processor core or with a different size threshold. In various embodiments, a small cache type may have a memory size in a range of 1 to 10 megabytes (MB), while a giant cache type may have a memory size in a range of 10-100 MB or in the gigabyte (GB) range. In various embodiments, a small cache type may have a memory size of approximately less than 2 MB to approximately 12 MB.In some embodiments, a small cache type may have a memory size of approximately 500 kilobytes (KB), 1 MB, 2 MB, 4 MB, 6 MB, 8 MB, 10 MB, 12 MB, 16 MB, and ranges and values ​​between any two of these values ​​(including the endpoints). In various embodiments, a huge cache type may have a memory size of greater than approximately 10 MB to greater than approximately 10 GB. In some embodiments, a huge cache type may have a memory size of approximately 10 MB, 20 MB, 30 MB, 40 MB, 50 MB, 100 MB, 200 MB, 500 MB, 1 GB, 5 GB, 10 GB, 25 GB, 50 GB, 100 GB, 500 GB, and ranges and values ​​between any two of these values ​​(including the endpoints). Embodiments are not limited in this regard.

[0011] In general, the memory size of a small cache type and a giant cache type may depend on the operating characteristics of the respective system, including any cache coherence processes. In general, as described in more detail below, certain cache coherence processes or portions thereof may operate efficiently for small cache types, while certain cache coherence processes or portions thereof may operate inefficiently (or not at all) for giant cache types. For example, in a first system using a first cache coherence protocol, a giant cache type may have a memory size of approximately 500 MB, while in a second system using a second cache coherence protocol, a giant cache type may have a memory size of approximately 10 GB.For example, in some embodiments, a small cache type and a huge cache type may be distinguished based on the impact that the memory size of a cache has on an operating characteristic of a cache coherence process, such as snoop latency.

[0012] For a snoop-based cache coherence protocol, the snoop latency for huge cache types (e.g., an accelerator or an intelligent I / O device) can be very high compared to small cache types of a processor core. Accordingly, huge cache types can be associated with "high snoop latency," and small cache types can be associated with "low snoop latency." When agents with huge cache types are treated as conventional coherence agents, standard cache coherence mechanisms can cause the snoop latency of an agent associated with the huge cache type to become a limiter of the memory read latency, as seen, for example, by a processor core. Accordingly, standard cache coherence mechanisms for huge cache types typically have a significant negative impact on processor performance.

[0013] Conventional software-based methods for using standard cache coherence mechanisms for huge cache types generally involve managing data at a low level of granularity. For example, data can be managed at page granularity, where caching a page of data at a coherent agent may require that the page be deallocated from the page tables seen by processor cores. Additionally, certain operations, such as a reverse page table operation, may be required when returning a page from the cache to host memory. Accordingly, such conventional software-based methods for providing cache coherence for huge caches are extremely inefficient and provide poor performance.

[0014] For snoop-based cache coherence systems that employ small caches, a random-access memory (RAM)-based snoop filter, such as static RAM (SRAM), associated with a coherent agent can be implemented, for example, on a processor chip. However, SRAM-based snoop filters cannot be applied to huge caches due to the size of the snoop filter required to map such a huge cache. For example, such a snoop filter would consume an unacceptably large percentage of the area and power available to a processor chip.

[0015] Accordingly, some embodiments provide methods for supporting a cache coherence process for huge cache types. In some embodiments, the cache coherence process may include snoop-based coherence processes, directory-based coherence processes, combinations thereof, and / or variations thereof. In various embodiments, the cache coherence process may provide multiple coherence flows suitable for supporting small cache types and huge cache types within the same or related architecture. For example, in some embodiments, the cache coherence process may extend the state tracked in a coherence directory (e.g., an in-memory server coherence directory) to track multiple agents instead of a single class of agents as provided according to conventional methods.In various embodiments, the multiple agents may include agents for huge cache types (or high snoop latency) ("high-latency agents") and / or agents for small cache types (or low snoop latency) ("low-latency agents"). Agents, such as a processor core with a small cache type, may use low-latency directory bits to track coherency, while agents with a huge cache type may use high-latency directory bits to track coherency. In various embodiments, an agent may specify which agent type to use for a particular request. In this way, an agent may implement both a huge cache type (or high snoop latency) and a small cache type (or low snoop latency).Accordingly, in some embodiments, coherence for huge cache types may be tracked at least partially in hardware without the resource-intensive paging-related software activities required for software-based solutions. Therefore, in some embodiments, systems may achieve improved cache coherence performance for huge cache types. For example, some embodiments may improve cache coherence performance for dense computer logic devices, such as computer accelerators (e.g., graphics processing units (GPUs)) and / or high-gate-count field-programmable gate arrays (FPGAs) that operate on computations with a data footprint larger than the capacity of their associated memory.

[0016] Fig. 1 illustrates an example of operating environment 100 that may be representative of various embodiments. Fig. The operating environment 100 illustrated in Figure 1 may include a device 105 with a processing unit 120, such as a central processing unit (CPU). In some embodiments, the processing unit 120 may be implemented on a system-on-a-chip (SoC). In some embodiments, the processing unit 120 may be implemented as a standalone processor chip. The processing unit 120 may include one or more processing cores 122a-n, such as 1, 2, 4, 6, 8, 10, 12, or 16 processing cores. Embodiments are not limited in this context.The processing unit 120 may include any type of computer-based element, such as, but not limited to, a microprocessor, a microcontroller, a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a virtual processor (e.g., a VCPU), or any other type of processor or processing circuitry. In some embodiments, the processing unit 120 may be one or more processors in the family of Intel® processors available from Intel® Corporation, Santa Clara, California, such as an Intel® Xeon® processor. Although only one processing unit 120 is shown in FIG. Fig. 1, the device may include multiple processing units.

[0017] Each of the processor cores 122a-n may be connected to router block 160 via internal interconnect 112. Router block 160 and internal interconnect 112 may be generally representative of various circuits that support communication between components within processing unit 120, including buses, serial point-to-point connections, interconnect structures, and / or the like. Further details of such interconnects are not shown so as not to obscure the detail of operating environment 100.

[0018] Non-limiting examples of supported communication protocols may include Peripheral Component Interconnect (PCI) protocol, Peripheral Component Interconnect Express (PCIe or PCI-E) protocol, Universal Serial Bus (USB) protocol, Serial Peripheral Interface (SPI) protocol, Serial AT Attachment (SATA) protocol, Intel® QuickPath Interconnect (QPI) protocol, Intel® UltraPath Interconnect (UPI) protocol, Intel® Optimized Accelerator (OAP) protocol, Intel® Accelerator Link (IAL), Intra-Device Interconnect (IDI) protocol, Intel® On-Chip Scalable Fabric (IOSF) protocol, Scalable Memory Interconnect (SMI) protocol, SMI Generation 3 (SMI3), and / or the like. In some embodiments, interconnect 115 may support an intra-device protocol (e.g., IDI) and a memory interconnect protocol (e.g., SMI3).In various embodiments, the interconnect 115 may support an intra-device protocol (e.g., IDI), a memory interconnect protocol (e.g., SMI3), and a fabric-based protocol (e.g., IOSF).

[0019] As in Fig. 1, processing unit 120 may be communicatively coupled to logic device 180. In various embodiments, logic device 180 may include a hardware device. In some embodiments, logic device 180 may be implemented in hardware, software, or a combination thereof. In various embodiments, logic device 180 may include an accelerator. In some embodiments, logic device 180 may include a hardware accelerator. In various embodiments, logic device 180 may include an I / O device, such as an intelligent I / O device. In some embodiments, processing unit 120 may be coupled to multiple logic devices.Although an accelerator and an intelligent I / O device may be used as an example logic device 180, embodiments are not so limited, as logic device 180 may include any type of device, processor (e.g., a graphics processing unit (GPU)), logic unit, circuit, integrated circuit, application-specific integrated circuit (ASIC), FPGA, memory unit, compute unit, and / or the like suitable for operating in accordance with some embodiments. For example, in a multi-chip configuration, logic device 180 may include a processing unit that is the same, similar, or substantially similar to processing unit 120.

[0020] Logic device 180 may include or be associated with coherent agents 140n and one or more caches 184a-n. In some embodiments, at least one of caches 184a-n may include a giant cache type. In various embodiments, coherent agents 140a-n may be located on a same processor chip as coherent fabric 130 (e.g., coherent agent 140a) or on a companion chip coupled to the processor chip, for example, via a multi-chip package (MCP) or other off-package connection (e.g., coherent agent 140n).

[0021] As in Fig. 1, processing unit 120 may include a local cache configured as a hierarchy of, for example, various levels of caches 124a-n and last-level cache (LLC) 126. Generally, caches closest to processor cores 122a-n have the lowest latency and smallest size, and caches farther away are larger but have higher latency. For example, cores 122a-n may include caches 124a-n that may be deployed as a first-level (L1) cache or as an L1 cache and a second-level (L2) cache. Caches 124a-n, or portions thereof, within cores 122a-n may be "private" to each processor core 122a-n. In the context of operating environment 100, the highest-level cache is LLC 126.For example, the LLC 126 for a given core 122a-n may include a third-level (L3) cache if L1 and L2 caches are also deployed, or an L2 cache if the only other cache is an L1 cache. Embodiments are not limited in this regard, as the cache memory may be expanded to include additional cache levels.

[0022] Each of the processor cores 122a-n may also be connected to the LLC 126. The LLC 126 is in Fig. 1 as a single logical block; however, components related to LLC 126 may be distributed across processing unit 120 rather than implemented as a single monolithic block or box. In some embodiments, one or more of caches 124a-n and / or LLC 126 may include a small cache type. In some embodiments, one or more of caches 124a-n and / or LLC 126 may include a giant cache type. In various embodiments, caches 124a-n and LLC 126 may include small cache types, and at least one of caches 124a-n may include a giant cache type.Accordingly, the processing unit 120 may include a local cache formed from a private cache implemented via the caches 124a-n in each of the processor cores 122a-n and the LLC 126, which are implemented as a hierarchical layer sitting between a private core cache and system memory 156.

[0023] In various embodiments, LLC 126 may be implemented as a single block or box (e.g., shared across each of processor cores 122a-n) or a set of peer LLCs (e.g., shared across a subset of processor cores 122a-n). Each LLC 126 may be implemented as a monolithic agent or as a distributed set of sliced ​​agents. Each monolithic LLC or sliced ​​LLC agent may include or be associated with core / LLC coherence agents 114a-n configured to run in conjunction with cache coherence logic in each processor core 122a-n associated with a particular LLC peer of LLC 126. In some embodiments, the core / LLC coherence agent 114a-n may provide a coherence process to maintain cache coherence between an LLC peer and the associated processor cores 122a-n.In various embodiments, the core / LLC coherence agent 114a-n may provide cache coherence processes that are designed to be informed by LLC tag memories that track search lines stored in LLC data fields of the LLC 126, and snoop filter (SF) tag memories that track cache lines (e.g., at a granularity of 64B) stored in the caches 124a-n of the associated processor cores 122a-n. The LLC tags and / or the SF tags may be implemented using various types of memory, such as SRAM, on a processor chip of the processing unit 120. In general, a typical implementation of a high core count processor (e.g., 30-50 cores) in a 10 nm silicon processor chip may include a total of approximately 100 MB of LLC memory distributed across a number of LLC peers, and approximately100 MB of private core cache distributed across the cores (in one example including 30-50 cores, providing a private cache of approximately 1.5-2 MB of private cache per core).

[0024] Processing unit 120 may include caching agents 128 and / or coherent agents 140a-n. In some embodiments, coherent agents 140a-n may include an LLC / memory coherence agent disposed between a set of on-chip LLC peers and system memory 126. The LLC / memory coherence agent may provide cache coherence processes that, among other things, serve to maintain cache coherence between LLC peers of LLC 126. Coherence processes provided by the LLC / memory coherence agent may be informed by a caching status corresponding to the status captured by the LLC tag memory and the SF tag memory of the LLC peers of the processing chip of processing unit 120.

[0025] As in Fig. 1, in some embodiments, at least a portion of the caching agents 128 and / or the coherent agents 140a-n may be coupled to a coherent fabric 130, which may include various components, such as memory 134, storage (e.g., registers), interfaces, and cache coherence controller 132. In various embodiments, at least a portion of the caching agents 128 may be coupled to the coherent agents 140a-n without the use of a coherent fabric 130. During normal operation, the coherent fabric 130 may route traffic between system agents. The traffic may include snoops in response to cache requests. In some embodiments, the coherent fabric 130 may act as a primary on-chip interconnect between a plurality of different agents (e.g., the caching agents 128 and / or the coherent agents 140a-n) and other components.In various embodiments, the coherent fabric 130 may include logic for enforcing coherence and ordering on the fabric within the processing unit 120 based on a cache coherence process. Furthermore, the coherent fabric 130 may perform end-to-end routing of packets and / or other data and may also handle any associated functionality, such as network fairness, deadlock avoidance, traffic modulation, among other routing functions for agents of the processing unit 120.

[0026] In some embodiments, the caching agents 128 may be responsible for managing data delivery between the cores 122a-n and / or the logical device 180 and the shared cache 126 and / or the caches 184a-n. The caching agents 128 are also responsible for maintaining cache coherence between the cores 122a-n within a single socket (e.g., within the processing unit 120) and the logical device 180. This may include generating snoops and collecting snoop responses from the cores 122a-n according to a cache coherence protocol, such as MESI, MOSI, MOESI, or MESIF. Each of the multiple caching agents 128 may be assigned to manage a specific subset of the shared cache 126.

[0027] Caching agents 128 may include any agent capable of caching one or more cache blocks. For example, cores 122a-n may be caching agents, as may caches 124a-n. Coherent fabric 130 and / or cache coherence controller 132 may enable coherent communication within device 105. Coherent communication may be supported within device 105 during address, response, and data phases of transactions. Generally, a transaction is initiated by transmitting the address of the transaction in an address phase along with a command indicating which transaction is being initiated and various other control information. Coherent agents 140a-n use the response phase to maintain cache coherence.Each coherent agent 140a-n may respond with an indication of the status of the cache block addressed by the address and may also retry transactions for which a coherent response cannot be determined. Retrying transactions are aborted and may be retried later by the initiating agent. The order of successful (non-retrying) address phases may establish the order of transactions for coherence purposes (e.g., according to a sequentially consistent model). The data for a transaction is transferred in the data phase. Some transactions may not include a data phase. For example, some transactions may be used only to establish a change in the coherence status of a cached block.In general, the coherence state for a cache block can define the permissible operations the caching agent can perform on the cache block (e.g., read, write, and / or the like). Common coherence state schemes include MESI, MOESI, and variations of these schemes.

[0028] The coherent agents 140a-n (one or more of which may be referred to as a "home" agent) may be connected to memory controller 150. In some embodiments, the caching agents 128, one or more of the home agents 140a-n, and the memory controller 150 may work in concert to manage access to the system memory 156. The actual physical memory of memory 156 may be stored in one or more memory modules (not shown) and accessed by the memory controller 150 via a memory interface. For example, in one embodiment, a memory interface may include one or more Double Data Rate (DDR) interfaces, such as a DDR Type 3 (DDR3) interface. In some embodiments, the memory controller 150 may include a RAM memory controller (e.g., Dynamic RAM (DRAM)).

[0029] Coherent agents 140a-n may cooperate with caching agents 128 to manage the use of cache lines by various memory consumers (e.g., processor cores 122a-n, logic device 180, and / or the like). In particular, coherent agents 140a-n and caching agents 128, alone or through coherent fabric 130, may support a coherent memory scheme under which memory may be shared in a coherent manner without data corruption.To support this functionality, the coherent agents 140a-n and the caching agents 128 access and update cache line usage data stored in directory 152, which is logically represented in the memory controller 150 (e.g., some of the logic used by the memory controller 150 relates to using data of the directory 152), but whose data may be stored in the system memory 156 (see, for example, . Fig. 5).

[0030] The cache memory in device 105 may be maintained coherent using various cache coherence protocols, such as a snoop-based protocol, a directory-based protocol, combinations thereof, and / or combinations thereof, such as MESI, MOSI, MOESI, MESIF, Directory-Assisted Snoop (DAS), I / O Directory Cache (IODC), Home Snoop, Early Snoop, Snoop with Directory and Opportunistic Snoop Transfer, Cluster-on-Die, and / or the like. In general, a system memory address may be associated with a specific location in the system. This location may generally be referred to as the "home node" of the memory address. In a directory-based protocol, the caching agents 128 may send requests to the home node (for example, via the coherent structure 130) to access a memory address associated with a coherent agent (or home agent) 140a-n.The caching agents 128, in turn, are responsible for ensuring that the most recent copy of the requested data is returned to the requester, either from the storage 156 or from a caching agent 128 that owns the requested data. The coherent agent 140a-n may also be responsible for invalidating copies of data at other caching agents 128, for example, if the request is for an exclusive copy. For these purposes, a coherent agent 140a-n may generally snoop on each caching agent 128 or rely on the directory 152 to keep track of a set of caching agents 128 where data may be located.

[0031] In some embodiments, device 105 may provide coherency by using a cache coherency process. The cache coherency process may include a variety of cache coherency protocols, functions, data flows, and / or the like. In various embodiments, the cache coherency process may include a directory-based cache coherency protocol. In a directory-based cache coherency protocol, agents that protect memory, such as coherency agents 140a-n (or home agents), collectively maintain directory 152, which tracks where and in what state each cache line is cached in the system.Caching agents 128 seeking to acquire a cache line may send a request to coherent agents 140a-n (e.g., via coherent fabric 130), which perform a lookup in directory 152 and send messages commonly referred to as snoops only to those caching agents 128 for which directory 152 indicates that they may have cached copies of the cache line.

[0032] The directory 152 may store directory information (or coherency information). The memory controller 150 and / or the cache coherency controller 132 may maintain the directory status of each line in memory in the directory 152. In response to some transactions, the directory status of a corresponding line may change. The memory controller 150 and / or the cache coherency controller 132 may update records identifying the current directory status of each line in memory. In some cases, the memory controller 150 and / or the cache coherency controller 132 may receive data from a coherency agent 140a-n to indicate a change to the directory status of a line of memory.In other cases, the memory controller 150 and / or the cache coherence controller 132 may include logic to automatically update the directory status (e.g., without a host indicating the change) based on the nature of a corresponding request. In various embodiments, a directory read is performed by using a request address to determine whether or not the request address hits the directory 152 and, if the memory address hits, to determine the coherence status of the target memory block. The directory information, which may include a directory status for each cache line in the system, may be returned by the directory 152 in a subsequent cycle.

[0033] In various embodiments, directory information may include bits or other data, such as information stored in error correction code (ECC) bits of memory entries corresponding to requested data. In some embodiments, directory information may include 2 bits encoded as follows: invalid - no cache copies, clean - clean copies cached in a memory location, dirty - dirty copy cached in a memory location and not in use. In various embodiments, the directory information may include two sets of a two-bit directory status. In some embodiments, the directory information may include 2 bits or 4 bits stored with the data, for example, as part of the ECC bits.The number of bits for an entry in directory 152 may include 2 bits, 4 bits, 6 bits, 8 bits, 10 bits, 12 bits, 16 bits, 20 bits, 32 bits, and any value or range between any two of these values ​​(including endpoints). In some embodiments, the directory information may include various fields. Non-limiting examples of fields may include a tag field that identifies the real address of the memory block held in the corresponding cache line, a status field that indicates the coherency status of the cache line, inclusivity bits that indicate whether the memory block is held in the associated L1 cache, and cache status information.In some embodiments, the directory may include a status (or status information) to identify the exact one cache that has a copy of data if only one copy exists, a specific set of caching agents in which a copy of data can be found, a status (or status information) to indicate a number of copies (e.g., 1 copy, 2 copies, multiple copies, and / or the like), and / or an indicator of the type of caching agent.

[0034] In some embodiments, the directory information may include cache status or cache status information. In some embodiments, the cache status information may indicate a type of memory operation, cache coherence process, coherence operation, and / or the like. For example, the cache status information may indicate whether a cache coherence process should support a huge cache type, a small cache type, and / or the like. In this way, the directory 152 may support tracking the coherence of multiple classes of agents (e.g., coherent agents or home agents). Non-limiting classes of agents include huge cache agents (high-latency agents) designed to support coherence for huge cache types, and cache agents (low-latency agents) designed to support coherence for small cache types.In some embodiments, huge cache agents and / or small cache agents may include coherent agents 140a-n (or home agents). In various embodiments, a cache coherency process may include a small cache coherency process and a huge cache coherency process. Device 105 or components thereof may follow the small cache coherency process in response to the cache status indicating a small cache status, and device 105 or components thereof may follow a huge cache coherency process in response to the cache status indicating a huge cache status.

[0035] In some embodiments, a giant cache agent may include a coherent agent 140a-n that serves as an agent for a giant cache type. For example, coherent agent 140n may be a giant cache agent for a giant cache of logical device 180. In another example, coherent agent 140a may be a small cache agent for a small cache of core 122a (e.g., cache 124a). In some embodiments, processing unit 120 may include a giant cache. In some embodiments, logical device 180 may include a giant cache. In some embodiments, processing unit 120 and logical device 180 may include a giant cache. Embodiments are not limited in this context.

[0036] In various embodiments, the cache status information may include or use additional directory bits per directory entry to encode a cache status (e.g., a small cache status or a huge cache status) per directory entry. For example, the cache status information may include 2 additional directory bits per directory entry to provide a 2-bit directory encoding field per agent class per directory entry. In some embodiments, 4 bits or encodings (such as common Xeon® directory encodings) may track huge cache status, two bits or encodings may track small cache status, and the remaining bits or encodings may track "uncached" statuses.

[0037] In some embodiments, device 105 may be arranged in a multi-chip (or multi-socket) configuration, in which, for example, logic device 180 is a processor, processing unit, processor chip, CPU, and / or the like similar or substantially similar to processing unit 120. In some embodiments, logic device 180 may be or include multiple companion processor chips, such as 2, 4, or 8 chips. In some embodiments where the logic device is a companion processor chip, all local processor sets of peer LLCs may be combined to create a single large set of LLC peers across all chips. In such embodiments, LLC / memory coherence agents may be responsible for maintaining coherence across all peers by using cache coherence processes, according to some embodiments.For example, LLC / memory coherence agents may use a UPI, UPI-based, or UPI-derived cache coherence process according to some embodiments.

[0038] In a multi-chip configuration, LLC / memory coherence agents (or the cache coherence processes executed thereby) must be informed by the same equivalent of the LLC tag and SF tag information, which is the same or similar to that described above. However, in a multi-chip configuration, each cache coherence process on each LLC / memory coherence agent on each chip must be informed by the equivalent of the LLC status and SF status on each chip in the configuration.

[0039] In a conventional system, if each chip of a multi-chip system contained the tag memories for every other chip, the size of the total tag memory per chip could grow by a factor of 8 (for an 8-processor chip architecture), forcing either the chip to grow or reducing the amount of SRAM that could be used as cache (and adding additional power to each chip). Also, in terms of scalability, if a chip was built to support up to 8 chip configurations but was also used in 2- and 4-chip configurations, half or three-quarters of the total tag state (and allocated area) would be wasted in 4- and 2-chip configurations, respectively (for example, would not be used for cache or other functional uses).

[0040] Because of such scalability and cost concerns, most solutions known in the art use various different schemes, such as tag (directory) caching, coarse-grain coherence, full-directory caching, and / or the like. In tag caching, each chip maintains a relatively small cache of the peer LLC's tag state. Tag caching can address certain aspects of the cost and scalability problem, but it limits the amount of data that can be cached in other chips and constrains performance, while still requiring the addition of additional tag state on the processor chip. In coarse-grain coherence, coherence between peer LLCs can be maintained by, for example, using 1 KB cache lines instead of the 64B used between the LLC and the cores.In this model, cache misses are more expensive (requiring the full 1KB data block to be moved), and the chances of cache conflicts (hot sets) are much greater. Both of these phenomena create performance obligations, while still requiring an on-chip trace structure, only a reduced-size structure. In full directory schemes, trace state is stored in DRAM on a per-memory-line basis (compared to per-cache-line basis). This solution adds memory read latency to some operations to allow access to the trace state, which has some performance impact, but it requires no additional on-chip structure, allows caching at 64B granularity, and places no limits on how much data is cached in companion processor chips.In some embodiments, the device may implement a form of full directory caching by using a single directory state to represent the tag state of companion chips (such as companion chips 1 through 7).

[0041] As described above, in some embodiments, logic device 180 may be an accelerator, accelerator chip, or the like. If logic device 180 is a companion chip in the form of a coherent accelerator chip, the chip may be connected to the cache hierarchy in the manner of a processor core 122a-n participating in the cache coherence process managed by core / LLC coherence agents 114a-n, and its cached lines are tracked in the same snoop filter as processor chip cores 122a-n. This is advantageous because it allows the coherent accelerator to use the processors LLC (or an LLC peer), and it allows the processor core's snoop filter to be used to filter snoops to the accelerator as well, which can be particularly beneficial when the accelerator device has poor snoop latency.

[0042] The above arrangement works well when the size of the accelerator's cache memory (e.g., caches 184a-n of logic device 180) is the same size as the cache 124a-n of processor core 122a-n. In such a case, each accelerator attachment point requires an amount of snoop filter capacity similar to that of a single core to be added to the SF of the associated LLC peer. This is a relatively small adder and requires a relatively low cost for the entire chip.However, such a configuration can become problematic, similar to the problems with multi-socket systems, when the accelerator cache is an order of magnitude larger than the cache of a processor core (for example, 10 MB), and can become inoperable or essentially inoperable when the accelerator cache grows by several orders of magnitude (for example, 10 GB), which has become possible with the advent of memory technologies such as multi-channel DRAM (MCDRAM) and high-bandwidth memory (HBM).

[0043] Accordingly, the apparatus 105 and / or components thereof may be configured to operate cache coherence processes according to some embodiments to, among other things, address the problem of providing coherence state tracking for a chip / agent that interfaces with the cache hierarchy in the core / LLC coherence agent layer and that has a cache that is too large to be tracked in the LLC peer's allocated SF.For example, according to some embodiments, cache coherence processes may use a mechanism deployed by the LLC / memory coherence agent in a full-directory or full-directory-based scheme with a second field to the full-directory in memory, where the first field tracks the caching of a memory line in an LLC peer to inform the LLC / memory coherence agent, and where the second field tracks the caching of an accelerator agent to inform the core / LLC coherence agent. In various embodiments, one of each field type may be included, such that all LLC peers on all companion processor chips are tracked by a single instance of the first field type, and all accelerators attached to all companion processors in a multi-processor chip system are tracked by a single instance of the second field type.In some embodiments, N of each field type may be implemented such that each instance of each field type may be mapped to exactly one logic device 180 (e.g., exactly one companion processor chip, peer LLC, and / or accelerator chip).

[0044] Included herein are one or more logical flows representative of example methodologies for carrying out new aspects of the disclosed architecture. While, for the sake of simplicity of explanation, the one or more methodologies shown herein are shown and described as a series of operations, those skilled in the art will understand and appreciate that the methodologies are not limited by the order of operations. Accordingly, some operations may occur in a different order and / or concurrently with other operations than those shown and described herein. For example, those skilled in the art will understand and appreciate that a methodology could alternatively be represented as a series of interrelated states or events, such as in a state diagram. Furthermore, not all operations shown in a methodology may be required for a novel implementation.

[0045] A logic flow may be implemented in software, firmware, hardware, or any combination thereof. In software and firmware embodiments, a logic flow may be implemented by computer-executable instructions stored on a non-transitory computer-readable medium or a machine-readable medium, such as optical, magnetic, or semiconductor storage. The embodiments are not limited in this context.

[0046] Fig. 2 illustrates one embodiment of a logic flow 200. Logic flow 200 may be representative of some or all of the operations performed by one or more embodiments described herein, such as apparatus 105, 405, and / or 505 and / or components thereof. In some embodiments, logic flow 200 may be representative of some or all of the operations for a multi-processor chip, single-processor core read request.

[0047] As in Fig. 2, at block 202, logic flow 200 may issue a read. For example, one of cores 122a-n may issue a data read. At block 204, logic flow 200 may look up tags for the read. For example, the read process may use the core / LLC memory coherence agent to look up the LLC and SF tags. At block 206, logic flow 200 may determine whether there are copies of the available data. For example, the core / LLC memory coherence agent may determine, based on the LLC tags and the SF tags, whether there are copies of a line cache in an LLC peer or a child processor core. If logic flow 200 determines at block 206 that there are available copies of the data, logic flow 200 may return the data at block 216.

[0048] If logic flow 200 determines at block 206 that there are available copies of the data, logic flow 200 may issue a read to memory at block 210. For example, the core / LLC memory coherence agent may issue a read to DRAM. At block 212, logic flow 200 may return a memory line and a trace field. For example, DRAM may return a memory line of data and a full directory trace field. Logic flow 200 may determine at block 214 whether a copy of the data is available in a companion chip. For example, the full directory field may indicate that a companion processor chip may have a copy of the memory line. If logic flow 200 determines at block 214 that a copy of the data is available in a companion chip, logic flow 200 may issue a snoop to all processor chips at block 220.For example, the LLC / memory coherence agent may issue a snoop to all companion processor chips. If logic flow 200 determines at block 214 that a copy of the data is unavailable in a companion chip, logic flow 200 may return the data to the requesting core at block 216 and update the fields and tags at block 218. For example, the full directory, LLC tags, and SF tags may be updated to reflect that the memory line is now cached in the requesting core.

[0049] Fig. 3 illustrates one embodiment of a logic flow 300. Logic flow 300 may be representative of some or all of the operations performed by one or more embodiments described herein, such as apparatus 105, 405, and / or 505 and / or components thereof. In some embodiments, logic flow 300 may be representative of some or all of the operations for a multi-processor chip single-core processor read request involving at least one accelerator chip.

[0050] As in Fig. 3, at block 302, logic flow 300 may issue a read. For example, one of cores 122a-n may issue a data read. At block 304, logic flow 300 may look up tags for the read. For example, the read process may use the core / LLC memory coherence agent to look up the LLC and SF tags. At block 306, logic flow 300 may determine whether there are copies of the available data. For example, the core / LLC memory coherence agent may determine, based on the LLC tags and the SF tags, whether there are copies of a line cache in an LLC peer or a child processor core. If logic flow 300 determines at block 306 that there are available copies of the data, logic flow 300 may return the data at block 322.

[0051] If logic flow 300 determines at block 306 that there are available copies of the data, logic flow 300 may issue a read to memory at block 310. For example, the core / LLC memory coherence agent may issue a read to DRAM. At block 312, logic flow 300 may return a memory line and a trace field. For example, DRAM may return a memory line of data and a full directory trace field. Logic flow 300 may determine at block 314 whether a copy of the data is available in a companion chip. For example, the full directory field may indicate that a companion processor chip may have a copy of the memory line. If logic flow 300 determines at block 314 that a copy of the data is available in a companion chip, logic flow 300 may issue a snoop to all processor chips at block 315.For example, the LLC / memory coherence agent may issue a snoop to all companion processor chips. If logic flow 300 determines at block 314 that a copy of the data is unavailable in a companion chip, logic flow 300 may determine at block 318 whether a copy of the data is available in an accelerator chip. If logic flow 300 determines at block 318 that a copy of the data is available in an accelerator chip, logic flow 300 may issue a snoop to all accelerator chips at block 320. If logic flow 200 determines at block 314 that a copy of the data is not available in a companion processor chip and that a copy of the data is not available in a companion accelerator chip at block 314, logic flow 300 may return the data to the requesting core at block 322 and update the fields and tags at block 324.

[0052] Fig. 4 illustrates an example of an operating environment 400 that may be representative of various embodiments. In Fig. Operating environment 400 illustrated in Figure 4 may include apparatus 405 with processing unit 420 and logic device 480. In some embodiments, the processing unit may be a CPU and / or may be implemented as a SoC. In various embodiments, logic device 480 may include an accelerator, an I / O device, an intelligent I / O device, and / or the like. Logic device 480 may include attached or direct-attached memory. In various embodiments, logic device 480 may use at least a portion of the direct-attached memory as one or more caches. For example, logic device 480 may include and / or be operatively coupled to direct-attached memory serving as one or more caches, such as small cache 484a and / or huge cache 484b.In some embodiments, the small cache 484a may include a small cache type and the huge cache 484b may include a huge cache type.

[0053] Logic device 480 may be coupled to processing unit 420 via one or more interfaces, including, without limitation, a coherency interface or a bus (e.g., a "request bus") 406. In some embodiments, logic device 480 may be coupled directly to the processing unit via bus 406. In various embodiments, logic device 480 may be coupled to processing unit 420 via bus 406 via a coherency fabric (such as coherency fabric 130). In some embodiments, logic device 480 may be in communication with coherency controller 422 of processing unit 420. Coherency controller 422 may include various components to implement cache coherency in device 405. For example, coherence controller 422 may include home agent 428, coherence engine 424, snoop filter 426 (or directory), and / or the like.Processing unit 420 may include request processing agent 430 configured to process memory operation requests from logic device 480 or other elements (e.g., caching agents 128). In some embodiments, request processing agent 430 may be a component of a coherent fabric (e.g., coherent fabric 130), serving, for example, as an interface for interconnect 406. In general, a memory operation request may include a request for processing unit 420 to perform a memory operation (e.g., read, write, and / or the like).

[0054] In some embodiments, when a request for a store operation arrives at coherency controller 422, coherency engine 424 may determine where to forward the request. A store operation may generally refer to a transaction requesting access to host memory 456 or one of caches 484a, 484b and / or a processing unit cache (not shown). In some embodiments, host memory 456 may include directory 452. Coherency engine 424 may perform a lookup of snoop filter 426 or directory 452 to determine if snoop filter 426 or directory 452 has directory information of the requested line. In some embodiments, snoop filter 426 may be a "sparse" or "partial" directory used to maintain a partial directory (e.g., of directory 452).For example, directory information in snoop filter 426 may track all caches in the system, while directory 452 may track all or substantially all of the physical address space in the system. For example, directory information in snoop filter 426 may indicate caches (such as cache 484a, 484b, cache 124a-n, cache 126, and / or the like) that have a copy of a cache block associated with a memory operation request. In general, snoop filters may be used to avoid sending unnecessary snoop traffic, thereby removing snoop traffic from the critical path and reducing the amount of traffic and cache activity in the overall system.In various embodiments, the snoop filter 426 and / or the directory 452 may be an array organized in a set-associative manner, where each path entry may contain one or more valid bits (depending on the configuration—either line- or region-based), agent class bits, and a tag field. In one embodiment, a valid bit indicates whether the corresponding entry has valid information, the agent status bits are used to indicate whether the cache line is present in the corresponding cache agent, and the tag field is used to store the cache line's tag. In one embodiment, the agent class bits are configured to indicate a class of agent associated with a cache line and / or memory operation.

[0055] If the snoop filter 426 and / or the directory 452 have directory information, the coherence engine 424 forwards the request to the cache that has a current copy of the line based on the line's presence vector. If the transaction may potentially change the status of the requested line, the coherence engine 424 updates the directory information in the snoop filter 426 and / or the directory 452 to reflect the changes. If the snoop filter 426 and / or the directory 452 does not have information for the line, the coherence engine 424 may add an entry to the snoop filter 426 and / or the directory 452 to record directory information of the requested line. In some embodiments, directory information may include cache status information to indicate a type of memory operation, cache coherence process, coherence operation, and / or the like.For example, the cache status information may indicate whether a cache coherence process should support a huge cache type, a small cache type, and / or the like. In some embodiments, the logic device 480 and / or the coherence controller 422 may be operable to specify the cache status information associated with an entry in the snoop filter 426 and / or the directory 452. For example, entries in the snoop filter 426 and / or the directory 452 associated with cache lines of the small cache 484a may include cache status information indicating a small cache status. In another example, entries in the snoop filter 426 and / or the directory 452 associated with cache lines of the huge cache 484b may include cache status information indicating a huge cache status.In this way, the logic device 480 in combination with the coherence controller 422 can support both the small cache 484a and the giant cache 484b by using the same hardware and / or software architecture.

[0056] In some embodiments, agent class qualifiers may be included and / or determined from memory requests of logic device 480 on bus 406 (as a request bus). In various embodiments, agent class qualifiers on bus 406 during a read request or a write request from logic device 480 may specify which of two levels of coherency support are required by the logic device for a particular request, including support for a huge cache type or support for a small cache type.Accordingly, in various embodiments, memory operation requests generated by logic device 480, processor cores (e.g., processor cores 122a-n), caching agents 128, and / or other elements may include agent class qualifiers configured to specify the agent class to be used for the particular memory operation. In some embodiments, elements sending memory requests, such as logic device 480 (e.g., serving as an accelerator, I / O device, and / or the like), may modulate the active class on a request-by-request basis, allowing a single logic device 480 to employ one or more small caches and one or more giant caches.

[0057] In some embodiments, a data requester (e.g., logic device 480) may transmit a data request over bus 406. The data request may include data request information, such as data, a memory address, cache status information, and / or the like. In various embodiments, the cache status information may indicate an agent class associated with the request, such as whether the request is associated with a huge cache status or a small cache status. For example, store requests from logic device 480 associated with small cache 484a may include cache status information indicating a small cache status, and store requests from logic device 480 associated with huge cache 484b may include cache status information indicating a huge cache status.

[0058] The request processing agent 430 may receive or otherwise access memory requests and determine cache status information from the memory requests. In some embodiments, for requests indicating a small cache status, the request processing agent 430 may execute the coherence controller 422, cause it to execute, and / or instruct it to operate a small cache coherence process, for example, by using flows that invoke or use the on-chip snoop filter 426 of the processing unit 420 and the directory field associated with the small cache status.In various embodiments, for requests indicating a huge cache status, the request processing agent 430 may execute, cause the coherence controller 422 to execute, and / or instruct the coherence controller 422 to operate a huge cache coherence process and instruct the coherence controller 422 to operate using flows that use the directory field associated with the huge cache status, for example, the directory 452.

[0059] In some embodiments, in response to a cache line access directory status, the coherence controller 422 may forward snoop messages directly to the logic device 480 in response to determining a huge cache status (and therefore the huge cache coherence process is active) associated with a store operation, and may indicate that the logic device 480 may have a cached copy of the associated cache line. In some embodiments, the coherence controller 422 may forward snoop messages directly to the on-chip processor snoop filters 426 in response to a cache line access directory status if the small cache status indicates that a small cache agent has a cached copy of the cache line associated with a store operation (and therefore the small cache coherence process is active).In some embodiments, the coherence controller 422 at the on-chip snoop filter 426 may forward snoops to the logic device 480 when the on-chip snoop filter indicates that the logic device 480 may have a cached copy of the cache line associated with a memory operation.

[0060] Fig. 5 illustrates an example of operating environment 500 that may be representative of various embodiments. Fig. Operating environment 500 illustrated in Figure 5 may include device 505 having system memory 510 for storing directory 520. Directory 520 may include one or more entries 530a-n. In some embodiments, each entry 530a-n may correspond to a cache line of data stored in a cache of device 505 (e.g., caches 124a-n, LLC 126, caches 184a-n, and / or the like). Each entry 530a-n may include directory information, such as cache line information 540a-n (e.g., an associated cache line and / or physical memory address indicating entry 530a-n), coherency status 542a-n (e.g., invalid, locked, owned, modified, exclusive, shared, forward, and / or the like), and / or cache status 544a-n.In various embodiments, cache status 544a-n may indicate whether a cache entry 530a-n is associated with a small cache status (and therefore a small cache type) or a huge cache status (and therefore a huge cache type).

[0061] In some embodiments, for each memory operation, a coherency controller (such as coherency controller 422) may access entries 530a-n of directory 520 to determine the cache status. For example, the coherency controller may determine the cache status to execute the corresponding cache coherency process. For example, for a read operation, the coherency controller may execute a huge cache coherency process in response to determining that cache status 544a-n is a huge cache status, and the coherency controller may execute a small cache coherency process in response to determining that cache status 544a-n is a small cache status.In another case, for a write operation, the cache state may be determined from the store operation request, and the corresponding cache state 544a-n may be written to a new entry 530a-n if no previous entry exists for the associated cache line, or the cache state 544a-n may be updated to match the cache state specified in the store operation request.

[0062] Included herein are one or more logical flows representative of example methodologies for carrying out new aspects of the disclosed architecture. While, for the sake of simplicity of explanation, the one or more methodologies shown herein are shown and described as a series of operations, those skilled in the art will understand and appreciate that the methodologies are not limited by the order of operations. Accordingly, some operations may occur in a different order and / or concurrently with other operations than those shown and described herein. For example, those skilled in the art will understand and appreciate that a methodology could alternatively be represented as a series of interrelated states or events, such as in a state diagram. Furthermore, not all operations shown in a methodology may be required for a novel implementation.

[0063] A logic flow may be implemented in software, firmware, hardware, or any combination thereof. In software and firmware embodiments, a logic flow may be implemented by computer-executable instructions stored on a non-transitory computer-readable medium or a machine-readable medium, such as optical, magnetic, or semiconductor storage. The embodiments are not limited in this context.

[0064] Fig. 6 illustrates one embodiment of a logical flow 600. The logical flow 600 may be representative of some or all of the operations performed by one or more embodiments described herein, such as device 105, 405, and / or 605. In some embodiments, the logical flow 600 may be representative of some or all of the operations for a cache coherence process according to some embodiments.

[0065] As in Fig. 6, at block 602, logic flow 600 may receive a memory operation request. For example, request processing agent 430 may receive or otherwise access memory requests and determine cache status information from the memory requests. At block 604, logic flow 600 may determine a cache status associated with the memory operation request. For example, request processing agent 430 may determine whether the memory operation requests are associated with a huge cache status or a small cache status.In various embodiments, agent class qualifiers on bus 406 during a read request or a write request from logic device 480 may specify which of two levels of coherency support are required by the logic device for a particular request, including support for a huge cache type or support for a small cache type. In various embodiments, memory operation requests generated by logic device 480, processor cores (e.g., processor cores 122a-n), caching agents 128, and / or other elements may include agent class qualifiers configured to specify the agent class to be used for the particular memory operation.

[0066] If the determination block 606 determines that the cache status is a small cache status, the logic flow 600 may invoke the snoop filter at block 608 and use the directory field associated with the small cache status. For example, for requests indicating a small cache status, the request processing agent 430 may instruct the coherence controller 422 to operate using flows that invoke the on-chip snoop filter 426 of the processing unit 420 and the directory field associated with the small cache status.

[0067] If determination block 606 determines that the cache status is a huge cache status, logic flow 600 may invoke the directory field associated with the huge cache status. For example, for requests indicating a huge cache status, request processing agent 430 may instruct coherence controller 422 to operate using flows that invoke the directory field associated with the huge cache status, such as directory 452.

[0068] Fig. 7 illustrates one embodiment of a logical flow 700. The logical flow 700 may be representative of some or all of the operations performed by one or more embodiments described herein, such as device 105, 405, and / or 505. In some embodiments, the logical flow 700 may be representative of some or all of the operations for a cache coherence process according to some embodiments.

[0069] As shown in FIG. 7, at block 702, logic flow 700 may determine a cache status associated with a cache line access. For example, in some embodiments, for each store operation, a coherence controller (such as coherence controller 422) may access entries 530a-n of directory 520 to determine the cache status. If determination block 704 determines that the cache status is a small cache status, logic flow 700 may determine at determination block 706 whether a small cache agent may have a cached copy of a cache line associated with the store operation. If a small cache agent has a copy of the cache line, as indicated by determination block 706, snoop messages may be forwarded to the snoop filter at block 708.For example, in response to a cache line access directory status, the coherence controller 422 may forward snoop messages directly to the on-chip processor snoop filters 426 if the small cache status indicates that a small cache agent has a cached copy of the cache line associated with a store operation. If the snoop filter indicates that a giant cache agent of a logical device may have a cached copy of the cache line associated with a store operation at the determination block 712, snoop messages may be forwarded to the logical device at block 714. For example, the coherence controller 422 at the on-chip snoop filter 426 may forward snoops to the logic device 480 when the on-chip snoop filter indicates that the logic device 480 may have a cached copy of the cache line associated with a memory operation.

[0070] If determination block 704 determines that the cache status is a giant cache status, logic flow 700 may determine at determination block 712 whether a logical device cache (e.g., a giant cache device) may have a cached copy of the cache line associated with the store operation. If logic flow 700 determines at block 712 that a logical device cache may have a cached copy of the cache line, logic flow 700 may forward snoop messages to a logical device at block 714. For example, coherence controller 422 at on-chip snoop filter 426 may forward snoops to logic device 480 if the on-chip snoop filter indicates that logic device 480 may have a cached copy of the cache line associated with a store operation.

[0071] Fig. 8 illustrates an example of a storage medium 800. The storage medium 800 may comprise an article of manufacture. In some examples, the storage medium 800 may include any non-transitory computer-readable medium or machine-readable medium, such as optical, magnetic, or semiconductor storage. The storage medium 800 may store various types of computer-executable instructions, such as instructions for implementing the logic flow 100, 200, 600, and / or the logic flow 700. Examples of a computer-readable or machine-readable storage medium may include any tangible media capable of storing electronic data, including volatile memory or non-volatile memory, removable or non-removable memory, erasable or non-erasable memory, writable or rewritable memory, and so on.Examples of computer-executable instructions may include any suitable type of code, such as source code, compiled code, translated code, executable code, static code, dynamic code, object-oriented code, visual code, and the like. The examples are not limited in this context.

[0072] Fig. 9 illustrates one embodiment of an exemplary computer architecture 900 suitable for implementing various embodiments as described above. In various embodiments, computer architecture 900 may include or be implemented as part of an electronic device. In some embodiments, computer architecture 900 may be representative of, for example, device 105, 405, and / or 505. The embodiments are not limited in this context.

[0073] As used in this application, the terms "system," "component," and "module" are intended to refer to a computer-related entity, either hardware, a combination of hardware and software, software, or software in execution, examples of which are provided by example computer architecture 900. For example, a component may be, but is not limited to, a process running on a processor, a processor, a hard disk drive, multiple storage drives (of optical and / or magnetic storage media), an object, an executable file, a thread of execution, a program, and / or a computer. For illustrative purposes, both an application running on a server and the server may be a component.One or more components may reside within a process and / or thread of execution, and a component may be localized on one computer and / or distributed between two or more computers. Further, components may be communicatively coupled to one another through various types of communication media to coordinate operations. The coordination may involve the unidirectional or bidirectional exchange of information. For example, the components may communicate information in the form of signals communicated over the communication media. The information may be implemented as signals assigned to different signal lines. In such assignments, each message is a signal. However, other embodiments may alternatively employ data messages. Such data messages may be sent over various connections.Example connections include parallel interfaces, serial interfaces, and bus interfaces.

[0074] Computer architecture 900 includes various common computer elements, such as one or more processors, multi-core processors, co-processors, memory units, chipsets, controllers, peripherals, interfaces, oscillators, timing devices, video cards, audio cards, multimedia input / output (I / O) components, power supplies, and so forth. However, embodiments are not limited to implementation by computer architecture 900.

[0075] As in Fig. 9, the computer architecture 900 includes processing unit 904, system memory 906, and system bus 999. The processing unit 904 may be any of various commercially available processors, including, but not limited to, AMD® Athlon®, Duron®, and Opteron® processors; ARM® application, embedded, and secure processors; IBM® and Motorola® DragonBall® and PowerPC® processors; IBM and Sony® Cell processors; Intel® Celeron®, Core(2) Duo®, Itanium®, Pentium®, Xeon®, and XScale® processors; and similar processors. Dual microprocessors, multi-core processors, and other multi-processor architectures may also be employed as the processing unit 904.

[0076] System bus 999 provides an interface for system components, including, but not limited to, system memory 906 to processing unit 904. System bus 999 may be any of various types of bus structures, which may further interface with a memory bus (with or without a memory controller), a peripheral bus, and a local bus using any of several commercially available bus architectures. Interface adapters may connect to system bus 999 via a slot architecture. Example slot architectures may include, without limitation, Accelerated Graphics Port (AGP), Cardbus, (Extended) Industry Standard Architecture ((E)ISA), Micro-Channel Architecture (MCA), NuBus, Peripheral Component Interconnect (Extended) (PCI(X)), PCI Express, Personal Computer Memory Card International Association (PCMCIA), and the like.

[0077] The system memory 906 may include various types of computer-readable storage media in the form of one or more higher-speed memory devices, such as read-only memory (ROM), random-access memory (RAM), dynamic RAM (DRAM), double-data-rate DRAM (DDRAM), synchronous DRAM (SDRAM), static RAM (SRAM), programmable ROM (PROM), erasable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory, polymer memory, such as ferroelectric polymer memory, ovonic memory, phase-change or ferroelectric memory, silicon oxide nitride oxide silicon (SONOS) memory, magnetic or optical cards, a variety of devices such as redundant array of independent disks (RAID) drives, solid-state storage devices (e.g., USB memory, solid-state drives (SSD)), and any other type of storage media suitable for storing information. In the illustrated Fig.In the embodiment shown in Figure 9, system memory 906 may include non-volatile memory 910 and / or volatile memory 912. A basic input / output system (BIOS) may be stored in non-volatile memory 910.

[0078] Computer 902 may include various types of computer-readable storage media in the form of one or more lower-speed storage devices, including internal (or external) hard disk drive (HDD) 914, magnetic floppy disk drive (FDD) 916 for reading from or writing to removable magnetic disk 919, and optical disk drive 920 for reading from or writing to removable optical disk 922 (e.g., a CD-ROM or DVD). HDD 914, FDD 916, and optical disk drive 920 may be connected to system bus 999 through HDD interface 924, FDD interface 926, and optical drive interface 928, respectively. HDD interface 924 for external drive implementations may include at least one or both Universal Serial Bus (USB) and IEEE 1394 interface technologies.

[0079] The drives and associated computer-readable media provide volatile and / or non-volatile storage of data, data structures, computer-executable instructions, and so forth. For example, a number of program modules may be stored in the drives and storage units 910, 912, including the operating system 930, one or more application programs 932, other program modules 934, and program data 936. In one embodiment, the one or more application programs 932, the other program modules 934, and the program data 936 may include, for example, the various applications and / or components of the device 105, 305, 405, and / or 505.

[0080] A user can enter commands and information into computer 902 through one or more wired / wireless input devices, for example, keyboard 938 and a pointing device such as mouse 940. Other input devices may include microphones, infrared (IR) remote controls, radio frequency (RF) remote controls, game consoles, stylus pens, card readers, dongles, fingerprint readers, gloves, graphics tablets, joysticks, keyboards, retina readers, touchscreens (e.g., capacitive, resistive, etc.), trackballs, trackpads, sensors, styluses, and the like. These and other input devices are often connected to the processing unit 904 through input device interface 942, which is coupled to the system bus 999, but they may be connected through other interfaces, such as a parallel port, IEEE 1394 serial port, a gaming port, a USB port, an IR interface, and so on.

[0081] Monitor 944 or another type of display device is also connected to system bus 999 via an interface, such as video adapter 946. Monitor 944 may be internal or external to computer 902. In addition to monitor 944, a computer typically includes other peripheral output devices, such as speakers, printers, and so on.

[0082] Computer 902 may operate in a network environment by using logical connections via wired and / or wireless communication with one or more remote computers, such as remote computer 948. Remote computer 948 may be a workstation, server computer, router, personal computer, portable computer, microprocessor-based entertainment application, peer device, or other common network node, and typically includes many or all of the described elements in association with computer 902, although for the sake of brevity, only memory / storage device 950 is illustrated. The illustrated logical connections include wired / wireless connectivity to local area network (LAN) 952 and / or larger networks, for example, wide area network (WAN) 954.Such LAN and WAN network environments are commonplace in offices and companies and enable company-wide computer networks, such as intranets, all of which can connect to a global communications network, such as the Internet.

[0083] When used in a LAN network environment, computer 902 is connected to LAN 952 through a wired and / or wireless communications network interface or adapter 956. Adapter 956 may enable wired and / or wireless communications to LAN 952, which may also include a wireless access point disposed thereon to communicate with the wireless functionality of adapter 956.

[0084] When used in a WAN network environment, computer 902 may include modem 959, or be connected to a communications server in WAN 954, or have other means for establishing communications over WAN 954, for example, over the Internet. Modem 959, which may be internal or external and a wired and / or wireless device, connects to system bus 999 via input device interface 942. In a network environment, program modules shown in association with computer 902 or portions thereof may be stored in remote memory / storage device 950. It is understood that the network connections shown are exemplary, and other means for establishing a communication link between the computers may be used.

[0085] The computer 902 is designed to communicate with wired and wireless devices or units using the IEEE 902 family of standards, such as wireless devices operatively arranged in wireless communication (e.g., IEEE 802.16 radio modulation methods). This includes, among others, Wi-Fi (or Wireless Fidelity), WiMax, and Bluetooth™ wireless technologies. Thus, communication can be a predefined structure, as in a conventional network, or simply ad hoc communication between two or more devices. Wi-Fi networks use radio technologies known as IEEE 802.11x (a, b, g, n, etc.) to provide secure, reliable, and high-speed wireless connectivity. A Wi-Fi network can be used to connect computers to each other, to the Internet, and to wired networks (using IEEE 802.3-related media and features).

[0086] One or more aspects of at least one embodiment may be implemented by representative instructions stored on a machine-readable medium representing various logic within the processor that, when read by a machine, causes the machine to fabricate logic for performing the methods described herein. Such representations, known as "IP cores," may be stored on a tangible, machine-readable medium and delivered to various customers or manufacturing facilities for loading into the manufacturing machines that actually manufacture the logic or processor.For example, some embodiments may be implemented using a machine-readable medium or article that can store an instruction or set of instructions that, when executed by a machine, can cause the machine to perform a method and / or operations according to the embodiments. Such a machine may, for example, include any suitable processing platform, computing platform, computing device, processing device, computing system, processing system, computer, processor, or the like, and may be implemented using any suitable combination of hardware and / or software.The machine-readable medium or article may, for example, include any suitable type of storage unit, storage device, storage article, storage medium, storage device, storage article, storage medium and / or storage unit, for example, memory, removable or non-removable media, erasable or non-erasable media, writable or rewritable media, digital or analog media, hard disk, floppy disk, compact disk read-only memory (CD-ROM), compact disk recordable (CD-R), compact disk rewritable (CD-RW), optical disk, magnetic media, magneto-optical media, removable memory cards or floppy disks, various types of digital versatile disk (DVD), a tape, a cassette, or the like.The instructions may include any suitable type of code, such as source code, compiled code, interpreted code, executable code, static code, dynamic code, encrypted code, and the like, implemented using any suitable high-level, low-level, object-oriented, visual, compiled and / or interpreted programming language. Example 1 is an apparatus comprising at least one processor, at least one cache, and logic embodied at least partially in hardware, wherein the logic is to receive a memory operation request associated with the at least one cache, determine a cache status of the memory operation request, the cache status indicating either a huge cache status or a small cache status, execute the memory operation request via a small cache coherence process in response to a determination that the cache status is a small cache status, and execute a memory operation request via a huge cache coherence process responsive to determining that the cache status is a small cache status. Example 2 is the apparatus of Example 1, wherein the at least one cache memory comprises at least one small cache and at least one giant cache. Example 3 is the apparatus of Example 1, wherein the at least one cache memory comprises at least one giant cache having a memory size greater than 1 gigabyte (GB). Example 4 is the apparatus of Example 1, wherein the at least one cache memory comprises at least one small cache having a memory size of less than or equal to 10 megabytes (MB). Example 5 is the apparatus of Example 1, further comprising a logic device, wherein the at least one cache memory is operatively coupled to the logic device. Example 6 is the apparatus of Example 1, further comprising an accelerator, wherein the at least one cache memory is operatively coupled to the accelerator. Example 7 is the apparatus of Example 1, further comprising system memory storing a directory including cache status information to indicate whether the at least one cache is a giant cache or a small cache. Example 8 is the apparatus of Example 1, further comprising a coherence controller having a snoop filter, wherein the snoop filter includes cache status information to indicate whether the at least one cache is a giant cache or a small cache. Example 9 is the apparatus of Example 1, wherein the logic is to specify the cache status of the memory operation request via an agent class qualifier of the memory operation request. Example 10 is the apparatus of Example 1, wherein the logic is to use a snoop filter to perform the small cache coherence process. Example 11 is the apparatus of Example 1, wherein the logic is to use a directory array to perform the giant cache coherency process. Example 12 is the apparatus of Example 1, wherein the memory operation request comprises a request to access a cache line, wherein the cache status indicates a small cache status based on a copy of the cache line stored in a small cache agent. Example 13 is the apparatus of Example 1, wherein the memory operation request comprises a request to access a cache line, wherein the cache status indicates the status of a giant cache based on a copy of the cache line stored in a giant cache agent. Example 14 is a system comprising the device according to any one of claims 1 to 13 and at least one transceiver. Example 15 is a method comprising receiving a memory operation request associated with at least one cache, determining a cache status of the memory operation request, the cache status indicating either a huge cache status or a small cache status, executing the memory operation request via a small cache coherence process in response to the cache status being a small cache status, and executing the memory operation request via a huge cache coherence process responsive to the cache status being a small cache status. Example 16 is the method of Example 15, wherein the at least one cache memory comprises at least one small cache and at least one giant cache. Example 17 is the method of Example 15, wherein the at least one cache memory comprises at least one giant cache having a memory size greater than 1 gigabyte (GB). Example 18 is the method of Example 15, wherein the at least one cache memory comprises at least one small cache having a memory size of less than or equal to 10 megabytes (MB). Example 19 is the method of Example 15, wherein the at least one cache memory is operably coupled to a logic device of the computing device. Example 20 is the method of Example 15, wherein the at least one cache memory is operably coupled to an accelerator of the computing device. Example 21 is the method of Example 15, including storing a directory including cache status information to indicate whether the at least one cache is a huge cache or a small cache. Example 22 is the method of Example 15, further comprising a snoop filter stored in a coherence controller of the computing device, the snoop filter comprising cache status information to indicate whether the at least one cache is a giant cache or a small cache. Example 23 is the method of Example 15, including specifying the cache status of the memory operation request via an agent class qualifier of the memory operation request. Example 24 is the method of Example 15, including using a snoop filter to perform the small cache coherency process. Example 25 is the method of Example 15, including using a directory field to perform the giant cache coherency process. Example 26 is the method of Example 15, wherein the memory operation request comprises a request to access a cache line, the cache status indicating the status of the small cache based on a copy of the cache line stored in an agent of the small cache. Example 27 is the method of Example 15, wherein the memory operation request comprises a request to access a cache line, wherein the cache status indicates the status of the huge cache based on a copy of the cache line stored in an agent of the huge cache. Example 28 is a computer-readable storage medium storing instructions for execution by processing circuitry of a computing device, the instructions to cause the computing device to receive a memory operation request associated with at least one cache memory, determine a cache status of the memory operation request, the cache status indicating one of a huge cache status and a small cache status, execute the memory operation request via a small cache coherence process in response to a determination that the cache status is a small cache status, and execute the memory operation request via a huge cache coherence process responsive to determining that the cache status is a small cache status. Example 29 is the computer-readable storage medium of Example 28, wherein the at least one cache memory comprises at least one small cache and at least one giant cache. Example 30 is the computer-readable storage medium of Example 28, wherein the at least one cache memory comprises at least one giant cache having a memory size greater than 1 gigabyte (GB). Example 31 is the computer-readable storage medium of Example 28, wherein the at least one cache memory comprises at least one small cache having a memory size of less than or equal to 10 megabytes (MB). Example 32 is the computer-readable storage medium of Example 28, wherein the at least one cache memory is operably coupled to a logic device of the computing device. Example 33 is the computer-readable storage medium of Example 28, wherein the at least one cache memory is operably coupled to an accelerator of the computing device. Example 34 is the computer-readable storage medium of Example 28, wherein the instructions are to cause the computing device to store a directory comprising cache status information to indicate whether the at least one cache is a giant cache or a small cache. Example 35 is the computer-readable storage medium of Example 28, wherein the instructions are to cause the computing device to provide a snoop filter stored in a coherence controller of the computing device, the snoop filter comprising cache status information to indicate whether the at least one cache is a giant cache or a small cache. Example 36 is the computer-readable storage medium of Example 28, wherein the instructions are to cause the computing device to specify the cache status of the memory operation request via an agent class qualifier of the memory operation request. Example 37 is the computer-readable storage medium of Example 28, wherein the instructions are to cause the computing device to use a snoop filter to perform the small cache coherency process. Example 38 is the computer-readable storage medium of Example 28, wherein the instructions are to cause the computing device to use a directory field to perform the giant cache coherency process. Example 39 is the computer-readable storage medium of Example 28, wherein the memory operation request comprises a request to access a cache line, the cache status indicating the status of the small cache based on a copy of the cache line stored in an agent of the small cache. Example 40 is the computer-readable storage medium of Example 28, wherein the memory operation request comprises a request to access a cache line, the cache status indicating the status of the giant cache based on a copy of the cache line stored in an agent of the giant cache. Example 41 is an apparatus comprising means for at least one cache memory and means for memory operation management to receive a memory operation request associated with the at least one cache memory, determine a cache status of the memory operation request, the cache status indicating either a huge cache status or a small cache status, execute the memory operation request via a small cache coherency process in response to a determination that the cache status is a small cache status, and execute the memory operation request via a huge cache coherency process responsive to the determination that the cache status is a small cache status. Example 42 is the apparatus of Example 41, wherein the at least one cache memory comprises at least one small cache and at least one giant cache. Example 43 is the apparatus of Example 41, wherein the at least one cache memory comprises at least one giant cache having a memory size greater than 1 gigabyte (GB). Example 44 is the apparatus of Example 41, wherein the at least one cache memory comprises at least one small cache having a memory size of less than or equal to 10 megabytes (MB). Example 45 is the apparatus of Example 41, further comprising means for a logic device, wherein the means for the at least one cache memory is operatively coupled to the logic device. Example 46 is the apparatus of Example 41, further comprising means for an accelerator, wherein the means for the at least one cache memory is operatively coupled to the means for the accelerator. Example 47 is the apparatus of Example 41, further comprising means for a system memory storing a directory comprising cache status information to indicate whether the means for the at least one cache memory is a giant cache or a small cache. Example 48 is the apparatus of Example 41, further comprising means for a coherence controller having a snoop filter, wherein the snoop filter includes cache status information to indicate whether the means for the at least one cache memory is a giant cache or a small cache. Example 49 is the apparatus of Example 41, wherein the means for memory operation management is to specify the cache status of the memory operation request via an agent class qualifier of the memory operation request. Example 50 is the apparatus of Example 41, wherein the means for memory operations management is to use a snoop filter to perform the small cache coherency process. Example 51 is the apparatus of Example 41, wherein the means for memory operations management is to use a directory array to perform the giant cache coherency process. Example 52 is the apparatus of Example 41, wherein the memory operation request comprises a request to access a cache line, wherein the cache status indicates the status of the small cache based on a copy of the cache line stored in an agent of the small cache. Example 53 is the apparatus of Example 41, wherein the memory operation request comprises a request to access a cache line, wherein the cache status indicates the status of the giant cache based on a copy of the cache line stored in an agent of the giant cache. Example 54 is a system comprising the apparatus of any one of claims 41 to 53 and at least one transceiver.

Claims

[1] Device (104, 405, 505) comprising: at least one processor (120, 420); at least one cache memory (124a ... 124n, 184a, ..., 184n, 484a, 484b); at least one system memory (156, 510) storing a directory (152, 520) comprising cache status information to indicate whether the at least one cache memory (124a ... 124n, 184a, ..., 184n, 484a, 484b) is a huge cache (484b) or a small cache (484a). and logic (180) at least partially contained in hardware, wherein the logic (180) is designed to: receive a memory operation request associated with the at least one cache memory (124a ... 124n, 184a, ..., 184n, 484a, 484b), determine a cache status of the memory operation request, wherein the cache status indicates either a huge cache status (484b) or a small cache status (484a), execute the memory operation request via a small cache coherence process (484a) in response to a determination that the cache state is a small cache state (484a), and execute the memory operation request via a giant cache coherency process (484b) in response to determining that the cache state is a giant cache state (484b). [2] The apparatus (104, 405, 505) of claim 1, wherein the at least one cache memory (124a ... 124n, 184a, ..., 184n, 484a, 484b) comprises at least one small cache (484a) and at least one giant cache (484b). [3] The apparatus (104, 405, 505) of claim 1, wherein the at least one cache memory (124a ... 124n, 184a, ..., 184n, 484a, 484b) comprises at least one giant cache (484b) having a memory size greater than 1 gigabyte (GB). [4] The device (104, 405, 505) of claim 1, wherein the at least one cache memory (124a ... 124n, 184a, ..., 184n, 484a, 484b) comprises at least one small cache (484a) having a memory size of less than or equal to 10 megabytes (MB). [5] The apparatus (104, 405, 505) of claim 1, further comprising a logic device (180), wherein the at least one cache memory (124a ... 124n, 184a, ..., 184n, 484a, 484b) is operatively coupled to the logic device (180). [6] The apparatus (104, 405, 505) of claim 1, further comprising an accelerator, wherein the at least one cache memory (124a ... 124n, 184a, ..., 184n, 484a, 484b) is operatively coupled to the accelerator. [7] Method (600, 700) comprising: Storing a directory comprising cache status information to indicate whether the at least one cache memory (124a ... 124n, 184a, ..., 184n, 484a, 484b) is a huge cache (484b) or a small cache (484a). Receiving a memory operation request associated with at least one cache memory (124a ... 124n, 184a, ..., 184n, 484a, 484b); Determining a cache status of the memory operation request, the cache status indicating either a huge cache status (484b) or a small cache status (484a); Executing the memory operation request via a small cache coherency process (484a) in response to the cache status being a small cache status (484a); and Executing the memory operation request via a giant cache coherency process (484b) in response to the cache state being a giant cache state (484b). [8] The method (600, 700) of claim 7, wherein the at least one cache memory comprises at least one small cache (484a) and at least one giant cache (484b). [9] The method (600, 700) of claim 7, wherein the at least one cache memory (124a ... 124n, 184a, ..., 184n, 484a, 484b) comprises at least one giant cache (484b) having a memory size greater than 1 gigabyte (GB). [10] The method (600, 700) of claim 7, wherein the at least one cache memory (124a ... 124n, 184a, ..., 184n, 484a, 484b) comprises at least one small cache (484a) having a memory size of less than or equal to 10 megabytes (MB). [11] A computer-readable storage medium (800) storing instructions for execution by processing circuitry of a computing device, the instructions causing the computing device to: provide a snoop filter (426) stored in a coherence controller (132, 422) of the computing device, wherein the snoop filter (426) includes cache status information to indicate whether the at least one cache memory (124a ... 124n, 184a, ..., 184n, 484a, 484b) is a huge cache (484b) or a small cache (484a), receive a memory operation request associated with at least one cache memory (124a ... 124n, 184a, ..., 184n, 484a, 484b); determine a cache status of the memory operation request, the cache status indicating either a huge cache status (484b) or a small cache status (484a); execute the memory operation request via a small cache coherency process (484a) in response to a determination that the cache state is a small cache state (484a); and execute the memory operation request via a giant cache coherency process (484b) in response to a determination that the cache state is a giant cache state (484b). [12] The computer-readable storage medium (800) of claim 11, wherein the at least one cache memory (124a ... 124n, 184a, ..., 184n, 484a, 484b) comprises at least one small cache (484a) and at least one giant cache (484b). [13] The computer-readable storage medium (800) of claim 11, wherein the at least one cache memory (124a ... 124n, 184a, ..., 184n, 484a, 484b) comprises at least one giant cache (484b) having a memory size greater than 1 gigabyte (GB). [14] The computer-readable storage medium (800) of any one of claims 11 to 13, wherein the at least one cache memory (124a ... 124n, 184a, ..., 184n, 484a, 484b) comprises at least one small cache (484a) having a memory size of less than or equal to 10 megabytes (MB).

Citation Information

Patent Citations

  • Adaptive hierarchical cache policy in a microprocessor

    US20140281239A1

  • Multi-regime caching in a virtual file system for cloud-based shared content

    US20160321288A1

  • Near-memory accelerator for offloading pointer chasing operations from a processing element

    US20170139836A1