Regulating operations between a central processing unit cache and a memory module

By determining and managing cache lines associated with the same memory block and writing them to the memory module, the method addresses the issues of high read and write amplification, improving memory bandwidth and reducing latency and wear.

WO2025119473A1PCT designated stage expired Publication Date: 2025-06-12HUAWEI TECH CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2023/084598
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-06
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Current CPU cache management leads to high read and write amplification, reducing usable memory bandwidth, increasing latency, and shortening memory module lifespan, especially with random memory access patterns.

Method used

The method involves determining cache lines associated with the same memory block in the CPU cache and writing these cache lines to the memory module in response to modifications, thereby reducing the number of read/write operations and improving bandwidth utilization.

Benefits of technology

This approach reduces memory medium wear, decreases latency, and enhances CPU cache management efficiency, allowing for larger ECC blocks without performance degradation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2023084598_12062025_PF_FP_ABST
    Figure EP2023084598_12062025_PF_FP_ABST
Patent Text Reader

Abstract

In some examples, a method for regulating operations between a central processing unit (CPU) cache and a memory module of an apparatus, the memory module comprising a plurality of memory blocks, comprises determining, by the CPU cache, a plurality of cache lines associated with the same memory block of the plurality of memory blocks, the plurality of cache lines stored in the CPU cache, and, in response to modifying a cache line of the plurality of cache lines associated with the same memory block, writing, by the CPU cache, the plurality of cache lines associated with the same memory block to the memory module.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] REGULATING OPERATIONS BETWEEN A CENTRAL PROCESSING UNIT CACHE AND A MEMORY MODULE

[0002] TECHNICAL FIELD

[0003] The present disclosure relates, in general, to regulating operations between a central processing unit (CPU) cache and a memory module of an apparatus. Aspects of the disclosure relate to CPU cache write-back and eviction policies.

[0004] BACKGROUND

[0005] Modern computers feature multiple memory modules that are often based on different media (e.g., DRAM, PCM, 3D XPoint, MRAM, ReRAM). The memory modules are either directly attached to the CPU’s memory controller or are provided as separate devices and connected to the computer through a fast interconnect, e.g., based on the Compute Express Link (CXL) standard. Access to the memory happens through memory read and write operations issued by the CPU’s memory controller. Data is moved between the memory and CPU’s caches in blocks called cache lines, which are typically 64 bit (B) in size.

[0006] Some computer memory modules feature fast on-board buffers (or cache), which are used to speed-up access to the medium (which can be considerably slower than, e.g., Dynamic Random Access Memory, DRAM). For example, 3D XPoint-based Intel Optane DC Persistent Memory Module, PMM, features 16 kilobyte (kB) buffer for each memory module (Dual In-line Memory Module, DIMM). The buffers are managed by a separate memory controller located on the memory module itself (denoted MMC) and are transparent to computer CPUs and their cache controllers.

[0007] The fast buffers located on a memory module help to reduce the latency of memory accesses because computer programs typically exhibit temporal and spatial locality in the memory access patterns. As such, the cost of accessing slow memory medium can be amortised over multiple read operations. Similarly, multiple write operations can be coalesced in the buffer, so that costly write operations to the medium do not happen upon every write operation issued by the CPU memory controller to the memory module. Due to technical reasons, accesses to the memory medium sometimes are performed in blocks (called Error Correction Code, ECC, blocks), which are larger than the size of a single cache line. For example, Intel Optane DC PMM internally uses 256B blocks, which are four times larger than a typical cache line. Some of the future memory expansion cards by various vendors, such as Samsung, Panasonic, HP, will also internally use ECC blocks that are larger than 64B (so, e.g., 128B, 256B, 512B, etc.). In such case, the buffers located on a memory module aid the MMC in servicing smaller, cache line-sized read and write operations. To this end, the buffers store the entire ECC blocks, but allow the MMC to read and write portions of the ECC blocks, which correspond to single cache lines.

[0008] The overall costs (expressed in terms of, e.g., latency and bandwidth) of memory accesses heavily depend, among others, on the relative speed of read / write operations to the buffer and to the memory medium, the size of the buffer, the buffer eviction policy (i.e., which ECC block should be evicted to the memory medium to make room in the buffer for an ECC block needed to service the incoming read or write operation) as well as memory access patterns, which are application dependent.

[0009] In particular, random memory access patterns that involve reads and writes scattered across very different memory locations are problematic as they incur high read and write amplification on the memory medium:

[0010] • to service a read of a single cache line, the entire ECC block needs to be read from the memory medium and placed in the buffer,

[0011] • to service a write of a single cache line, the entire ECC block needs to be read from the memory medium, placed in the buffer, and modified in part according to the write operation,

[0012] • each ECC block in the buffer is only rarely reused to service later read / write operations.

[0013] High read and write amplification not only can significantly reduce the useable memory bandwidth and increase the latency of memory accesses. In case of some memory media, such as PCM, which have lower write endurance compared to, e.g., DRAM, high write amplification can cause an increased wear of the memory medium, and, in consequence, shorten the memory module lifespan. SUMMARY

[0014] An objective of the present disclosure is to improve the management of cache lines in the CPU cache.

[0015] The foregoing and other objectives are achieved by the features of the independent claims.

[0016] Further implementation forms are apparent from the dependent claims, the description and the Figures.

[0017] A first aspect of the present disclosure provides a method for regulating operations between a central processing unit, CPU, cache and a memory module of an apparatus, the memory module comprising a plurality of memory blocks, the method comprising determining, by the CPU cache, a plurality of cache lines associated with the same memory block of the plurality of memory blocks, the plurality of cache lines stored in the CPU cache, and, in response to modifying a cache line of the plurality of cache lines associated with the same memory block, writing, by the CPU cache, the plurality of cache lines associated with the same memory block to the memory module.

[0018] Accordingly, the number of read / write operations from / to the memory medium can be reduced. Furthermore, the utilisation of memory read / write bandwidth can be improved, increasing performance. In addition, by reducing the number of operations, memory medium wear can be reduced. Furthermore, the latency of the eviction of some cache lines from the CPU cache can be reduced, thereby increasing the performance of a related application.

[0019] Determining, by the CPU cache, the plurality of cache lines associated with the same memory block of the plurality of memory block may comprise acquiring, by a CPU comprising the CPU cache, information relating to the layout of the multiple memory blocks in the memory module, wherein the information relating to the layout of the multiple memory blocks comprises at least one of information about size of the multiple memory blocks and physical addresses of the multiple memory blocks, processing the acquired information to thereby determine the layout of the multiple memory blocks, the layout comprising a physical and / or a logical layout, and determining, by the CPU cache, the plurality of cache lines associated with the same memory block of the plurality of memory blocks based on the determined layout of the multiple memory blocks. The method may further comprise, in response to determining the plurality of cache lines associated with the same memory block of the plurality of memory blocks, storing, by the CPU cache, metadata indicating which memory block of the plurality of memory blocks each of the plurality of cache lines is associated with.

[0020] The method may further comprise, in response to evicting, by the CPU cache, the cache line of the plurality of cache lines associated with the same memory block from the CPU cache, or writing, by the CPU cache, the cache line of the plurality of cache lines associated with the same memory block to the memory module, writing and / or evicting the remaining cache lines of the plurality of cache lines associated with the same memory block.

[0021] The writing and the evicting may be performed substantially simultaneously.

[0022] A second aspect of the present disclosure provides a computer readable storage medium comprising computer program code, accessible by an apparatus comprising a processor, to provide instructions and / or data to the apparatus, the computer program code configured to, with the processor, cause the apparatus to determine, by a central processing unit, CPU, cache, a plurality of cache lines associated with the same memory block of a plurality of memory blocks comprised by a memory module, wherein the plurality of cache lines is stored in the CPU cache, and , in response to modifying a cache line of the plurality of cache lines associated with the same memory block, write, by the CPU cache, the plurality of cache lines associated with the same memory block to the memory module.

[0023] In order to determine, by the CPU cache, the plurality of cache lines associated with the same memory block of the plurality of memory blocks, the computer program code may be further configured to, with the processor, cause the apparatus to acquire, by a CPU comprising the CPU cache, information relating to the layout of the multiple memory blocks in the memory module, wherein the information relating to the layout of the multiple memory blocks comprises at least one of information about size of the multiple memory blocks and physical addresses of the multiple memory blocks, process the acquired information to thereby determine the layout of the multiple memory blocks, the layout comprising a physical and / or a logical layout, and determine, by the CPU cache, the plurality of cache lines associated with the same memory block of the plurality of memory blocks based on the determined layout of the multiple memory blocks. The computer program code may be further configured to, with the processor, cause the apparatus to, in response to determining the plurality of cache lines associated with the same memory block of the plurality of memory blocks, store, by the CPU cache, metadata indicating which memory block of the plurality of memory blocks each of the plurality of cache lines is associated with.

[0024] The computer program code may be further configured to, with the processor, cause the apparatus to, in response to evicting, by the CPU cache, the cache line of the plurality of cache lines associated with the same memory block from the CPU cache, or writing, by the CPU cache, the cache line of the plurality of cache lines associated with the same memory block to the memory module, write and / or evict the remaining cache lines of the plurality of cache lines associated with the same memory block.

[0025] The computer program code may be further configured to, with the processor, cause the apparatus to perform the writing and the evicting substantially simultaneously.

[0026] A third aspect of the present disclosure provides an apparatus comprising a memory module comprising a plurality of memory blocks, and a central processing unit, CPU, comprising a CPU cache, wherein the CPU cache is arranged to determine a plurality of cache lines associated with the same memory block of the plurality of memory blocks, the plurality of cache lines stored in the CPU cache, and, in response to modifying a cache line of the plurality of cache lines associated with the same memory block, write the plurality of cache lines associated with the same memory block to the memory module.

[0027] In order to determine the plurality of cache lines associated with the same memory block of the plurality of memory blocks, the CPU may be further arranged to acquire information relating to the layout of the multiple memory blocks in the memory module, wherein the information relating to the layout of the multiple memory blocks comprises at least one of information about size of the multiple memory blocks and physical addresses of the multiple memory blocks, and process the acquired information to thereby determine the layout of the multiple memory blocks, the layout comprising a physical and / or a logical layout, wherein the CPU cache may be arranged to determine the plurality of cache lines associated with the same memory block of the plurality of memory blocks based on the determined layout of the multiple memory blocks.

[0028] The CPU cache may be further arranged to, in response to determining the plurality of cache lines associated with the same memory block of the plurality of memory blocks, store metadata indicating which memory block of the plurality of memory blocks each of the plurality of cache lines is associated with.

[0029] The CPU cache may be further arranged to, in response to evicting the cache line of the plurality of cache lines associated with the same memory block from the CPU cache, or writing the cache line of the plurality of cache lines associated with the same memory block to the memory module, write and / or evict the remaining cache lines of the plurality of cache lines associated with the same memory block.

[0030] The CPU cache may be arranged to perform the writing and the evicting substantially simultaneously.

[0031] These and other aspects of the invention will be apparent from the embodiment(s) described below.

[0032] BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order that the present invention may be more readily understood, embodiments of the invention will now be described, by way of example, with reference to the accompanying drawings, in which:

[0034] Figs, la-lm are a step-by-step depiction of a read-modify-write operation according to the prior art;

[0035] Fig. 2 is a flow chart of a method for regulating operations between a CPU cache and a memory module of an apparatus according to an example;

[0036] Figs. 3a-3i are a step-by-step depiction of a read-modify-write operation according to an embodiment;

[0037] Fig. 4 is a flow chat of an extended CPU cache management policy according to an example; and

[0038] Fig. 5 is a schematic depiction of an apparatus according to an example. DETAILED DESCRIPTION

[0039] Example embodiments are described below in sufficient detail to enable those of ordinary skill in the art to embody and implement the systems and processes herein described. It is important to understand that embodiments can be provided in many alternate forms and should not be construed as limited to the examples set forth herein.

[0040] Accordingly, while embodiments can be modified in various ways and take on various alternative forms, specific embodiments thereof are shown in the drawings and described in detail below as examples. There is no intent to limit to the particular forms disclosed. On the contrary, all modifications, equivalents, and alternatives falling within the scope of the appended claims should be included. Elements of the example embodiments are consistently denoted by the same reference numerals throughout the drawings and detailed description where appropriate.

[0041] The terminology used herein to describe embodiments is not intended to limit the scope. The articles “a,” “an,” and “the” are singular in that they have a single referent, however the use of the singular form in the present document should not preclude the presence of more than one referent. In other words, elements referred to in the singular can number one or more, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises,” “comprising,” “includes,” and / or “including,” when used herein, specify the presence of stated features, items, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, items, steps, operations, elements, components, and / or groups thereof.

[0042] Unless otherwise defined, all terms (including technical and scientific terms) used herein are to be interpreted as is customary in the art. It will be further understood that terms in common usage should also be interpreted as is customary in the relevant art and not in an idealized or overly formal sense unless expressly so defined herein.

[0043] Typically, CPU caches help a lot in reducing the required memory bandwidth, reducing the latency of memory accesses, and similar. Unfortunately, in some cases, the way in which data is managed in CPU caches may contribute to the reduction in the useable memory bandwidth and increases the latency of memory accesses, as well as shortens the memory module lifespan. In particular, issues may arise due to the fact that the CPU caches are not synchronised with the buffer of the memory module, causing the above mentioned issues. More precisely, each cache line in CPU caches is considered independent and thus is fetched and evicted by the CPU cache controller as needed by issuing read and write operations that typically concern the entire cache lines. There is no additional coordination between the CPU cache controller and the MMC in regard to the prefetching ECC blocks to the buffer and evicting them to the memory medium later on. In contrary, the multi-level CPU caches of modem computers are always tightly coordinated. In principle, random memory accesses that concern objects that occupy at most one cache line will incur read and write amplification that directly corresponds to the ratio of the size of an ECC block on the memory module and the size of a single cache line. However, even for objects that occupy multiple (i.e., consecutive) cache lines or for objects allocated in memory next to each other and often accessed at the same time or in quick succession (as applications typically exhibit temporal and spatial locality), the read and write amplification is still high. It is because the CPU cache controller considers all cache lines independent, even though they might belong to the same ECC block and be modified together during computation. Moreover, CPU caches are typically much larger compared to the size of the buffers on-board the memory modules (tens of MB vs tens of KB). Hence, an ECC block is typically evicted from the memory module buffer well before the cache lines that are part of the ECC block are evicted from the CPU caches. This increases the read and write amplification even further.

[0044] Figs, la-lm are a step-by-step depiction of a read-modify-write operation according to the prior art. In Fig. la, a memory block #42 (e.g., ECC block) is already present in the buffer 130 of the memory module 120 (i.e., the memory block #42 has already been fetched into the buffer 130 from memory 140 of the memory module 120), and the cache lines #42a and #42c are already present in a CPU cache 111 of the CPU 110. The CPU 110 performs a read-modify -write operation with a random-access pattern and modifies two cache lines present in the CPU cache 111 - #42c and #42a - at the same time or in quick succession. In particular, in Fig. la, the CPU 110 reads (depicted as arrows pointing from the cache lines) the cache lines #42c and #42a from its cache 111. In Fig. lb, the CPU modifies the previously read cache lines and writes (depicted as arrows pointing to the cache lines) the modified cache lines #42c and #42a to its cache 111. In the figures, the modified cache lines are depicted as having a shaded background.

[0045] In Fig. 1c, the memory block #42 (containing the cache lines #42a and #42c) is evicted from the buffer 130 of the memory module 120. The memory block #42 was initially placed in the buffer 130 to serve the read operations performed by the CPU 110. The memory block #42 is evicted from the buffer 130 early, because the buffer 130 is much smaller than the CPU cache 111. In Fig. Id, the CPU evicts the modified cache line #42c from its cache 111, and writes the modified cache line #42c to the memory module 120. However, eviction of each of the modified cache lines #42a and #42c from the CPU cache 111 requires fetching the memory block #42 into the buffer 130 again. As such, in Fig. le, the memory block #42 is fetched into the buffer 130 again, in order to write the modified cache line #42c thereto. In Fig. If, the modified cache line #42c is updated in the buffer 130. Fig. 1g shows the modified cache line #42c now present in the memory block #42 in the buffer 130. In Fig. Ih, the memory block #42 is once again evicted from the buffer 130 and replaced by some other memory block. In Fig. li, the CPU 110 evicts the remaining modified cache line #42a from its cache 111 and writes the modified cache line #42a to the memory module 120. In Fig. Ij, the memory block #42 is fetched into the buffer 130 again, similar to step Fig. le. In Fig. Ik, the modified cache line #42a is written to the buffer 130. Fig. 11 shows the modified cache line #42a now present in the memory block #42 in the buffer 130. Finally, in Fig. Im, the memory block #42c is once again evicted from the buffer 130.

[0046] In other words, even though the modifications of two cache lines (64B each) that belong to the same memory block (256B) happened at the same time or in quick succession, two additional reads from the memory medium were required, thus requiring 512B of additional read bandwidth (768B in total, including the initial read, omitted from Figs, la-lm for brevity). Additionally, two writes to the memory medium were required, requiring 512B of bandwidth.

[0047] According to an example, there is provided a mechanism to reduce the number of read / write operations from / to the memory medium that happen, in particular, with random memory access patterns. Aspects of the invention relate to allowing some caches to be evicted more quickly and efficiently from the CPU cache. In turn, the invention increases the usable read / write bandwidth of the memory module, reduces the memory medium wear, and makes CPU cache management more efficient. Moreover, the improvement in utilisation of memory read / write bandwidth enables the use of larger memory (ECC) blocks to store data in the memory medium without degrading the performance. This, in turn, improves the effective capacity of the memory module.

[0048] The invention concerns computer systems that feature memory modules with on-board buffers. The memory modules can be based on different media (e.g., DRAM, PCM, 3D XPoint, MRAM, ReRAM) and can be either directly attached to the CPU’s memory controller, or be provided as separate devices and connected to the computer through a fast interconnect, e.g., based on the Compute Express Link (EXL) standard. In particular, the invention affects the computer CPU cache subsystem, i.e., some or all of the CPU caches (e.g., LI, L2, L3) and the CPU cache controller, as well as the memory module controller aboard the memory modules.

[0049] Examples in the present disclosure can be provided as methods, systems or machine-readable instructions, such as any combination of software, hardware, firmware or the like. Such machine-readable instructions may be included on a computer readable storage medium (including but not limited to disc storage, CD-ROM, optical storage, etc.) having computer readable program codes therein or thereon.

[0050] The present disclosure is described with reference to flow charts and / or block diagrams of the method, devices and systems according to examples of the present disclosure. Although the flow diagrams described above show a specific order of execution, the order of execution may differ from that which is depicted. Blocks described in relation to one flow chart may be combined with those of another flow chart. In some examples, some blocks of the flow diagrams may not be necessary and / or additional blocks may be added. It shall be understood that each flow and / or block in the flow charts and / or block diagrams, as well as combinations of the flows and / or diagrams in the flow charts and / or block diagrams can be realized by machine readable instructions.

[0051] The machine-readable instructions may, for example, be executed by a machine such as a general-purpose computer, user equipment such as a smart device, e.g., a smart phone, a special purpose computer, an embedded processor or processors of other programmable data processing devices to realize the functions described in the description and diagrams. In particular, a processor or processing apparatus may execute the machine-readable instructions. Thus, modules of apparatus (for example, a module implementing a comparator unit, or a firewall structure and so on) may be implemented by a processor executing machine readable instructions stored in a memory, or a processor operating in accordance with instructions embedded in logic circuitry. The term 'processor' is to be interpreted broadly to include a CPU, processing unit, ASIC, logic unit, or programmable gate set etc. The methods and modules may all be performed by a single processor or divided amongst several processors.

[0052] Such machine-readable instructions may also be stored in a computer readable storage that can guide the computer or other programmable data processing devices to operate in a specific mode. For example, the instructions may be provided on a non-transitory computer readable storage medium encoded with instructions, executable by a processor.

[0053] Fig. 2 is a flow chart of a method for regulating operations between a CPU cache and a memory module of an apparatus according to an example. In block 201, the method comprises determining, by the CPU cache, a plurality of cache lines associated with the same memory block of a plurality of memory blocks, the plurality of cache lines stored in the CPU cache, wherein the memory module comprises the plurality of memory blocks. In block 202, in response to modifying a cache line of the plurality of cache lines associated with the same memory block, the method comprises writing, by the CPU cache, the plurality of cache lines associated with the same memory block to the memory module.

[0054] To better illustrate this, reference is made to Figs. 3a-3i, which are a step-by-step depiction of a read-modify-write operation according to an embodiment. Figs. 3a, 3b and 3c correspond to Figs, la, lb and 1c, respectively, with no changes to the CPU / memory module behaviour. Same elements in Figs, la-lm and Figs. 3a-3i are denoted using the same reference numerals and function likewise.

[0055] In Fig. 3a, a memory block #42 (e.g., ECC block) is already present in the buffer 130 of the memory module 120, and the cache lines #42a and #42c are already present in a CPU cache 111 of the CPU 110. The CPU 110 performs a read-modify-write operation with a randomaccess pattern and modifies two cache lines present in the CPU cache 111 - #42c and #42a - at the same time or in quick succession. In particular, in Fig. 3a, the CPU 110 reads (depicted as arrows pointing from the cache lines) the cache lines #42c and #42a from its cache 111. In Fig. 3b, the CPU modifies the previously read cache lines and writes (depicted as arrows pointing to the cache lines) the modified cache lines #42c and #42a to its cache 111. In Fig. 1c, the memory block #42 (containing the cache lines #42a and #42c) is evicted from the buffer 130 of the memory module 120. The memory block #42 was initially placed in the buffer 130 to serve the read operations performed by the CPU 110. The memory block #42 is evicted from the buffer 130 early, because the buffer 130 is much smaller than the CPU cache 111.

[0056] In Fig. 3d, the modified cache line #42c is to be evicted from the CPU cache 111. The eviction of the modified cache line #42c triggers the write-back of the modified cache line #42a, associated with the same memory block #42 as the cache line #42c to be evicted. As such, in Fig. 3e, the modified cache line #42a and the modified cache line #42c are to be written to the memory module 120. In order to enable the cache lines to be written to the memory block they are associated with, in Fig. 3f, the memory module fetches the memory block #42 into the buffer 130, and the CPU 110 marks #42a as written-back (depicted as a darker shade of grey in the Figure). At this point, the modified cache line #42c is no longer present in the CPU cache.

[0057] In Fig. 3g, the modified cache lines #42a and #42c are updated in the buffer 130 of the memory module. Fig. 3h shows the buffer 130, in which the modified cache lines #42a and #42c are present. Finally, in Fig. 3i, the memory block #42 (associated with the modified cache lines #42a and #42c) is evicted from the buffer 130 and replaced by some other ECC block. At this point, whenever the cache controller decides to evict the modified cache line #42a from the CPU cache 111, it can do so easily by simply replacing it with some other cache line, as the modified cache line #42a has already been written to the memory module 120.

[0058] In order to be able to choose the quick path during cache line eviction, the CPU should know which cache lines have already been written back to the memory module. This capability may be realised by a CPU cache controller and a cache coherence protocol (e.g., MESI) that maintains a consistent state of all cache lines across all CPU caches. The details of the implementation of the existing cache coherence protocol are not crucial to the invention.

[0059] While the above example relates to a read-modify-write operation, the invention could also be used, for example, to enable the CPU cache controller to prefetch into the CPU cache multiple cache lines, for example, in anticipation for an operation to be performed by the CPU on the cache lines. With the invention, if deemed useful by the cache controller, some or all cache lines associated with the same memory block could be fetched into the CPU cache, while making sure that the memory block is fetched from the memory of the memory module into the buffer of the memory module only once. The above mechanism can be used in conjunction with other mechanisms already supported by the cache controller, such as, allocating the cache line with zeroes, thereby preparing a cache line before overwriting its contents completely by the CPU.

[0060] As described above, in block 201 of Fig. 2, the method comprises determining, by the CPU cache, the plurality of cache lines (stored in the CPU cache) associated with the same memory block. To achieve this, the CPU may acquire, from the memory module, information relating to the layout of the multiple memory blocks stored in the memory module. The information relating to the layout of the multiple memory blocks may comprise, for example, information about the size of the multiple memory blocks, type of the memory blocks, and / or physical addresses of the multiple memory blocks.

[0061] The memory module may provide the information relating to the layout of the multiple memory blocks (i.e., additional metadata) to the CPU cache. The additional metadata may be provided as a part of each memory (for example, read and / or write) operation response. Alternatively, the additional metadata may be provided to the CPU cache controller by the memory module during the memory initialisation process, to be stored in the CPU cache and used for all cache lines whose addresses fall within the address range associated by the computer with the memory module. In such case, advantageously, the metadata does not need to be shipped as all memory operation read / write responses, thereby not occupying the bandwidth of the memory bus. Furthermore, using this approach, the metadata information does not need to be stored in each CPU cache entry, which decreases the size requirement for the CPU cache. The management of this metadata might be problematic in particular in relation to multi-level CPU caches that are not fully inclusive.

[0062] To aid understanding of the invention, detail will now be provided about the additional metadata, i.e., the information relating to the layout of the memory blocks stored in the memory module. Assuming that the CPU operates on 64-bit cache lines, the additional metadata may utilise two bits in the following way:

[0063] 00 - the cache line can be accessed directly as a 64B memory block (i.e., the ECC block encompassing the cache line is 64B aligned)

[0064] 01 - the cache line is part of a 128B-sized memory block (i.e., the ECC block encompassing the cache line is 128B aligned)

[0065] 10 - the cache line is part of a 256B-sized memory block (i.e., the ECC block encompassing the cache line is 256B aligned)

[0066] 11 - the cache line is part of a 512B-sized memory block (i.e., the ECC block encompassing the cache line is 512B aligned).

[0067] In other words, cache lines in the CPU cache, whose metadata values are other than 00, belong to ECC blocks several times larger than the cache lines themselves. For such cache lines it is necessary to calculate - which other entries in the CPU cache have to be checked for cache lines that belong to the same memory block. In order to aid understanding of the invention, an example CPU cache line will be described. For example, a 64-bit machine (48-bit virtual addresses), with 2MB, 4-way associative CPU cache, which supports memory modules that use standard 64B ECC blocks (e.g., DRAM) as well as memory modules that use 128B, 256B and 512B ECC blocks, may be considered. Since the CPU cache is 2MB in size and each cache line is 64B in size, the CPU cache features 32768 entries. Since the CPU cache is 4-way associative (each cache line can be stored only in 4 different entries whose location in the CPU cache is known in advance), it is divided into 8192 sets, each of which contains exactly 4 entries. Finally, each entry may feature:

[0068] • Valid bit - which signifies whether the entry is occupied or not,

[0069] • Dirty bit - which signifies whether the cache line stored in the entry has been written back to the memory or not,

[0070] • ECC block alignment bits - which signify whether a cache line stored in the entry is part of an ECC block that is 64B, 128B, 256B or 512B in size (as we described above),

[0071] • Recency metadata - some additional bits that are updated upon accesses to the cache line stored in the entry and then used by the CPU cache controller to decide, which entry to empty in order to make space for other cache lines required by the CPU,

[0072] • Tag - a portion of the address of a cache line that allows one to identify the cache line stored in the entry,

[0073] • Data - the 64B cache line stored in the entry.

[0074] As such, according to the above scenario, each 64-bit memory address processed by the CPU cache (in particular, by the CPU cache controller) may consist of:

[0075] The first 16 bits may be unused because the above-described machine utilises 48-bit virtual addresses. The following 42 bits used for the cache line address (i.e., the tag and the set) may be used by the CPU cache controller to find, in the CPU cache, the appropriate set of four entries where the cache line can be stored (or is already stored), as well as to identify the given cache line within the set using the tag. Finally, the last 6 bits may be used to identify individual bits within the cache line.

[0076] In order to find, in the CPU cache, all the other cache lines from the same memory block as a cache line with tag t, memory block alignment xx (expressed as bits), that is present in set s (expressed as an integer number), the method may comprise:

[0077] Calculating the base set number - base = (s » dec(xx)) « dec(xx), where « and » represent standard bit shift operations and dec is a function that converts a binary number into a 10-base integer, and

[0078] Checking 2dec(xx)- i.e., 2 to the power of dec(xx) - consecutive sets, starting from base, for cache lines with tag t and the valid bit set (i.e. signifying that the entry is occupied).

[0079] All of the cache lines found using the above-described mechanism may belong to the same memory block having a size of 26+dec(xx)bytes.

[0080] To aid understanding of the above-described method, an example is provided, in which a cache line 12000 with tag 0xlF8D01D5 in the set 5818, having a cache line alignment of 10 (i.e., 2 in decimal notation) is provided. As described above, according to the example, the method may comprise:

[0081] Calculating the base set number - (5818 » 2) « 2 = 5816,

[0082] Checking 22= 4 consecutive set, starting from 5816, for cache lines with tag 0xlF8D01D5 and the valid bit set.

[0083] In the example above, all of the cache lines found using the method belong to the same 256B- sized memory block.

[0084] Referring back to block 202 of Fig. 2, in response to modifying the cache line of the plurality of cache lines associated with the same memory block, the CPU cache may write the plurality of cache lines associated with the same memory block to the memory module. As was described in relation to Fig. 3, the invention may allow the write-back and / or eviction of multiple cache lines together, if the cache lines belong to the same memory block. In order to achieve that, the CPU cache management policy may be extended. Fig. 4 is a flow chat of an extended CPU cache management policy according to an example. In block 401, a cache line is to be evicted from the CPU cache. In block 402, the method may check whether the dirty bit for the cache line is set. If the dirty bit is not set, the line may be evicted from the CPU cache. If the dirty bit is set, a write-back may be performed on the cache line, as shown in block 403. As an extension to the CPU cache management policy, in addition, if block 402 reveals that the dirty bit is set, the method may, in block 404, find (in other sets) all cache lines from the same memory block which also have the dirty bit set. Then, in block 405, line write-back may be performed.

[0085] However, the extension to the CPU cache management policy may involve any combination of a trigger action (e.g., the eviction of a cache line, the write-back of a cache line) and actions on the other cache lines that are present in the CPU cache and belong to the same ECC block as the cache line in question. To this extent, the method may comprise, for example, evicting all the other cache lines belonging to the same memory block, write-back of all the other cache lines belonging to the same memory block, and / or performing the write-back / eviction on only some of the other cache lines belonging to the same memory block.

[0086] Additionally, the changes to the write-back and eviction policy may involve an algorithm used for selecting a cache line for eviction (e.g., from a set of four possible candidates in a 4-way associative cache). In particular, an eviction of a cache line that triggers the most write-backs of cache lines that belong to the same memory block as the evicted cache line may be preferred over other options.

[0087] Fig. 5 is a schematic depiction of an apparatus according to an example. The apparatus 500 comprises a memory module 520 comprising a plurality of memory blocks and a central processing unit (CPU) 510 comprising a CPU cache 511. The memory module 520 may further comprise a memory 540 where the plurality of memory blocks is stored. The CPU cache 511 is arranged to determine a plurality of cache lines associated with the same memory block of the plurality of memory blocks, the plurality of cache lines stored in the CPU cache, and, in response to modifying a cache line of the plurality of cache lines associated with the same memory block, write the plurality of cache lines associated with the same memory block to the memory module.

[0088] According to an example, machine-readable instructions can be loaded onto a computer or other programmable data processing devices, so that the computer or other programmable data processing devices perform a series of operations to produce computer-implemented processing, thus the instructions executed on the computer or other programmable devices provide an operation for realizing functions specified by flow(s) in the flow charts and / or block(s) in the block diagrams.

[0089] Further, the teachings herein may be implemented in the form of a computer or software product, such as a non-transitory machine-readable storage medium, the computer software or product being stored in a storage medium and comprising a plurality of instructions, e.g., machine readable instructions, for making a computer device implement the methods recited in the examples of the present disclosure.

[0090] In some examples, some methods can be performed in a cloud-computing or network-based environment. Cloud-computing environments may provide various services and applications via the Internet. These cloud-based services (e.g., software as a service, platform as a service, infrastructure as a service, etc.) may be accessible through a web browser or other remote interface of the user equipment for example. Various functions described herein may be provided through a remote desktop environment or any other cloud-based computing environment.

[0091] While various embodiments have been described and / or illustrated herein in the context of fully functional computing systems, one or more of these exemplary embodiments may be distributed as a program product in a variety of forms, regardless of the particular type of computer- readable-storage media used to actually carry out the distribution. The embodiments disclosed herein may also be implemented using software modules that perform certain tasks. These software modules may include script, batch, or other executable files that may be stored on a computer-readable storage medium or in a computing system. In some embodiments, these software modules may configure a computing system to perform one or more of the exemplary embodiments disclosed herein. In addition, one or more of the modules described herein may transform data, physical devices, and / or representations of physical devices from one form to another.

[0092] The preceding description has been provided to enable others skilled in the art to best utilize various aspects of the exemplary embodiments disclosed herein. This exemplary description is not intended to be exhaustive or to be limited to any precise form disclosed. Many modifications and variations are possible without departing from the spirit and scope of the instant disclosure. The embodiments disclosed herein should be considered in all respects illustrative and not restrictive. Reference should be made to the appended claims and their equivalents in determining the scope of the instant disclosure.

Claims

CLAIMS1. A method for regulating operations between a central processing unit, CPU, cache and a memory module of an apparatus, the memory module comprising a plurality of memory blocks, the method comprising: determining, by the CPU cache, a plurality of cache lines associated with the same memory block of the plurality of memory blocks, the plurality of cache lines stored in the CPU cache (201); and in response to modifying a cache line of the plurality of cache lines associated with the same memory block, writing, by the CPU cache, the plurality of cache lines associated with the same memory block to the memory module (202).

2. The method of claim 1, wherein determining, by the CPU cache, the plurality of cache lines associated with the same memory block of the plurality of memory block comprises: acquiring, by a CPU comprising the CPU cache, information relating to the layout of the multiple memory blocks in the memory module, wherein the information relating to the layout of the multiple memory blocks comprises at least one of information about size of the multiple memory blocks and physical addresses of the multiple memory blocks; processing the acquired information to thereby determine the layout of the multiple memory blocks, the layout comprising a physical and / or a logical layout; and determining, by the CPU cache, the plurality of cache lines associated with the same memory block of the plurality of memory blocks based on the determined layout of the multiple memory blocks.

3. The method of claim 2, further comprising: in response to determining the plurality of cache lines associated with the same memory block of the plurality of memory blocks, storing, by the CPU cache, metadata indicating which memory block of the plurality of memory blocks each of the plurality of cache lines is associated with.

4. The method of any preceding claim, further comprising:in response to evicting, by the CPU cache, the cache line of the plurality of cache lines associated with the same memory block from the CPU cache, or writing, by the CPU cache, the cache line of the plurality of cache lines associated with the same memory block to the memory module, writing and / or evicting the remaining cache lines of the plurality of cache lines associated with the same memory block.

5. The method of claim 4, wherein the writing and the evicting are performed substantially simultaneously.

6. A computer readable storage medium comprising computer program code, accessible by an apparatus comprising a processor, to provide instructions and / or data to the apparatus, the computer program code configured to, with the processor, cause the apparatus to: determine, by a central processing unit, CPU, cache, a plurality of cache lines associated with the same memory block of a plurality of memory blocks comprised by a memory module, wherein the plurality of cache lines is stored in the CPU cache; and in response to modifying a cache line of the plurality of cache lines associated with the same memory block, write, by the CPU cache, the plurality of cache lines associated with the same memory block to the memory module.

7. The computer readable storage medium of claim 6, wherein, in order to determine, by the CPU cache, the plurality of cache lines associated with the same memory block of the plurality of memory blocks, the computer program code is further configured to, with the processor, cause the apparatus to: acquire, by a CPU comprising the CPU cache, information relating to the layout of the multiple memory blocks in the memory module, wherein the information relating to the layout of the multiple memory blocks comprises at least one of information about size of the multiple memory blocks and physical addresses of the multiple memory blocks; process the acquired information to thereby determine the layout of the multiple memory blocks, the layout comprising a physical and / or a logical layout; and determine, by the CPU cache, the plurality of cache lines associated with the same memory block of the plurality of memory blocks based on the determined layout of the multiple memory blocks.

8. The computer readable storage medium of claim 7, wherein the computer program code is further configured to, with the processor, cause the apparatus to: in response to determining the plurality of cache lines associated with the same memory block of the plurality of memory blocks, store, by the CPU cache, metadata indicating which memory block of the plurality of memory blocks each of the plurality of cache lines is associated with.

9. The computer readable storage medium of any one of claims 6 to 8, wherein the computer program code is further configured to, with the processor, cause the apparatus to: in response to evicting, by the CPU cache, the cache line of the plurality of cache lines associated with the same memory block from the CPU cache, or writing, by the CPU cache, the cache line of the plurality of cache lines associated with the same memory block to the memory module, write and / or evict the remaining cache lines of the plurality of cache lines associated with the same memory block.

10. The computer readable storage medium of claim 9, wherein the computer program code is further configured to, with the processor, cause the apparatus to: perform the writing and the evicting substantially simultaneously.

11. An apparatus, comprising: a memory module (520) comprising a plurality of memory blocks; and a central processing unit (510), CPU, comprising a CPU cache (511), wherein the CPU cache (511) is arranged to: determine a plurality of cache lines associated with the same memory block of the plurality of memory blocks, the plurality of cache lines stored in the CPU cache (511); and in response to modifying a cache line of the plurality of cache lines associated with the same memory block, write the plurality of cache lines associated with the same memory block to the memory module (520).

12. The apparatus of claim 11, wherein, in order to determine the plurality of cache lines associated with the same memory block of the plurality of memory blocks, the CPU (510) is further arranged to:acquire information relating to the layout of the multiple memory blocks in the memory module (520), wherein the information relating to the layout of the multiple memory blocks comprises at least one of information about size of the multiple memory blocks and physical addresses of the multiple memory blocks; and process the acquired information to thereby determine the layout of the multiple memory blocks, the layout comprising a physical and / or a logical layout, wherein the CPU cache (511) is arranged to: determine the plurality of cache lines associated with the same memory block of the plurality of memory blocks based on the determined layout of the multiple memory blocks.

13. The apparatus of claim 12, wherein the CPU cache (511) is further arranged to: in response to determining the plurality of cache lines associated with the same memory block of the plurality of memory blocks, store metadata indicating which memory block of the plurality of memory blocks each of the plurality of cache lines is associated with.

14. The apparatus of any one of claims 11 to 13, wherein the CPU cache (511) is further arranged to: in response to evicting the cache line of the plurality of cache lines associated with the same memory block from the CPU cache (511), or writing the cache line of the plurality of cache lines associated with the same memory block to the memory module (520), write and / or evict the remaining cache lines of the plurality of cache lines associated with the same memory block.

15. The apparatus of claim 14, wherein the CPU cache (511) is arranged to perform the writing and the evicting substantially simultaneously.

Citation Information

Patent Citations

  • Method and apparatus for memory management

    US20150067264A1

  • Rinsing cache lines from a common memory page to memory

    US20190179770A1

  • Translation cache and configurable ECC memory for reducing ECC memory overhead

    US20220012126A1