Improving the management of a buffer of a memory module through coordination with the CPU cache

By coordinating CPU cache and memory module buffer operations, the method addresses the issue of read and write amplification, improving memory bandwidth, reducing wear, and lowering latency for enhanced system performance.

WO2025119465A1PCT designated stage expired Publication Date: 2025-06-12HUAWEI TECH CO LTD +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2023/084554
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-06
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Current memory management systems face challenges in reducing read and write amplification, leading to increased latency and memory medium wear, especially with random memory access patterns.

Method used

The proposed solution involves coordinating operations between the CPU cache and the buffer of a memory module by querying the CPU cache for modified cache lines before eviction or write-back, allowing for efficient updating and reduction in unnecessary read/write operations.

Benefits of technology

This approach improves memory read/write bandwidth utilization, reduces memory medium wear, and decreases latency by optimizing cache line eviction and management, thereby enhancing overall system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2023084554_12062025_PF_FP_ABST
    Figure EP2023084554_12062025_PF_FP_ABST
Patent Text Reader

Abstract

In some examples, a method for regulating operations between a central processing unit, CPU, cache and a buffer of a memory module of an apparatus comprises placing, by the memory module, a memory block comprising multiple cache lines for the apparatus into the buffer of the memory module, querying, by the memory module, the CPU cache to obtain a plurality of modified cache lines associated with the memory block, and updating the multiple cache lines in the memory block based on the obtained plurality of modified cache lines, wherein the querying occurs before eviction and / or a write-back of the memory block from the buffer of the memory module of the apparatus.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] IMPROVING THE MANAGEMENT OF A BUFFER OF A MEMORY MODULE THROUGH COORDINATION WITH THE CPU CACHE

[0002] TECHNICAL FIELD

[0003] The present disclosure relates, in general, to regulating operations between a central processing unit (CPU) cache and a buffer of a memory module of an apparatus. Aspects of the disclosure relate to improving the management of a memory module’s buffer through coordination with the CPU cache.

[0004] BACKGROUND

[0005] Modern computers feature multiple memory modules that are often based on different media (e.g., DRAM, PCM, 3D XPoint, MRAM, ReRAM). The memory modules are either directly attached to the CPU’s memory controller or are provided as separate devices and connected to the computer through a fast interconnect, e.g., based on the Compute Express Link (CXL) standard. Access to the memory happens through memory read and write operations issued by the CPU’s memory controller. Data is moved between the memory and CPU’s caches in blocks called cache lines, which are typically 64 bit (B) in size.

[0006] Some computer memory modules feature fast on-board buffers (or cache), which are used to speed-up access to the medium (which can be considerably slower than, e.g., Dynamic Random Access Memory, DRAM). For example, 3D XPoint-based Intel Optane DC Persistent Memory Module, PMM, features 16 kilobyte (kB) buffer for each memory module (Dual In-line Memory Module, DIMM). The buffers are managed by a separate memory controller located on the memory module itself (denoted MMC) and are transparent to computer CPUs and their cache controllers.

[0007] The fast buffers located on a memory module help to reduce the latency of memory accesses because computer programs typically exhibit temporal and spatial locality in the memory access patterns. As such, the cost of accessing slow memory medium can be amortised over multiple read operations. Similarly, multiple write operations can be coalesced in the buffer, so that costly write operations to the medium do not happen upon every write operation issued by the CPU memory controller to the memory module. Due to technical reasons, accesses to the memory medium sometimes are performed in blocks (called Error Correction Code, ECC, blocks), which are larger than the size of a single cache line. For example, Intel Optane DC PMM internally uses 256B blocks, which are four times larger than a typical cache line. Some of the future memory expansion cards by various vendors, such as Samsung, Panasonic, HP, will also internally use ECC blocks that are larger than 64B (so, e.g., 128B, 256B, 512B, etc.). In such case, the buffers located on a memory module aid the MMC in servicing smaller, cache line-sized read and write operations. To this end, the buffers store the entire ECC blocks, but allow the MMC to read and write portions of the ECC blocks, which correspond to single cache lines.

[0008] The overall costs (expressed in terms of, e.g., latency and bandwidth) of memory accesses heavily depend, among others, on the relative speed of read / write operations to the buffer and to the memory medium, the size of the buffer, the buffer eviction policy (i.e., which ECC block should be evicted to the memory medium to make room in the buffer for an ECC block needed to service the incoming read or write operation) as well as memory access patterns, which are application dependent.

[0009] In particular, random memory access patterns that involve reads and writes scattered across very different memory locations are problematic as they incur high read and write amplification on the memory medium:

[0010] • to service a read of a single cache line, the entire ECC block needs to be read from the memory medium and placed in the buffer,

[0011] • to service a write of a single cache line, the entire ECC block needs to be read from the memory medium, placed in the buffer, and modified in part according to the write operation,

[0012] • each ECC block in the buffer is only rarely reused to service later read / write operations.

[0013] High read and write amplification not only can significantly reduce the useable memory bandwidth and increase the latency of memory accesses. In case of some memory media, such as Phase Change Memory, PCM, which have lower write endurance compared to, e.g., DRAM, high write amplification can cause an increased wear of the memory medium, and, in consequence, shorten the memory module lifespan. SUMMARY

[0014] An objective of the present disclosure is to improve the management of a memory module’s buffer through coordination with the CPU cache.

[0015] The foregoing and other objectives are achieved by the features of the independent claims.

[0016] Further implementation forms are apparent from the dependent claims, the description and the Figures.

[0017] A first aspect of the present disclosure provides a method for regulating operations between a central processing unit (CPU) cache and a buffer of a memory module of an apparatus, the method comprising placing, by the memory module, a memory block comprising multiple cache lines for the apparatus into the buffer of the memory module, querying, by the memory module, the CPU cache to obtain a plurality of modified cache lines associated with the memory block, and updating the multiple cache lines in the memory block based on the obtained plurality of modified cache lines, wherein the querying occurs before eviction and / or a writeback of the memory block from the buffer of the memory module of the apparatus.

[0018] Importantly, the utilisation of memory read / write bandwidth can be improved, increasing performance. In addition, by reducing the number of operations, memory medium wear can be reduced, as the memory medium wear is closely associated with the utilised memory write bandwidth. Finally, the invention offers a reduction of the latency of the eviction of some cache lines from the CPU, thus further increasing application performance.

[0019] The method may further comprise performing, using the CPU cache, an operation on a plurality of cache lines in the CPU cache, wherein the plurality of cache lines in the CPU cache is associated with the memory block, whereby to acquire the plurality of modified cache lines.

[0020] In response to performing, using the CPU cache, the operation on the plurality of cache lines in the CPU cache, the method may further comprise changing, by the CPU cache, a state of the plurality of cache lines in the CPU cache to thereby indicate that the plurality of cache lines in the CPU cache has been modified.

[0021] In response to detecting, by the memory module, the change of state of the multiple cache lines in the CPU cache, the method may further comprise issuing, by the memory module, a request to fetch the plurality of modified cache lines from the CPU cache. The method may further comprise assigning, by the memory module, respective unique states to a plurality of the multiple cache lines comprised in the memory block in the buffer of the memory module, wherein the unique state is used only by the memory module and not by the CPU cache, wherein the unique state indicates that the state of the plurality of the multiple cache lines is not known to the memory module, communicating, with the CPU cache, to determine a current state of each of the plurality of modified cache lines in the CPU cache, and changing, by the memory module, respective states of those cache lines of the multiple cache lines corresponding to the plurality of modified cache lines from the unique state based on the current states of each of the plurality of modified cache lines.

[0022] The method may further comprise, before the eviction and / or the write-back of the memory block from the buffer of the memory module, querying, by the memory module, the CPU cache to determine a current state of those cache lines of the multiple cache lines whose state has not been changed from the unique state.

[0023] A second aspect of the present disclosure provides an apparatus comprising a memory module comprising a buffer and a plurality of memory blocks, and a central processing unit (CPU) cache, wherein the memory module is arranged to place a memory block comprising multiple cache lines for the apparatus into the buffer of the memory module, query the CPU cache to obtain a plurality of modified cache lines associated with the memory block, and update the multiple cache lines in the memory block based on the obtained plurality of modified cache lines, wherein the querying occurs before eviction and / or a write-back of the memory block from the buffer of the memory module of the apparatus.

[0024] The CPU cache may be further arranged to perform an operation on a plurality of cache lines in the CPU cache, wherein the plurality of cache lines in the CPU cache is associated with the memory block, whereby to acquire the plurality of modified cache lines.

[0025] The CPU cache may be further arranged to, in response to performing the operation on the plurality of cache lines in the CPU cache, change a state of plurality of cache lines to thereby indicate that plurality of cache lines has been modified.

[0026] The memory module may be further arranged to detect the change of state of the multiple cache lines in the CPU cache and issue a request to fetch the plurality of modified cache lines from the CPU cache. The memory module may be further arranged to assign respective unique states to a plurality of the multiple cache lines comprised in the memory block in the buffer of the memory module, wherein the unique state is used only by the memory module and not by the CPU cache, wherein the unique state indicates that the state of the plurality of the multiple cache lines is not known to the memory module, communicate with the CPU cache to determine a current state of each of the plurality of modified cache lines in the CPU cache, and change respective states of those cache lines of the multiple cache lines corresponding to the plurality of modified cache lines from the unique state based on the current states of each of the plurality of modified cache lines.

[0027] The memory module may be further arranged to, before the eviction and / or the write-back of the memory block from the buffer of the memory module, query the CPU cache to determine a current state of those cache lines of the multiple cache lines whose state has not been changed from the unique state.

[0028] A third aspect of the present disclosure provides a computer readable storage medium comprising computer program code, accessible by an apparatus comprising a processor, to provide instructions and / or data to the apparatus, the computer program code configured to, with the processor, cause the apparatus to place, by a memory module, a memory block comprising multiple cache lines for the apparatus into the buffer of the memory module, query, by the memory module, a CPU cache to obtain a plurality of modified cache lines associated with the memory block, and update the multiple cache lines in the memory block based on the obtained plurality of modified cache lines, wherein the querying occurs before eviction and / or a write-back of the memory block from the buffer of the memory module of the apparatus.

[0029] The computer readable storage medium may further comprise the computer program code configured to, with the processor, cause the apparatus to perform, using the CPU cache, an operation on a plurality of cache lines in the CPU cache, wherein the plurality of cache lines in the CPU cache is associated with the memory block, whereby to acquire the plurality of modified cache lines.

[0030] The computer readable storage medium may further comprise the computer program code configured to, with the processor, cause the apparatus to, in response to performing, using the CPU cache, the operation on the plurality of cache lines in the CPU cache, change, by the CPU cache, a state of the plurality of cache lines in the CPU cache to thereby indicate that the plurality of cache lines in the CPU cache has been modified. These and other aspects of the invention will be apparent from the embodiment s) described below.

[0031] BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order that the present invention may be more readily understood, embodiments of the invention will now be described, by way of example, with reference to the accompanying drawings, in which:

[0033] Figs, la- Im are a step-by-step depiction of a read-modify-write operation according to the prior art;

[0034] Fig. 2 is a flow chart of a method for regulating operations between a CPU cache and a buffer of a memory module of an apparatus according to an example;

[0035] Figs. 3a-3g are a step-by-step depiction of a read-modify-write operation according to an example;

[0036] Fig. 4 is a schematic depiction of an example communication between the CPU and the memory module according to an example;

[0037] Fig. 5 is a schematic depiction of an example communication between the CPU and the memory module according to another example;

[0038] Fig. 6 is a schematic depiction of an example communication between the CPU and the memory module according to yet another example; and

[0039] Fig. 7 is a schematic depiction of an apparatus according to an example.

[0040] DETAILED DESCRIPTION

[0041] Example embodiments are described below in sufficient detail to enable those of ordinary skill in the art to embody and implement the systems and processes herein described. It is important to understand that embodiments can be provided in many alternate forms and should not be construed as limited to the examples set forth herein.

[0042] Accordingly, while embodiments can be modified in various ways and take on various alternative forms, specific embodiments thereof are shown in the drawings and described in detail below as examples. There is no intent to limit to the particular forms disclosed. On the contrary, all modifications, equivalents, and alternatives falling within the scope of the appended claims should be included. Elements of the example embodiments are consistently denoted by the same reference numerals throughout the drawings and detailed description where appropriate.

[0043] The terminology used herein to describe embodiments is not intended to limit the scope. The articles “a,” “an,” and “the” are singular in that they have a single referent, however the use of the singular form in the present document should not preclude the presence of more than one referent. In other words, elements referred to in the singular can number one or more, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises,” “comprising,” “includes,” and / or “including,” when used herein, specify the presence of stated features, items, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, items, steps, operations, elements, components, and / or groups thereof.

[0044] Unless otherwise defined, all terms (including technical and scientific terms) used herein are to be interpreted as is customary in the art. It will be further understood that terms in common usage should also be interpreted as is customary in the relevant art and not in an idealized or overly formal sense unless expressly so defined herein.

[0045] Typically, CPU caches help a lot in reducing the required memory bandwidth, reducing the latency of memory accesses, and similar. Unfortunately, in some cases, the way in which data is managed in CPU caches may contribute to the reduction in the useable memory bandwidth and increases the latency of memory accesses, as well as shortens the memory module lifespan. In particular, issues may arise due to the fact that the CPU caches are not synchronised with the buffer of the memory module, causing the abovementioned issues. More precisely, each cache line in CPU caches is considered independent and thus is fetched and evicted by the CPU cache controller as needed by issuing read and write operations that typically concern the entire cache lines. There is no additional coordination between the CPU cache controller and the MMC in regard to the prefetching ECC blocks to the buffer and evicting them to the memory medium later on. In contrary, the multi-level CPU caches of modern computers are always tightly coordinated. In principle, random memory accesses that concern objects that occupy at most one cache line will incur read and write amplification that directly corresponds to the ratio of the size of an ECC block on the memory module and the size of a single cache line. However, even for objects that occupy multiple (i.e., consecutive) cache lines or for objects allocated in memory next to each other and often accessed at the same time or in quick succession (as applications typically exhibit temporal and spatial locality), the read and write amplification is still high. It is because the CPU cache controller considers all cache lines independent, even though they might belong to the same ECC block and be modified together during computation. Moreover, CPU caches are typically much larger compared to the size of the buffers on-board the memory modules (tens of MB vs tens of kB). Hence, an ECC block is typically evicted from the memory module buffer well before the cache lines that are part of the ECC block are evicted from the CPU caches. This increases the read and write amplification even further.

[0046] Figs, la- Im are a step-by-step depiction of a read-modify-write operation according to the prior art. In Fig. la, a memory block #42 (e.g., ECC block) is already present in the buffer 130 of the memory module 120 (i.e., it has been fetched into the buffer 130 from memory 140 of the memory module, the memory 140 comprising a plurality of memory blocks), and the cache lines #42a and #42c are already present in a CPU cache 111 of the CPU 110. The CPU 110 performs a read-modify-write operation with a random-access pattern and modifies two cache lines present in the CPU cache 111 - #42c and #42a - at the same time or in quick succession. In particular, in Fig. la, the CPU 110 reads (depicted as arrows pointing from the cache lines) the cache lines #42c and #42a from its cache 111. In Fig. lb, the CPU modifies the previously read cache lines and writes (depicted as arrows pointing to the cache lines) the modified cache lines #42c and #42a to its cache 111. In the figures, the modified cache lines are depicted as having a shaded background.

[0047] In Fig. 1c, the memory block #42 (containing the cache lines #42a and #42c) is evicted from the buffer 130 of the memory module 120. The memory block #42 was initially placed in the buffer 130 to serve the read operations performed by the CPU 110. The memory block #42 is evicted from the buffer 130 early, because the buffer 130 is much smaller than the CPU cache 111. In Fig. Id, the CPU evicts the modified cache line #42c from its cache 111 and writes the modified cache line #42c to the memory module 120. However, eviction of each of the modified cache lines #42a and #42c from the CPU cache 111 requires fetching the memory block #42 into the buffer 130 again. As such, in Fig. le, the memory block #42 is fetched from the memory 140 into the buffer 130 again, in order to write the modified cache line #42c thereto. In Fig. If, the modified cache line #42c is updated in the buffer 130. Fig. 1g shows the modified cache line #42c now present in the memory block #42 in the buffer 130. In Fig. Ih, the memory block #42 is once again evicted from the buffer 130 and replaced by some other memory block. In Fig. li, the CPU 110 evicts the remaining modified cache line #42a from its cache 111, and writes the modified cache line #42a to the memory module 120. In Fig. Ij, the memory block #42 is fetched into the buffer 130 again, similar to step Fig. le. In Fig. Ik, the modified cache line #42a is written to the buffer 130. Fig. 11 shows the modified cache line #42a now present in the memory block #42 in the buffer 130. Finally, in Fig. Im, the memory block #42c is once again evicted from the buffer 130.

[0048] In other words, even though the modifications of two cache lines (64B each) that belong to the same memory block (256B) happened at the same time or in quick succession, two additional reads from the memory medium were required, thus requiring 512B of additional read bandwidth (768B in total, including the initial read, omitted from Figs, la-lm for brevity). Additionally, two writes to the memory medium were required, requiring 512B of bandwidth.

[0049] According to an example, there is provided a mechanism to reduce the number of read / write operations from / to the memory medium that happen, in particular, with random memory access patterns. Aspects of the invention relate to allowing some caches to be evicted more quickly and efficiently from the CPU cache. In turn, the invention increases the usable read / write bandwidth of the memory module, reduces the memory medium wear, and makes CPU cache management more efficient. Moreover, the improvement in utilisation of memory read / write bandwidth enables the use of larger memory (ECC) blocks to store data in the memory medium without degrading the performance. This, in turn, improves the effective capacity of the memory module.

[0050] The invention concerns computer systems that feature memory modules with on-board buffers. The memory modules can be based on different media (e.g., DRAM, PCM, 3D XPoint, MRAM, ReRAM) and can be either directly attached to the CPU’s memory controller, or be provided as separate devices and connected to the computer through a fast interconnect, e.g., based on the Compute Express Link (EXL) standard. In particular, the invention affects the memory module controller aboard the memory modules, as well as the CPU cache controller.

[0051] Examples in the present disclosure can be provided as methods, systems or machine-readable instructions, such as any combination of software, hardware, firmware or the like. Such machine-readable instructions may be included on a computer readable storage medium (including but not limited to disc storage, CD-ROM, optical storage, etc.) having computer readable program codes therein or thereon.

[0052] The present disclosure is described with reference to flow charts and / or block diagrams of the method, devices and systems according to examples of the present disclosure. Although the flow diagrams described above show a specific order of execution, the order of execution may differ from that which is depicted. Blocks described in relation to one flow chart may be combined with those of another flow chart. In some examples, some blocks of the flow diagrams may not be necessary and / or additional blocks may be added. It shall be understood that each flow and / or block in the flow charts and / or block diagrams, as well as combinations of the flows and / or diagrams in the flow charts and / or block diagrams can be realized by machine readable instructions.

[0053] The machine-readable instructions may, for example, be executed by a machine such as a general-purpose computer, user equipment such as a smart device, e.g., a smart phone, a special purpose computer, an embedded processor or processors of other programmable data processing devices to realize the functions described in the description and diagrams. In particular, a processor or processing apparatus may execute the machine-readable instructions. Thus, modules of apparatus (for example, a module implementing a comparator unit, or a firewall structure and so on) may be implemented by a processor executing machine readable instructions stored in a memory, or a processor operating in accordance with instructions embedded in logic circuitry. The term 'processor' is to be interpreted broadly to include a CPU, processing unit, ASIC, logic unit, or programmable gate set etc. The methods and modules may all be performed by a single processor or divided amongst several processors.

[0054] Such machine-readable instructions may also be stored in a computer readable storage that can guide the computer or other programmable data processing devices to operate in a specific mode. For example, the instructions may be provided on a non-transitory computer readable storage medium encoded with instructions, executable by a processor.

[0055] Fig. 2 is a flow chart of a method for regulating operations between a CPU cache and a buffer of a memory module of an apparatus according to an example. The method comprises, in block 201, placing, by the memory module, a memory block comprising multiple cache lines for the apparatus into the buffer of the memory module. In block 202, the memory module queries the CPU cache to obtain a plurality of modified cache lines associated with the memory block. In order to acquire the plurality of modified cache lines associated with the memory block that is in the buffer of the memory module, the CPU may perform an operation on a plurality of cache lines in its cache. For example, the CPU may modify a cache line by writing to it, thereby creating a modified cache line. The modified cache line may then need to be written to the buffer of the memory module in order to enable the memory module to store its value in its memory.

[0056] Advantageously, the memory module can fetch the modified cache lines from the CPU caches prior to the eviction of the ECC block which encompasses the cache lines from the buffer of the memory module, thereby avoiding the multiple eviction scenario depicted in Figs, la-lm. In block 203, the method comprises updating the multiple cache lines in the memory block based on the obtained plurality of modified cache lines, wherein the querying occurs before eviction and / or a write-back of the memory block from the buffer of the memory module of the apparatus.

[0057] To better illustrate this, reference is made to Figs. 3a-3g, which are a step-by-step depiction of a read-modify-write operation according to an example. Figs. 3a and 3b correspond to Figs, la and lb, respectively, with no changes to the CPU / memory module behaviour. Same elements in Figs, la-lm and Figs. 3a-3g are denoted using the same reference numerals and function likewise.

[0058] In Fig. 3a, a memory block #42 (e.g., ECC block) is already present in the buffer 130 of the memory module 120, and the cache lines #42a and #42c are already present in a CPU cache 111 of the CPU 110. The CPU 110 performs a read-modify-write operation with a randomaccess pattern and modifies two cache lines present in the CPU cache 111 - #42c and #42a - at the same time or in quick succession. In particular, in Fig. 3a, the CPU 110 reads (depicted as arrows pointing from the cache lines) the cache lines #42c and #42a from its cache 111. In Fig. 3b, the CPU modifies the previously read cache lines and writes (depicted as arrows pointing to the cache lines) the modified cache lines #42c and #42a to its cache 111. In Fig. 3c, just before the eviction of the memory block #42 from the memory, the memory module 120 queries the CPU cache 111 for all cache lines that belong to the memory block #42 which have not yet been written-back to the memory. In Fig. 3d, the modified cache lines #42a and #42c are written to the buffer 130. In Fig. 3e, the CPU cache marks the modified cache lines #42a and #42c as written back to the memory module, depicted in Fig. 3e using a different shade of grey. In Fig. 3f, the memory block #42 is evicted from the buffer and replaced by some other memory block. Finally, in Fig. 3g, whenever the cache controller decides to evict the modified cache lines #42a and #42c from the CPU cache 111 in order to make room for other cache lines, it can do so easily, as they have already both been written back to the memory module 120. The eviction of the cache lines in question can happen independently for each of the cache lines #42a and #42c. In other words, compared to the prior art solution, the number of read / write from / to the memory module can be reduced.

[0059] In order to be able to choose the quick path during cache line eviction, the CPU should know which cache lines have already been written back to the memory module. This capability may be realised by a CPU cache controller and a cache coherence protocol (e.g., MESI) that maintains a consistent state of all cache lines across all CPU caches. For example, using MESI: a cache line (of an address given) freshly fetched from memory may be marked in the CPU cache as Exclusive (E). The CPU’s write and read requests to this cache lines can be served without additional prior synchronisation with other CPU caches; when some other CPU wants to read a cache line of the same address, the cache line in the Exclusive state may be copied to the local CPU cache and then used to service the CPU’s request. Both copies of the cache line may now be marked as Shared (S); when a CPU modifies a cache line, the cache line may then be marked as Modified (M) in the CPU cache. At the same time, all cache lines of the same address in other CPU caches may be invalidated (i.e., marked as Invalid, I); in response to a read operation being performed on a cache line in an Invalid state, the local copy of the cache line may need to be updated. This can be achieved by copying the most recent value of the cache line in question from some other CPU cache. Both copies of the cache line may now be marked as Shared (S) in the CPU caches.

[0060] If a cache line is marked as Exclusive or Shared in any of the CPU caches, it means that the memory holds the most recent (up-to-date) value of the cache line.

[0061] When the CPU cache controller starts using a cache line originating from the memory module, it may monitor the coherence of the cache line. If any cache line is marked as Modified in any of the CPU caches, this can indicate that the memory module does not have the most recent version of this cache line. Just prior to eviction of a memory block from the buffer of the memory module, the memory module may issue a special request to the CPU cache controller to fetch all cache lines that belong to the memory block to be evicted. In particular, the memory module may fetch all modified (i.e., the cache lines not marked as Exclusive or Shared) cache lines that belong to the memory block to be evicted.

[0062] This process is illustrated in Fig. 4, which is a schematic depiction of an example communication between the CPU and the memory module according to an example. In the example of Fig. 4, the CPU performs a read operation on a cache line #42c. The CPU may check whether the cache line #42c is present in the CPU cache. If the cache line #42c is not in the local cache, the CPU cache may indicate that the requested cache line is not present and attempt to read the cache line #42c from the buffer of the memory module. As the buffer may be initially empty, the requested cache line may not be present in the buffer, in response to which the buffer may fetch the memory block associated with the requested cache line (e.g., memory block #42 for the cache line #42c) into the buffer. The buffer may then send the requested cache line #42c to the CPU cache. The CPU cache (e.g., the CPU cache controller) may discover that the state of the cache line #42c is Exclusive, i.e., that the cache line has been freshly fetched from the memory module. The CPU may then modify the cache line (e.g., by performing a write operation thereon), after which the state of the cache line will be changed to Modified. The CPU may then perform a read operation on a cache line #42a, associated with the same memory block #42 as the cache line #42c. Again, the CPU may attempt to read the cache line #42a from its cache, and, in response to the cache line #42a not being present in the CPU cache, the CPU cache may attempt to read the cache line #42a from the buffer of the memory module. As the memory block #42 has already been placed in the buffer as part of the operations being performed on cache line #42c, the cache line #42a may be fetched into the CPU cache from the buffer of the memory module, marked as Exclusive. Prior to evicting the memory block #42 from its buffer, the memory module may issue a special request to the CPU cache to fetch all modified cache lines belonging to memory block #42. Since the state of the cache line #42a is Exclusive, it will not be fetched into the buffer, in contrast to line #42c marked a Modified. The CPU cache may write the cache line #42c to the buffer of the memory module and change the state of the cache line #42c to Shared. After the cache line #42c has been updated in the buffer, the memory block #42 may then be written to the memory. That is, advantageously, since the cache line #42c has already been written back, its eviction can be performed quickly, without fetching the memory block #42 associated with the cache line into the buffer of the memory module again, as is the case with the solution proposed by the prior art. In an embodiment, the memory module (including the buffer) may be included in the CPU cache coherence domain. This is illustrated in Fig. 5, which is a schematic depiction of an example communication between the CPU and the memory module according to another example. The operations shown in Fig. 5 are largely similar to those described in relation to Fig. 4. As discussed above, in typical conditions, a cache line that has been fetched from the memory medium and placed in the buffer would be marked as Exclusive. However, if a cache line of the same address already exists in some other CPU cache, the cache line fetched from the memory must have a different state assigned in order to preserve cache coherence. Therefore, a query on the cache line state in the CPU cache may be required. The earlier request to read the cache line #42c (i.e., BusRd(S) #42c in the Figure) may indicate that there is no copy of #42c in any of the CPU’s caches. Once the CPU modifies the cache line #42c, the copy of cache line #42c in the CPU cache may change state to Modified. At the same time the copy of cache line #42c present in the buffer may change the state Invalid to preserve CPU cache coherence. When the memory module wants to bring up-to-date a cache line which is present in its buffer and marked as Invalid - i.e., in the example of Fig. 5, the cache line #42c - the memory module may use the cache coherence protocol and issue the request to read (BusRd(s)) for the cache line #42c, as depicted in the Figure. Consequently, the most recent value of the cache line #42c may be copied to the buffer from some other CPU cache, and both copies of the cache line #42c may be marked as Shared. For cache lines that belong to the same memory block, the memory module may discover their states in the CPU cache (or fetch them from the CPU cache) by issuing separate requests for each cache line, or by issuing a bulk request for all the cache lines.

[0063] Fig. 6 is a schematic depiction of an example communication between the CPU and the memory module according to yet another example. Here, compared to Fig. 5, an optimisation has been introduced to eschew the query of cache lines states which occurs when loading a memory block into the buffer. As discussed above in relation to Fig. 5, the earlier request to read the cache line #42c (i.e., BusRd(S) #42c) may indicate that there is no copy of #42c in any of the CPU’s caches. However, advantageously, the discovery of the coherence states of the remaining cache lines associated with the same memory block as #42c (i.e., cache lines associated with memory block #42) can be deferred until later, that is, just prior to the eviction of the memory block from the buffer of the memory module. In order to achieve that, a new unique state may be introduced, to be used only be the memory module and not by the CPU cache (that is, the CPU cache will not assign this state to any of the cache lines). In the Fig. 6, the unique state is referred to as Unknown (?), but the skilled person would appreciate that any unique state to be used by the memory module only could be used. Occasionally, the memory module will learn about the state of some cache line associated with one of the memory blocks in the buffer of the memory module in the course of normal operation. For example, as discussed before, a request to read the cache line #42a may indicate that there is no copy of #42a in any of the CPU’s caches, indicating to the memory module that it has the exclusive copy of the cache line #42a in its buffer. Subsequently, the memory module may downgrade the state of #42a to Shared once the cache line is copied into the CPU cache.

[0064] Finally, just prior to the eviction of the memory block #42 from the buffer, the memory module may attempt to fetch from the CPU cache(s) all the cache lines, associated with the memory block to be evicted, that are marked as Invalid, or their state is Unknown in the buffer. As before, no response for some of the cache lines may indicate that the buffer holds an exclusive copy of the cache line.

[0065] Fig. 7 is a schematic depiction of an apparatus according to an example. The apparatus 700 comprises a memory module 720 comprising a buffer 730 and a plurality of memory blocks. The apparatus 700 further comprises a central processing unit (CPU) 710, comprising a CPU cache 711. The memory module 720 is configured to place a memory block of the plurality of memory blocks into its buffer 730. The memory module 720 may further comprise a memory 740 where the plurality of memory blocks is stored. The memory block comprises multiple cache lines for the apparatus 700. The memory module 720 is further configured to query the CPU cache 711 to obtain a plurality of modified cache lines associated with the memory block in the buffer 730, and update the multiple cache lines in the memory block based on the plurality of modified cache lines, wherein the querying occurs before an eviction and / or a write-back of the memory block from the buffer 730 of the memory module 720 of the apparatus 700.

[0066] According to an example, machine-readable instructions can be loaded onto a computer or other programmable data processing devices, so that the computer or other programmable data processing devices perform a series of operations to produce computer-implemented processing, thus the instructions executed on the computer or other programmable devices provide an operation for realizing functions specified by flow(s) in the flow charts and / or block(s) in the block diagrams. Further, the teachings herein may be implemented in the form of a computer or software product, such as a non-transitory machine-readable storage medium, the computer software or product being stored in a storage medium and comprising a plurality of instructions, e.g., machine readable instructions, for making a computer device implement the methods recited in the examples of the present disclosure.

[0067] In some examples, some methods can be performed in a cloud-computing or network-based environment. Cloud-computing environments may provide various services and applications via the Internet. These cloud-based services (e.g., software as a service, platform as a service, infrastructure as a service, etc.) may be accessible through a web browser or other remote interface of the user equipment for example. Various functions described herein may be provided through a remote desktop environment or any other cloud-based computing environment.

[0068] While various embodiments have been described and / or illustrated herein in the context of fully functional computing systems, one or more of these exemplary embodiments may be distributed as a program product in a variety of forms, regardless of the particular type of computer- readable-storage media used to actually carry out the distribution. The embodiments disclosed herein may also be implemented using software modules that perform certain tasks. These software modules may include script, batch, or other executable files that may be stored on a computer-readable storage medium or in a computing system. In some embodiments, these software modules may configure a computing system to perform one or more of the exemplary embodiments disclosed herein. In addition, one or more of the modules described herein may transform data, physical devices, and / or representations of physical devices from one form to another.

[0069] The preceding description has been provided to enable others skilled in the art to best utilize various aspects of the exemplary embodiments disclosed herein. This exemplary description is not intended to be exhaustive or to be limited to any precise form disclosed. Many modifications and variations are possible without departing from the spirit and scope of the instant disclosure. The embodiments disclosed herein should be considered in all respects illustrative and not restrictive. Reference should be made to the appended claims and their equivalents in determining the scope of the instant disclosure.

Claims

CLAIMS1. A method for regulating operations between a central processing unit, CPU, cache and a buffer of a memory module of an apparatus, the method comprising: placing, by the memory module, a memory block comprising multiple cache lines for the apparatus into the buffer of the memory module (201); querying, by the memory module, the CPU cache to obtain a plurality of modified cache lines associated with the memory block (202); and updating the multiple cache lines in the memory block based on the obtained plurality of modified cache lines (203), wherein the querying occurs before eviction and / or a writeback of the memory block from the buffer of the memory module of the apparatus.

2. The method of claim 1, further comprising: performing, using the CPU cache, an operation on a plurality of cache lines in the CPU cache, wherein the plurality of cache lines in the CPU cache is associated with the memory block, whereby to acquire the plurality of modified cache lines.

3. The method as claimed in claim 2, further comprising: in response to performing, using the CPU cache, the operation on the plurality of cache lines in the CPU cache, changing, by the CPU cache, a state of the plurality of cache lines in the CPU cache to thereby indicate that the plurality of cache lines in the CPU cache has been modified.

4. The method as claimed in claim 3, further comprising: in response to detecting, by the memory module, the change of state of the multiple cache lines in the CPU cache, issuing, by the memory module, a request to fetch the plurality of modified cache lines from the CPU cache.

5. The method as claimed in claim 3 or 4, further comprising: assigning, by the memory module, respective unique states to a plurality of the multiple cache lines comprised in the memory block in the buffer of the memory module, wherein the unique state is used only by the memory module and not by the CPU cache, wherein the unique state indicates that the state of the plurality of the multiple cache lines is not known to the memory module; communicating, with the CPU cache, to determine a current state of each of the plurality of modified cache lines in the CPU cache; and changing, by the memory module, respective states of those cache lines of the multiple cache lines corresponding to the plurality of modified cache lines from the unique state based on the current states of each of the plurality of modified cache lines.

6. The method as claimed in claim 5, further comprising: before the eviction and / or the write-back of the memory block from the buffer of the memory module, querying, by the memory module, the CPU cache to determine a current state of those cache lines of the multiple cache lines whose state has not been changed from the unique state.

7. An apparatus (700), comprising: a memory module (720) comprising a buffer (730) and a plurality of memory blocks; and a central processing unit, CPU, cache (711), wherein the memory module (720) is arranged to: place a memory block comprising multiple cache lines for the apparatus into the buffer (730) of the memory module (720); query the CPU cache (711) to obtain a plurality of modified cache lines associated with the memory block; and update the multiple cache lines in the memory block based on the obtained plurality of modified cache lines, wherein the querying occurs before eviction and / or a write-back of the memory block from the buffer (730) of the memory module (720) of the apparatus.

8. The apparatus (700) as claimed in claim 7, wherein the CPU cache (711) is further arranged to: perform an operation on a plurality of cache lines in the CPU cache (711), wherein the plurality of cache lines in the CPU cache is associated with the memory block, whereby to acquire the plurality of modified cache lines.

9. The apparatus (700) as claimed in claim 8, wherein the CPU cache (711) is further arranged to: in response to performing the operation on the plurality of cache lines in the CPU cache (711), change a state of plurality of cache lines to thereby indicate that plurality of cache lines has been modified.

10. The apparatus (700) as claimed in claim 9, wherein the memory module (720) is further arranged to: detect the change of state of the multiple cache lines in the CPU cache (711) and issue a request to fetch the plurality of modified cache lines from the CPU cache (711).

11. The apparatus (700) as claimed in claim 9 or 10, wherein the memory module (720) is further arranged to: assign respective unique states to a plurality of the multiple cache lines comprised in the memory block in the buffer (730) of the memory module (720), wherein the unique state is used only by the memory module (720) and not by the CPU cache (711), wherein the unique state indicates that the state of the plurality of the multiple cache lines is not known to the memory module (720); communicate with the CPU cache (711) to determine a current state of each of the plurality of modified cache lines in the CPU cache (711); andchange respective states of those cache lines of the multiple cache lines corresponding to the plurality of modified cache lines from the unique state based on the current states of each of the plurality of modified cache lines.

12. The apparatus as claimed in claim 11, wherein the memory module (720) is further arranged to: before the eviction and / or the write-back of the memory block from the buffer (730) of the memory module, query the CPU cache (711) to determine a current state of those cache lines of the multiple cache lines whose state has not been changed from the unique state.

13. A computer readable storage medium comprising computer program code, accessible by an apparatus comprising a processor, to provide instructions and / or data to the apparatus, the computer program code configured to, with the processor, cause the apparatus to: place, by a memory module, a memory block comprising multiple cache lines for the apparatus into the buffer of the memory module; query, by the memory module, a CPU cache to obtain a plurality of modified cache lines associated with the memory block; and update the multiple cache lines in the memory block based on the obtained plurality of modified cache lines, wherein the querying occurs before eviction and / or a write-back of the memory block from the buffer of the memory module of the apparatus.

14. The computer readable storage medium of claim 13, further comprising the computer program code configured to, with the processor, cause the apparatus to: perform, using the CPU cache, an operation on a plurality of cache lines in the CPU cache, wherein the plurality of cache lines in the CPU cache is associated with the memory block, whereby to acquire the plurality of modified cache lines.

15. The computer readable storage medium of claim 14, further comprising the computer program code configured to, with the processor, cause the apparatus to: in response to performing, using the CPU cache, the operation on the plurality of cache lines in the CPU cache, change, by the CPU cache, a state of the plurality of cache lines in the CPU cache to thereby indicate that the plurality of cache lines in the CPU cache has been modified.

Citation Information

Patent Citations

  • Method and apparatus related to cache memory

    US20150019823A1

  • Rinsing cache lines from a common memory page to memory

    US20190179770A1

  • Method, device and computer program product for cache management

    US20200133870A1

  • Translation cache and configurable ECC memory for reducing ECC memory overhead

    US20220012126A1